Item 4: LiteLLM 3-way cap split — document context caps
- litellm-self-heal.prose.md: Add 3-way cap split (Abiba/Hermes=128K uncapped, Crew=64K via crew-auto alias) - Note: Do NOT document max_input_tokens as context window — use max_tokens/context_window_size
This commit is contained in:
@@ -68,7 +68,15 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|||||||
|
|
||||||
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
|
## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116)
|
||||||
|
|
||||||
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`.
|
`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20).
|
||||||
|
|
||||||
|
### Context Cap Split (2026-08-20)
|
||||||
|
|
||||||
|
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
|
||||||
|
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
|
||||||
|
- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit
|
||||||
|
|
||||||
|
Preferred implementation: uncap shared pool, add capped alias for crew-only.
|
||||||
|
|
||||||
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
|
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200).
|
||||||
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
|
- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.
|
||||||
|
|||||||
Reference in New Issue
Block a user