diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 6e95cef..c7680a8 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -68,7 +68,15 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) ## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116) -`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`. +`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20). + +### Context Cap Split (2026-08-20) + +- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads +- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped +- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit + +Preferred implementation: uncap shared pool, add capped alias for crew-only. - `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200). - `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.