From b7778c287b54f9940c44ef0756d97acf3f05444a Mon Sep 17 00:00:00 2001 From: root Date: Thu, 20 Aug 2026 07:52:44 +0000 Subject: [PATCH] =?UTF-8?q?Item=204:=20LiteLLM=203-way=20cap=20split=20?= =?UTF-8?q?=E2=80=94=20document=20context=20caps?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - litellm-self-heal.prose.md: Add 3-way cap split (Abiba/Hermes=128K uncapped, Crew=64K via crew-auto alias) - Note: Do NOT document max_input_tokens as context window — use max_tokens/context_window_size --- litellm-self-heal.prose.md | 10 +++++++++- 1 file changed, 9 insertions(+), 1 deletion(-) diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 6e95cef..c7680a8 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -68,7 +68,15 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) ## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116) -`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`. +`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20). + +### Context Cap Split (2026-08-20) + +- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads +- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped +- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit + +Preferred implementation: uncap shared pool, add capped alias for crew-only. - `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200). - `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively.