From baaac9d7c60be414706c480ea47affa21b87abfb Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 15:43:21 +0000 Subject: [PATCH] no-mistakes(review): delete remaining duplicated timeout and frozen model-list state --- gpu-fleet.prose.md | 5 ++--- litellm-health.prose.md | 18 +++++++++++------- litellm-self-heal.prose.md | 4 ++++ 3 files changed, 17 insertions(+), 10 deletions(-) diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index cebe3b4..04cd136 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -217,6 +217,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions. - **Alert migration**: All alerts now go to `#agent-hub` topics (`alerts-gpu`, `alerts-pm2`, `alerts-infra`) instead of DMs. Cross-agent visibility enabled. - **tok/s benchmarks**: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds. - **NetBird 502**: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down. +- **Alias-retirement follow-up (2026-09-12)**: the agent-facing templates (`hermes-config-template.prose.md`, `hermes-agent-baseline.prose.md`), `litellm-api-keys.prose.md`, and the executable `audit-hermes-config.py` still reference the retired `gpu-light`/`gemma-4-12b` names, and `audit-hermes-config.py` Rule 8 currently fails a config whose vision/web_extract model is `gpu-vision`. These are tracked separately and are intentionally NOT updated in this change. ## GPU Inference Benchmarks (Current) @@ -251,7 +252,7 @@ When the underlying model is swapped, only the LiteLLM config changes — agent - **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek) - Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling) - **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K) -- Mumuni compression model alias: `strix-moe` with 300s timeout +- Mumuni compression model alias: `strix-moe` ### Mumuni Agent Profile @@ -273,8 +274,6 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the | `memory.memory_char_limit` | 800 | Brief memory entries | | `personalities` | `creative` | Creative assistant personality | | Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms | -| Main model timeout | 300s | LiteLLM global timeout | -| Compression model timeout | 300s | strix-moe timeout increased from 120s | ### Agent Update Status (2026-07-15) diff --git a/litellm-health.prose.md b/litellm-health.prose.md index a28b555..471eb54 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -52,8 +52,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - Router REMOVED from request path — LiteLLM proxies directly to GPU - All GPUs at parallel 2 (was parallel 1) - NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17) -- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal) -- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s +- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here. ## Parameters @@ -92,6 +91,11 @@ CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those valu contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`, `crew-auto`). +> Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and +> per-model timeouts in this contract is intentionally superseded — that state is config, +> and this contract points at the CT 116 config instead. Step 7 likewise authenticates with +> the dedicated `monitor` key, not the master key, which is admin-only. + ## Containers on CT 116 | Container | Image | Port | Health Check | @@ -156,11 +160,11 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to this list. The RTX 5070 host now serves `gpu-vision`. - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so - the set depends on the key. Always state which key a model list was read with. This - probe uses the `monitor` key; on the backend surface - (`http://{{backend_host}}/litellm/v1/models`) that key returns `gpu-vision`, - `qwen3.6-27B-code`, `strix-moe`, `syslog-auto` (verified 2026-09-12). The master key - sees a larger registry — read that from the authority config, not from this probe. + the set depends on the key. Always state which key a model list was read with — a + snapshot without its key is not evidence. This probe uses the `monitor` key on the + backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model + registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather + than freezing a list here. 8. **Check agent keys**: - GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index f3f8347..28ca553 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -93,6 +93,10 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- - Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed — a monitor key, an agent key, and the master key can each return a different set. Agents should use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them. - **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`. +> Re-scope note (2026-09-12): the fallback-chain and per-model timeout tables were removed +> by the single-source-of-truth re-scope; read those values from the CT 116 config named +> above rather than from this contract. + ## Containers on CT 116 | Container | Image | Port | Health Check |