no-mistakes(review): delete remaining duplicated timeout and frozen model-list state

This commit is contained in:
root
2026-09-12 15:43:21 +00:00
parent 2f65c38213
commit baaac9d7c6
3 changed files with 17 additions and 10 deletions
+2 -3
View File
@@ -217,6 +217,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Alert migration**: All alerts now go to `#agent-hub` topics (`alerts-gpu`, `alerts-pm2`, `alerts-infra`) instead of DMs. Cross-agent visibility enabled.
- **tok/s benchmarks**: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds.
- **NetBird 502**: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down.
- **Alias-retirement follow-up (2026-09-12)**: the agent-facing templates (`hermes-config-template.prose.md`, `hermes-agent-baseline.prose.md`), `litellm-api-keys.prose.md`, and the executable `audit-hermes-config.py` still reference the retired `gpu-light`/`gemma-4-12b` names, and `audit-hermes-config.py` Rule 8 currently fails a config whose vision/web_extract model is `gpu-vision`. These are tracked separately and are intentionally NOT updated in this change.
## GPU Inference Benchmarks (Current)
@@ -251,7 +252,7 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling)
- **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K)
- Mumuni compression model alias: `strix-moe` with 300s timeout
- Mumuni compression model alias: `strix-moe`
### Mumuni Agent Profile
@@ -273,8 +274,6 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the
| `memory.memory_char_limit` | 800 | Brief memory entries |
| `personalities` | `creative` | Creative assistant personality |
| Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms |
| Main model timeout | 300s | LiteLLM global timeout |
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
### Agent Update Status (2026-07-15)
+11 -7
View File
@@ -52,8 +52,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
## Parameters
@@ -92,6 +91,11 @@ CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those valu
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
`crew-auto`).
> Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and
> per-model timeouts in this contract is intentionally superseded — that state is config,
> and this contract points at the CT 116 config instead. Step 7 likewise authenticates with
> the dedicated `monitor` key, not the master key, which is admin-only.
## Containers on CT 116
| Container | Image | Port | Health Check |
@@ -156,11 +160,11 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
this list. The RTX 5070 host now serves `gpu-vision`.
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
the set depends on the key. Always state which key a model list was read with. This
probe uses the `monitor` key; on the backend surface
(`http://{{backend_host}}/litellm/v1/models`) that key returns `gpu-vision`,
`qwen3.6-27B-code`, `strix-moe`, `syslog-auto` (verified 2026-09-12). The master key
sees a larger registry — read that from the authority config, not from this probe.
the set depends on the key. Always state which key a model list was read with — a
snapshot without its key is not evidence. This probe uses the `monitor` key on the
backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model
registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather
than freezing a list here.
8. **Check agent keys**:
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
+4
View File
@@ -93,6 +93,10 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed — a monitor key, an agent key, and the master key can each return a different set. Agents should use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them.
- **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`.
> Re-scope note (2026-09-12): the fallback-chain and per-model timeout tables were removed
> by the single-source-of-truth re-scope; read those values from the CT 116 config named
> above rather than from this contract.
## Containers on CT 116
| Container | Image | Port | Health Check |