From 1a598d0fcb747510c7cbd9975bb18bf48519688b Mon Sep 17 00:00:00 2001 From: abiba Date: Sat, 12 Sep 2026 15:11:24 +0000 Subject: [PATCH 1/8] docs(litellm-health): fix public-vs-backend probe surfaces and stale gemma model list - Execution step 2 now documents the public edge and the backend edge as two distinct surfaces: public serves /ui/ and /docs (404 on the /litellm/ prefix), backend http://192.168.68.116 serves /litellm/ui/ and /litellm/docs (with /ui/ and /docs as 301 helpers). Each probe names its surface. - GPU topology: ocu-llm RTX 5070 now serves gpu-vision (gemma-4-12b retired). - Fallback/timeout table rewritten to the live router_settings.fallbacks chains. - Step 7 model list: gemma-4-12b -> gpu-vision, with a key-scoped /v1/models note and the 2026-09-12 master-key registry snapshot. --- litellm-health.prose.md | 52 ++++++++++++++++++++++++++++------------- 1 file changed, 36 insertions(+), 16 deletions(-) diff --git a/litellm-health.prose.md b/litellm-health.prose.md index ab502c8..c4e383f 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -82,19 +82,20 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) | Host | IP | Hardware | Models Served | Engine | Context | Parallel | |------|-----|----------|---------------|--------|---------|----------| | llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 | -| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 | +| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gpu-vision | llama-server systemd | **128K** | 2 | | amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 | ## Model Fallback Chains (LiteLLM) -| Primary | Timeout | Fallback | Timeout | -|---------|---------|----------|---------| -| qwen3.6-27B-code | 300s | gemma-4-12b | 120s | -| gemma-4-12b | 120s | qwen3.6-27B-code | 300s | -| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — | -| syslog-auto (balanced) | 300s | qwen → gemma | — | +| Primary | Timeout | Fallback chain (live `router_settings.fallbacks`) | Timeout | +|---------|---------|--------------------------------------------------|---------| +| syslog-auto (balanced) | 300s | qwen3.6-27B-code → strix-moe → gpu-vision | 300s | +| qwen3.6-27B-code | 300s | strix-moe | 300s | +| strix-moe | 300s | qwen3.6-27B-code → gpu-vision | 300s | +| gpu-vision | 300s | — (leaf) | — | -> Global: request_timeout=300s, nginx proxy_read_timeout=600s +> Global: request_timeout=300s, nginx proxy_read_timeout=600s. `gemma-4-12b` was retired +> and is NOT in the registry — do not re-add it to this table. ## Containers on CT 116 @@ -112,10 +113,22 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) 1. **Read parameters** — Use provided values or defaults -2. **Check public endpoints**: - - GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") - - GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") - - GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11) +2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME + app under DIFFERENT paths. They are not interchangeable, so every probe below must name + the surface it targets. Never point a check at a path that only resolves on the other + surface. + + **Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the + ROOT; the `/litellm/` prefix does not exist there and 404s: + - GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard") + - GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI") + - GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge) + + **Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`: + - GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") + - GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") + - GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, + added 2026-09-11) 3. **Check LiteLLM health (no-auth)**: - GET http://{{backend_host}}/litellm/health/liveliness → expect 200 @@ -136,11 +149,18 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - Verify GPUs reporting status "healthy" - Check alerts array for active warnings/critical -7. **Check model inference via LiteLLM** — Test each model: - - POST /v1/chat/completions model=gemma-4-12b → expect 200 - - POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 - - POST /v1/chat/completions model=strix-moe → expect 200 +7. **Check model inference via LiteLLM** — Test one model on each GPU host: + - POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8) + - POST /v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) + - POST /v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) - Use master key for auth + - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to + this list. The RTX 5070 host now serves `gpu-vision`. + - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so + an agent key can return a different set than the master key. Always state which key a + model list was read with. Verified 2026-09-12 with the master key: + `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, + `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. 8. **Check agent keys**: - GET /key/list with master key → verify all 6 agents have keys -- 2.54.0 From 3f07b9bbccd961c54e963f248001706d44647a70 Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 15:22:02 +0000 Subject: [PATCH 2/8] no-mistakes(review): fix master-key inference and probe-surface drift in sibling contracts --- gpu-fleet.prose.md | 28 ++++++++++----------- litellm-health.prose.md | 19 ++++++++------ litellm-self-heal.prose.md | 51 ++++++++++++++++++++++---------------- 3 files changed, 56 insertions(+), 42 deletions(-) diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index e539ee4..6f2b80b 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -83,10 +83,11 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen |-------|-----|---------------|---------------| | `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | | `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 | -| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 | +| `gpu-vision` | RTX 5070 (.110) | gpu-vision | Whatever runs on RTX 5070 | -**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.5-9b-it) still work -but are deprecated for agent configs. Only the stable aliases survive model swaps. +**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) still work +but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b` +is retired and returns 400 `Invalid model name`. ## Current Model Assignments (2026-07-15) @@ -114,13 +115,12 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. |-------|---------|-----------|---------| | `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) | | `gpu-dense` | 500 | RTX 3090 | Heavy reasoning | -| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks | +| `gpu-vision` | 500 | RTX 5070 | Vision, web extract, light tasks | -### Fallback Chains -- gemma → qwen -- qwen → gemma -- strix-moe → qwen → gemma -- syslog-auto → qwen → gemma → qwen3.6-35B-udq4 +### Fallback Chains (live `router_settings.fallbacks`, request_timeout 300s throughout) +- syslog-auto → qwen3.6-27B-code → strix-moe → gpu-vision +- qwen3.6-27B-code → strix-moe +- strix-moe → qwen3.6-27B-code → gpu-vision ### Why Strix Halo RPM Is Capped - Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks @@ -266,9 +266,9 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window. All agent configs MUST use stable role-based aliases, never model-specific names: - `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`) -- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`) +- `auxiliary.vision.model: gpu-vision` (NOT `gemma-4-12b`) - `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`) -- `auxiliary.web_extract.model: gpu-light` +- `auxiliary.web_extract.model: gpu-vision` When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched. @@ -285,12 +285,12 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the | Setting | Value | Notes | |---------|-------|-------| -| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) | +| `model.default` | `syslog-auto` | Weighted pool (70% qwen, 20% strix, 10% gpu-vision) | | `model.provider` | `custom:litellm` | LiteLLM on CT116 | | `compression.model` | `strix-moe` | Stable alias — survives model swaps | | `aux.compression.model` | `strix-moe` | Compression auxiliary model | -| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) | -| `aux.web_extract.model` | `gpu-light` | Web extraction | +| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) | +| `aux.web_extract.model` | `gpu-vision` | Web extraction | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | | `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling | | `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context | diff --git a/litellm-health.prose.md b/litellm-health.prose.md index c4e383f..f0134a2 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -149,21 +149,26 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - Verify GPUs reporting status "healthy" - Check alerts array for active warnings/critical -7. **Check model inference via LiteLLM** — Test one model on each GPU host: - - POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8) - - POST /v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) - - POST /v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) - - Use master key for auth +7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health + check runs on the **backend edge**, not the public edge, so these paths carry the + `/litellm/` prefix: + - POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8) + - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) + - POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) + - Auth uses the dedicated `monitor` agent key, read on CT 116 from + `/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference — + the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to this list. The RTX 5070 host now serves `gpu-vision`. - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so an agent key can return a different set than the master key. Always state which key a - model list was read with. Verified 2026-09-12 with the master key: + model list was read with. Verified 2026-09-12 on the backend surface + (`http://{{backend_host}}/litellm/v1/models`) with the `monitor` key: `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. 8. **Check agent keys**: - - GET /key/list with master key → verify all 6 agents have keys + - GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys 9. **Check Grafana**: - GET {{grafana_url}}/api/health → expect 200 diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index f961460..3768941 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -61,14 +61,14 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) | Host | IP | Hardware | Models Served | Engine | Context | Parallel | |------|-----|----------|---------------|--------|---------|----------| | llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 | -| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 | +| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gpu-vision | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 | | amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 | > Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced. ## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116) -`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20). +`model_name`s served (live registry, verified 2026-09-12 with the master key): `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. `gemma-4-12b`, `gpu-light`, and `crew-auto` are retired and absent from the registry. ### Context Cap Split (2026-08-20) @@ -78,19 +78,19 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) Preferred implementation: uncap shared pool, add capped alias for crew-only. -- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200). -- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively. -- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','gemma-4-12b','strix-moe','gpu-dense','gpu-light']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard 6 only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them. +- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.70, rpm 500, api_base .8) + strix-moe (0.20, rpm 60, api_base .15) + gpu-vision (0.10, rpm 200, api_base .110). +- `gpu-dense` is the high-rpm alias (rpm 500) onto qwen3.6-27B-code; `gpu-light` is retired — the RTX 5070 alias is now `gpu-vision`. +- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','strix-moe','gpu-dense']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard set only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them. `/v1/models` is key-scoped, so a model list is only meaningful with the key it was read with. - **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`. ## Model Fallback Chains (LiteLLM) -| Primary | Timeout | Fallback | Timeout | -|---------|---------|----------|---------| -| qwen3.6-27B-code | 300s | gemma-4-12b | 120s | -| gemma-4-12b | 120s | qwen3.6-27B-code | 300s | -| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — | -| syslog-auto (balanced) | 300s | qwen → gemma | — | +| Primary | Timeout | Fallback chain (live `router_settings.fallbacks`) | Timeout | +|---------|---------|--------------------------------------------------|---------| +| syslog-auto (balanced) | 300s | qwen3.6-27B-code → strix-moe → gpu-vision | 300s | +| qwen3.6-27B-code | 300s | strix-moe | 300s | +| strix-moe | 300s | qwen3.6-27B-code → gpu-vision | 300s | +| gpu-vision | 300s | — (leaf) | — | > Global: request_timeout=300s, nginx proxy_read_timeout=600s @@ -137,10 +137,19 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only. Run this first on every cycle. Results feed into remediation rules below. -### 1. Check public endpoints -- GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") # served directly by nginx -- GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") # served directly by nginx -- GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11) +### 1. Check the end-user surfaces +The public edge and the backend edge serve the same app under different paths; they are not +interchangeable, so every probe names the surface it targets. + +Public edge — `{{public_url}}` serves the app at the ROOT (the `/litellm/` prefix 404s): +- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard") +- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI") +- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge) + +Backend edge — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`: +- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") +- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") +- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11) ### 2. Check LiteLLM health (no-auth) - GET http://{{backend_host}}/litellm/health/liveliness → expect 200 @@ -158,14 +167,14 @@ Run this first on every cycle. Results feed into remediation rules below. - Verify GPUs reporting status "healthy" - Check alerts array for active warnings/critical -### 5. Check model inference via LiteLLM — test each model -- POST /v1/chat/completions model=gemma-4-12b → expect 200 -- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 -- POST /v1/chat/completions model=strix-moe → expect 200 -- Use master key for auth +### 5. Check model inference via LiteLLM — test one model per GPU host (backend surface) +- POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8) +- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) +- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) +- Auth uses the dedicated `monitor` agent key, read on CT 116 from `/etc/litellm-monitor.env` (root-only 0600) — never the master key ### 6. Check agent keys -- GET /key/list with master key → verify all 6 agents have keys +- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys ### 7. Check Grafana - GET {{grafana_url}}/api/health → expect 200 -- 2.54.0 From d9eb18c024cf3d3a01af04562aa48fc1b7afd636 Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 15:30:14 +0000 Subject: [PATCH 3/8] no-mistakes(review): align sibling contracts on aliases, crew cap, and monitor key --- gpu-fleet.prose.md | 9 ++++++--- litellm-health.prose.md | 4 +++- litellm-self-heal.prose.md | 8 ++++---- 3 files changed, 13 insertions(+), 8 deletions(-) diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 6f2b80b..e952fea 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -6,9 +6,10 @@ description: > registration, health checks, LiteLLM sync, agent key management, GPU saturation watchdog, Prometheus/Grafana monitoring, and self-healing. UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe, - gpu-dense, gpu-light. These never change — only the underlying model does. + gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12). + These never change — only the underlying model does. Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). - RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster). + RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster). UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability. Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias). Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. @@ -100,7 +101,9 @@ is retired and returns 400 `Invalid model name`. | Model | GPU | Weight | RPM Cap | Timeout | |-------|-----|--------|---------|---------| -| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | +| `qwen3.6-27B-code` | RTX 3090 (.8:8080) | **0.70** | 500 | **300s** | +| `strix-moe` | Strix Halo (.15:8080) | **0.20** | 60 | **300s** | +| `gpu-vision` | RTX 5070 (.110:8080) | **0.10** | 200 | **300s** | Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path. diff --git a/litellm-health.prose.md b/litellm-health.prose.md index f0134a2..aca07e0 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -75,7 +75,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - SSH key access to backend_host for container checks - Network access to public_url, auth_host, and gpu_dashboard_url -- LiteLLM master key for key management endpoints +- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`) +- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for + model inference checks — the master key must never be used for inference ## GPU Fleet Topology diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 3768941..6101259 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -70,17 +70,17 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) `model_name`s served (live registry, verified 2026-09-12 with the master key): `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. `gemma-4-12b`, `gpu-light`, and `crew-auto` are retired and absent from the registry. -### Context Cap Split (2026-08-20) +### Context Cap Split (2026-08-20; crew cap RETIRED) - **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads - **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped -- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit +- **Crewmates** (ops, tune, verify, auth-keys, build): the 64K cap was retired together with the `crew-auto` alias; NO context cap is currently in force. -Preferred implementation: uncap shared pool, add capped alias for crew-only. +> The 64K crew cap was retired with `crew-auto` (2026-09-12). No limit is currently in force; reinstating one would need per-key model limits as a separate, deliberately-scoped change. - `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.70, rpm 500, api_base .8) + strix-moe (0.20, rpm 60, api_base .15) + gpu-vision (0.10, rpm 200, api_base .110). - `gpu-dense` is the high-rpm alias (rpm 500) onto qwen3.6-27B-code; `gpu-light` is retired — the RTX 5070 alias is now `gpu-vision`. -- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','strix-moe','gpu-dense']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard set only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them. `/v1/models` is key-scoped, so a model list is only meaningful with the key it was read with. +- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed. A representative agent key (read 2026-09-12) returns `deepseek-v4-pro`, `gpu-dense`, `gpu-vision`, `qwen3.6-27B-code`, `qwen3.6-35B-udq4`, `strix-moe`, `syslog-auto` — `gpu-vision` present, `gpu-light` and `gemma-4-12b` absent. Agents should still use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them. - **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`. ## Model Fallback Chains (LiteLLM) -- 2.54.0 From 2f65c38213518d662c61b5bce132c1dfa2cd2d0f Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 15:39:15 +0000 Subject: [PATCH 4/8] no-mistakes(review): delete duplicated config tables, point to CT 116 authority --- gpu-fleet.prose.md | 65 +++++++++++--------------------------- litellm-health.prose.md | 35 ++++++++------------ litellm-self-heal.prose.md | 43 ++++++++++++------------- 3 files changed, 53 insertions(+), 90 deletions(-) diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index e952fea..cebe3b4 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -80,56 +80,29 @@ triggers: Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names. When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched. -| Alias | GPU | Current Model | Will Route To | -|-------|-----|---------------|---------------| -| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | -| `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 | -| `gpu-vision` | RTX 5070 (.110) | gpu-vision | Whatever runs on RTX 5070 | +| Alias | Serves | Where | Kind | +|-------|--------|-------|------| +| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias | +| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member | +| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias | +| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router | + +Single source of truth for models, aliases, rpm caps, weights and fallback chains: +CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in +contracts — read them there. **Backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) still work but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b` is retired and returns 400 `Invalid model name`. -## Current Model Assignments (2026-07-15) - -| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | -|-------|-----|------|------|-----|----------|----------|-------------|--------| - ## Routing Configuration (LiteLLM — July 2026) -### syslog-auto Weighted Pool (Direct GPU — bypasses router) - -| Model | GPU | Weight | RPM Cap | Timeout | -|-------|-----|--------|---------|---------| -| `qwen3.6-27B-code` | RTX 3090 (.8:8080) | **0.70** | 500 | **300s** | -| `strix-moe` | Strix Halo (.15:8080) | **0.20** | 60 | **300s** | -| `gpu-vision` | RTX 5070 (.110:8080) | **0.10** | 200 | **300s** | +Single source of truth for models, aliases, rpm caps, weights and fallback chains: +CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in +contracts — read them there. Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path. -### Direct Model Endpoints - -| Model | RPM Cap | Notes | -|-------|---------|-------| - -### Stable Aliases (for agent configs — never change) - -| Alias | RPM Cap | Routes To | Purpose | -|-------|---------|-----------|---------| -| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) | -| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning | -| `gpu-vision` | 500 | RTX 5070 | Vision, web extract, light tasks | - -### Fallback Chains (live `router_settings.fallbacks`, request_timeout 300s throughout) -- syslog-auto → qwen3.6-27B-code → strix-moe → gpu-vision -- qwen3.6-27B-code → strix-moe -- strix-moe → qwen3.6-27B-code → gpu-vision - -### Why Strix Halo RPM Is Capped -- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks -- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously -- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C - ## Operations ### add-model @@ -179,9 +152,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act 2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts 3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!") 4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models` -5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` - - Qwen3.5-9B: 120s, qwen3.6-27B-code: 300s, Carnice-Qwen3.6-MoE-35B-A3B/strix-moe: 300s (strix-moe alias retained, legacy name qwen3.6-35B-udq4 deprecated) - - global request_timeout: 300s, nginx proxy_read_timeout: 600s +5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` — read the live values from the authority config; do not assert them from this contract. 6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power) 7. Check port conflicts: verify only one llama-server on :8080 per host 8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`) @@ -268,9 +239,9 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window. ### Stable Aliases — CRITICAL All agent configs MUST use stable role-based aliases, never model-specific names: -- `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`) -- `auxiliary.vision.model: gpu-vision` (NOT `gemma-4-12b`) -- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`) +- `compression.model: strix-moe` +- `auxiliary.vision.model: gpu-vision` +- `delegation.model: gpu-dense` - `auxiliary.web_extract.model: gpu-vision` When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched. @@ -288,7 +259,7 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the | Setting | Value | Notes | |---------|-------|-------| -| `model.default` | `syslog-auto` | Weighted pool (70% qwen, 20% strix, 10% gpu-vision) | +| `model.default` | `syslog-auto` | Balanced default (pool router) | | `model.provider` | `custom:litellm` | LiteLLM on CT116 | | `compression.model` | `strix-moe` | Stable alias — survives model swaps | | `aux.compression.model` | `strix-moe` | Compression auxiliary model | diff --git a/litellm-health.prose.md b/litellm-health.prose.md index aca07e0..a28b555 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -81,23 +81,16 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) ## GPU Fleet Topology -| Host | IP | Hardware | Models Served | Engine | Context | Parallel | -|------|-----|----------|---------------|--------|---------|----------| -| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 | -| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gpu-vision | llama-server systemd | **128K** | 2 | -| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 | +| Host | IP | Hardware | Role | +|------|-----|----------|------| +| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) | +| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) | +| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) | -## Model Fallback Chains (LiteLLM) - -| Primary | Timeout | Fallback chain (live `router_settings.fallbacks`) | Timeout | -|---------|---------|--------------------------------------------------|---------| -| syslog-auto (balanced) | 300s | qwen3.6-27B-code → strix-moe → gpu-vision | 300s | -| qwen3.6-27B-code | 300s | strix-moe | 300s | -| strix-moe | 300s | qwen3.6-27B-code → gpu-vision | 300s | -| gpu-vision | 300s | — (leaf) | — | - -> Global: request_timeout=300s, nginx proxy_read_timeout=600s. `gemma-4-12b` was retired -> and is NOT in the registry — do not re-add it to this table. +Single source of truth for models, aliases, rpm caps, weights and fallback chains: +CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in +contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`, +`crew-auto`). ## Containers on CT 116 @@ -163,11 +156,11 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to this list. The RTX 5070 host now serves `gpu-vision`. - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so - an agent key can return a different set than the master key. Always state which key a - model list was read with. Verified 2026-09-12 on the backend surface - (`http://{{backend_host}}/litellm/v1/models`) with the `monitor` key: - `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, - `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. + the set depends on the key. Always state which key a model list was read with. This + probe uses the `monitor` key; on the backend surface + (`http://{{backend_host}}/litellm/v1/models`) that key returns `gpu-vision`, + `qwen3.6-27B-code`, `strix-moe`, `syslog-auto` (verified 2026-09-12). The master key + sees a larger registry — read that from the authority config, not from this probe. 8. **Check agent keys**: - GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 6101259..f3f8347 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -58,17 +58,29 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) ## GPU Fleet Topology -| Host | IP | Hardware | Models Served | Engine | Context | Parallel | -|------|-----|----------|---------------|--------|---------|----------| -| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 | -| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gpu-vision | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 | -| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 | +| Host | IP | Hardware | Role | +|------|-----|----------|------| +| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) | +| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) | +| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) | -> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced. +> Verified on the ground 2026-07-16. The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced. -## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116) +## LiteLLM Model Surface -`model_name`s served (live registry, verified 2026-09-12 with the master key): `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. `gemma-4-12b`, `gpu-light`, and `crew-auto` are retired and absent from the registry. +Single source of truth for models, aliases, rpm caps, weights and fallback chains: +CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in +contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`, +`crew-auto`). + +### Alias Surface + +| Alias | Serves | Where | Kind | +|-------|--------|-------|------| +| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias | +| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member | +| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias | +| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router | ### Context Cap Split (2026-08-20; crew cap RETIRED) @@ -78,22 +90,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) > The 64K crew cap was retired with `crew-auto` (2026-09-12). No limit is currently in force; reinstating one would need per-key model limits as a separate, deliberately-scoped change. -- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.70, rpm 500, api_base .8) + strix-moe (0.20, rpm 60, api_base .15) + gpu-vision (0.10, rpm 200, api_base .110). -- `gpu-dense` is the high-rpm alias (rpm 500) onto qwen3.6-27B-code; `gpu-light` is retired — the RTX 5070 alias is now `gpu-vision`. -- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed. A representative agent key (read 2026-09-12) returns `deepseek-v4-pro`, `gpu-dense`, `gpu-vision`, `qwen3.6-27B-code`, `qwen3.6-35B-udq4`, `strix-moe`, `syslog-auto` — `gpu-vision` present, `gpu-light` and `gemma-4-12b` absent. Agents should still use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them. +- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed — a monitor key, an agent key, and the master key can each return a different set. Agents should use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them. - **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`. -## Model Fallback Chains (LiteLLM) - -| Primary | Timeout | Fallback chain (live `router_settings.fallbacks`) | Timeout | -|---------|---------|--------------------------------------------------|---------| -| syslog-auto (balanced) | 300s | qwen3.6-27B-code → strix-moe → gpu-vision | 300s | -| qwen3.6-27B-code | 300s | strix-moe | 300s | -| strix-moe | 300s | qwen3.6-27B-code → gpu-vision | 300s | -| gpu-vision | 300s | — (leaf) | — | - -> Global: request_timeout=300s, nginx proxy_read_timeout=600s - ## Containers on CT 116 | Container | Image | Port | Health Check | -- 2.54.0 From baaac9d7c60be414706c480ea47affa21b87abfb Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 15:43:21 +0000 Subject: [PATCH 5/8] no-mistakes(review): delete remaining duplicated timeout and frozen model-list state --- gpu-fleet.prose.md | 5 ++--- litellm-health.prose.md | 18 +++++++++++------- litellm-self-heal.prose.md | 4 ++++ 3 files changed, 17 insertions(+), 10 deletions(-) diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index cebe3b4..04cd136 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -217,6 +217,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions. - **Alert migration**: All alerts now go to `#agent-hub` topics (`alerts-gpu`, `alerts-pm2`, `alerts-infra`) instead of DMs. Cross-agent visibility enabled. - **tok/s benchmarks**: Measured every 5 min via LiteLLM proxy. Baselines tracked with 30%/50% degradation thresholds. - **NetBird 502**: Tanko routes through NetBird for litellm.sysloggh.net. Use direct IP if NetBird down. +- **Alias-retirement follow-up (2026-09-12)**: the agent-facing templates (`hermes-config-template.prose.md`, `hermes-agent-baseline.prose.md`), `litellm-api-keys.prose.md`, and the executable `audit-hermes-config.py` still reference the retired `gpu-light`/`gemma-4-12b` names, and `audit-hermes-config.py` Rule 8 currently fails a config whose vision/web_extract model is `gpu-vision`. These are tracked separately and are intentionally NOT updated in this change. ## GPU Inference Benchmarks (Current) @@ -251,7 +252,7 @@ When the underlying model is swapped, only the LiteLLM config changes — agent - **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek) - Compression threshold 0.60: fires at ~77K (~51K headroom before 128K ceiling) - **Pi agents (Abiba)**: `compaction.reserveTokens: 52739` (≈60% of 128K) -- Mumuni compression model alias: `strix-moe` with 300s timeout +- Mumuni compression model alias: `strix-moe` ### Mumuni Agent Profile @@ -273,8 +274,6 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the | `memory.memory_char_limit` | 800 | Brief memory entries | | `personalities` | `creative` | Creative assistant personality | | Platforms | cli, homeassistant, signal, telegram, zulip | All Hermes platforms | -| Main model timeout | 300s | LiteLLM global timeout | -| Compression model timeout | 300s | strix-moe timeout increased from 120s | ### Agent Update Status (2026-07-15) diff --git a/litellm-health.prose.md b/litellm-health.prose.md index a28b555..471eb54 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -52,8 +52,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - Router REMOVED from request path — LiteLLM proxies directly to GPU - All GPUs at parallel 2 (was parallel 1) - NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17) -- LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal) -- nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s +- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here. ## Parameters @@ -92,6 +91,11 @@ CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those valu contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`, `crew-auto`). +> Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and +> per-model timeouts in this contract is intentionally superseded — that state is config, +> and this contract points at the CT 116 config instead. Step 7 likewise authenticates with +> the dedicated `monitor` key, not the master key, which is admin-only. + ## Containers on CT 116 | Container | Image | Port | Health Check | @@ -156,11 +160,11 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to this list. The RTX 5070 host now serves `gpu-vision`. - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so - the set depends on the key. Always state which key a model list was read with. This - probe uses the `monitor` key; on the backend surface - (`http://{{backend_host}}/litellm/v1/models`) that key returns `gpu-vision`, - `qwen3.6-27B-code`, `strix-moe`, `syslog-auto` (verified 2026-09-12). The master key - sees a larger registry — read that from the authority config, not from this probe. + the set depends on the key. Always state which key a model list was read with — a + snapshot without its key is not evidence. This probe uses the `monitor` key on the + backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model + registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather + than freezing a list here. 8. **Check agent keys**: - GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index f3f8347..28ca553 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -93,6 +93,10 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- - Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed — a monitor key, an agent key, and the master key can each return a different set. Agents should use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them. - **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`. +> Re-scope note (2026-09-12): the fallback-chain and per-model timeout tables were removed +> by the single-source-of-truth re-scope; read those values from the CT 116 config named +> above rather than from this contract. + ## Containers on CT 116 | Container | Image | Port | Health Check | -- 2.54.0 From dc42ecc235816bde0b22e12fc9149bf10b37919f Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 15:56:16 +0000 Subject: [PATCH 6/8] no-mistakes(document): Align gpu-vision and single-source-of-truth documentation --- README.md | 2 +- docs/AUTHORING-GUIDE.md | 2 +- gpu-fleet.prose.md | 7 +++---- 3 files changed, 5 insertions(+), 6 deletions(-) diff --git a/README.md b/README.md index 46b8968..ac61914 100644 --- a/README.md +++ b/README.md @@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4 # Configure an agent with a different auxiliary model -prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gemma-4-12b +prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision ``` ### Option B: Manual Execution diff --git a/docs/AUTHORING-GUIDE.md b/docs/AUTHORING-GUIDE.md index f73129e..adf2f78 100644 --- a/docs/AUTHORING-GUIDE.md +++ b/docs/AUTHORING-GUIDE.md @@ -13,7 +13,7 @@ the description: 1. **What system does this contract touch?** Name the hosts, CTs, containers, and services explicitly. "The inference fleet" is vague. "GPU .8 (RTX 3090, - qwen), .110 (RTX 5070, gemma), .15 (Strix Halo, strix-moe), and LiteLLM on CT + qwen), .110 (RTX 5070, gpu-vision), .15 (Strix Halo, strix-moe), and LiteLLM on CT 116" is specific. 2. **Who runs this contract, and when?** State the agent, the trigger (cron, diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 04cd136..fefebba 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -25,7 +25,7 @@ triggers: ## Maintains -- gpu_roster: { models: map, hosts: map } — Single source of truth for all GPU models +- gpu_roster: { models: map, hosts: map } — GPU host/model roster; the authoritative alias/weight/fallback registry is CT 116 `litellm_config.yaml` - router: { status: "healthy", roster_loaded: bool, models: array } - litellm: { status: "healthy", keys: array, models: array } - agent_keys: { agent: api_key } — All agent API keys registered in LiteLLM DB @@ -97,9 +97,8 @@ is retired and returns 400 `Invalid model name`. ## Routing Configuration (LiteLLM — July 2026) -Single source of truth for models, aliases, rpm caps, weights and fallback chains: -CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in -contracts — read them there. +Model, alias, rpm/weight and fallback values are owned by CT 116 +`/opt/inference-harness/litellm_config.yaml` (see § Stable Role-Based Aliases above). Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path. -- 2.54.0 From bb17c2120fa6a23360225a64fc130da7458bbe03 Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 16:00:56 +0000 Subject: [PATCH 7/8] no-mistakes(document): Make litellm-health the live owner; dedupe probes --- README.md | 4 +-- contract-registry.yaml | 4 +-- litellm-health.prose.md | 9 +---- litellm-self-heal.prose.md | 68 ++++++-------------------------------- 4 files changed, 15 insertions(+), 70 deletions(-) diff --git a/README.md b/README.md index ac61914..c3860c2 100644 --- a/README.md +++ b/README.md @@ -116,7 +116,7 @@ Run on trigger or schedule. Maintain persistent world-model state across runs. | `zulip-health` | Zulip | Checks Zulip connectivity, message flow, and bot responsiveness. | | `zulip-mention-reliability` | Zulip | Diagnoses and fixes @mention detection issues in Zulip. | | `zulip-approval-fix` | Zulip | Fixes broken /approve and /deny slash commands for Hermes agents. | -| `litellm-self-heal` | LiteLLM | Consolidated health check + self-healing for the full nginx → LiteLLM → GPU chain. Verifies 8 containers, 3 GPUs, model inference, and agent keys. Applies 9 remediation rules. (litellm-health merged into this contract 2026-07-09.) | +| `litellm-self-heal` | LiteLLM | Applies remediation rules for LiteLLM stack failures detected by `litellm-health` (full nginx → LiteLLM → GPU chain). 9 remediation rules. | | `gpu-fleet` | GPU | Manages the GPU inference fleet: model deployment, registration, health checks, LiteLLM sync. | | `gpu-monitor` | GPU | Comprehensive GPU fleet monitor — polls sidecars, router, LiteLLM every 15s, renders SSE dashboard. | | `proxmox-monitor` | Infra | Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack on CT 116. | @@ -149,7 +149,7 @@ Called on-demand as single-render tools. | Contract | Description | |---|---| | `litellm-api-keys` | Manages LiteLLM API keys for agent identity. Create, rotate, verify, and list agent keys. References gpu-fleet for current key inventory. | -| `litellm-health` | ⚠️ **DEPRECATED** — consolidated into `litellm-self-heal` (2026-07-09). Retained for reference only. | +| `litellm-health` | LiteLLM health check: public vs backend surfaces, CT 116 containers, GPU fleet, one model per GPU host, and agent keys. Owner of the probes; `litellm-self-heal` owns remediation. | | `infrastructure-monitoring` | Target-state for Prometheus + GPU exporters + Grafana. Core stack deployed, GPU exporters NOT live. | | `stirling-pdf-agent-access` | Documents the Stirling-PDF API access pattern for agents — global API key, 12 operations, curl examples. Agents use the `stirling-pdf-api` shared skill for templates. | | `hello-world` | Minimal test contract — verifies the OpenProse execution pipeline works. | diff --git a/contract-registry.yaml b/contract-registry.yaml index 93f398a..019be2b 100644 --- a/contract-registry.yaml +++ b/contract-registry.yaml @@ -693,8 +693,8 @@ contracts: version: 1.0.0 trigger: type: scheduled - cadence: '*/10 * * * *' - description: "Every 10 minutes \u2014 LiteLLM proxy health" + cadence: '5 3,7,11,15,19,23 * * *' + description: "4-hourly staggered dispatch via fm-send (run contract litellm-health)" cron_job_id: null execution: agent: abiba diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 471eb54..4a430ca 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -1,14 +1,7 @@ --- kind: function name: litellm-health -status: deprecated -deprecated_on: 2026-07-09 -replaced_by: litellm-self-heal.prose.md -note: > - Consolidated into litellm-self-heal.prose.md to eliminate duplication - of architecture diagrams, GPU topology, timeout tables, and container - lists. Health check is now § Health Check within litellm-self-heal. - This file is retained for reference only — use litellm-self-heal instead. +status: active description: > Verifies the LiteLLM inference stack health. Current architecture (2026-07-09): nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy. diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 28ca553..189cdb0 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -11,22 +11,22 @@ note: > Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`). Reports to /var/log/litellm/health-*.json and Gitea (SyslogSolution/health-logs). GPU monitoring integrated from gpu-monitor on .24:9100. - - Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate - duplication of architecture diagrams, GPU topology, timeout tables, and container - lists. Health check is now § Health Check within this contract. - + + Health probes are owned by litellm-health.prose.md (dispatched as + `run contract: litellm-health`). This contract owns remediation only — it does not + re-specify the probes. + Source of truth for GPU topology and keys: gpu-fleet.prose.md Last verified: 2026-07-12 description: > - LiteLLM inference stack health monitoring + self-healing. Verifies the full - nginx → LiteLLM → GPU chain, 12 containers on CT 116, 3 GPU hosts, model - inference, and agent keys. Applies remediation rules for common failures. + LiteLLM inference stack remediation. Applies remediation rules for failures detected + by litellm-health.prose.md (nginx → LiteLLM → GPU chain, CT 116 containers, GPU hosts, + model inference, and agent keys). Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs). --- --- -# LiteLLM Operations — Health Check + Self-Heal +# LiteLLM Operations — Self-Heal (Remediation) ## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU) @@ -138,55 +138,7 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- ## Health Check -Run this first on every cycle. Results feed into remediation rules below. - -### 1. Check the end-user surfaces -The public edge and the backend edge serve the same app under different paths; they are not -interchangeable, so every probe names the surface it targets. - -Public edge — `{{public_url}}` serves the app at the ROOT (the `/litellm/` prefix 404s): -- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard") -- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI") -- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge) - -Backend edge — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`: -- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") -- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") -- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11) - -### 2. Check LiteLLM health (no-auth) -- GET http://{{backend_host}}/litellm/health/liveliness → expect 200 - -### 3. Check backend container health -- SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11) -- Critical: harness-litellm, harness-nginx, harness-postgres -- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, - harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter, - trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11) -- Decommissioned 2026-09-11: harness-router (container, image and config removed) - -### 4. Check GPU fleet health (via fleet dashboard) -- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON -- Verify GPUs reporting status "healthy" -- Check alerts array for active warnings/critical - -### 5. Check model inference via LiteLLM — test one model per GPU host (backend surface) -- POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8) -- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) -- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) -- Auth uses the dedicated `monitor` agent key, read on CT 116 from `/etc/litellm-monitor.env` (root-only 0600) — never the master key - -### 6. Check agent keys -- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys - -### 7. Check Grafana -- GET {{grafana_url}}/api/health → expect 200 - -### 8. Compile overall status -Determine overall_status from individual check results: -- "healthy" — all checks pass -- "degraded" — 1-2 non-critical checks fail -- "down" — critical checks fail +Health probes are owned by `litellm-health.prose.md` (dispatched as `run contract: litellm-health`). This contract owns remediation only — it does not re-specify the probes. --- --- -- 2.54.0 From 21f7b6171ccb3137e3f49c7eef65c870da7f447c Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 16:02:15 +0000 Subject: [PATCH 8/8] no-mistakes(document): Dedupe cadence copies; registry remains authoritative --- cron-prompts-review.md | 8 ++++++-- litellm-client-timeouts.prose.md | 5 +++-- 2 files changed, 9 insertions(+), 4 deletions(-) diff --git a/cron-prompts-review.md b/cron-prompts-review.md index d942546..3d2151e 100644 --- a/cron-prompts-review.md +++ b/cron-prompts-review.md @@ -2,6 +2,10 @@ Generated: 2026-07-13 20:59:18 ET +> **Point-in-time snapshot.** Schedules and cadences are authoritative in +> `contract-registry.yaml`; any schedule quoted below may be stale. Do not use +> this file as the source of truth for a contract's trigger. + --- ## hermes-key-enforcement @@ -442,7 +446,7 @@ IMPORTANT: If the contract file does not exist in prose-contracts/main, report f ## litellm-health -**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** */10 * * * * +**Category:** monitoring | **Domain:** litellm | **Owner:** abiba | **Schedule:** see contract-registry.yaml (authoritative) ``` Contract Enforcement: litellm-health @@ -450,7 +454,7 @@ Contract Enforcement: litellm-health Category: monitoring Domain: litellm Owner: abiba -Schedule: Every 10 minutes — LiteLLM proxy health +Schedule: see contract-registry.yaml (authoritative) This is a monitoring contract. Execute the monitoring checks defined in the contract. Report any deviations from expected state. diff --git a/litellm-client-timeouts.prose.md b/litellm-client-timeouts.prose.md index 106fe37..4e7048f 100644 --- a/litellm-client-timeouts.prose.md +++ b/litellm-client-timeouts.prose.md @@ -79,8 +79,9 @@ proxy queuing. Send a real model name and use a real (probe-designated) key. - Probe timeout: 30s. A probe that takes longer than 30s IS the alert — report "backend slow (>30s)" rather than hanging. -- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the - standard; sub-hourly synthetic traffic distorts latency baselines. +- Probe cadence: at most hourly. The litellm-health cron cadence (authoritative + trigger in contract-registry.yaml) is the standard; sub-hourly synthetic + traffic distorts latency baselines. ### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk -- 2.54.0