diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index e539ee4..6f2b80b 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -83,10 +83,11 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen |-------|-----|---------------|---------------| | `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | | `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 | -| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 | +| `gpu-vision` | RTX 5070 (.110) | gpu-vision | Whatever runs on RTX 5070 | -**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.5-9b-it) still work -but are deprecated for agent configs. Only the stable aliases survive model swaps. +**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) still work +but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b` +is retired and returns 400 `Invalid model name`. ## Current Model Assignments (2026-07-15) @@ -114,13 +115,12 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. |-------|---------|-----------|---------| | `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) | | `gpu-dense` | 500 | RTX 3090 | Heavy reasoning | -| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks | +| `gpu-vision` | 500 | RTX 5070 | Vision, web extract, light tasks | -### Fallback Chains -- gemma → qwen -- qwen → gemma -- strix-moe → qwen → gemma -- syslog-auto → qwen → gemma → qwen3.6-35B-udq4 +### Fallback Chains (live `router_settings.fallbacks`, request_timeout 300s throughout) +- syslog-auto → qwen3.6-27B-code → strix-moe → gpu-vision +- qwen3.6-27B-code → strix-moe +- strix-moe → qwen3.6-27B-code → gpu-vision ### Why Strix Halo RPM Is Capped - Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks @@ -266,9 +266,9 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window. All agent configs MUST use stable role-based aliases, never model-specific names: - `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`) -- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`) +- `auxiliary.vision.model: gpu-vision` (NOT `gemma-4-12b`) - `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`) -- `auxiliary.web_extract.model: gpu-light` +- `auxiliary.web_extract.model: gpu-vision` When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched. @@ -285,12 +285,12 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the | Setting | Value | Notes | |---------|-------|-------| -| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) | +| `model.default` | `syslog-auto` | Weighted pool (70% qwen, 20% strix, 10% gpu-vision) | | `model.provider` | `custom:litellm` | LiteLLM on CT116 | | `compression.model` | `strix-moe` | Stable alias — survives model swaps | | `aux.compression.model` | `strix-moe` | Compression auxiliary model | -| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) | -| `aux.web_extract.model` | `gpu-light` | Web extraction | +| `aux.vision.model` | `gpu-vision` | Vision tasks (RTX 5070) | +| `aux.web_extract.model` | `gpu-vision` | Web extraction | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | | `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling | | `compression.threshold` | 0.60 | Triggers at ~77K (~60% of 128K) — optimized for 128K context | diff --git a/litellm-health.prose.md b/litellm-health.prose.md index c4e383f..f0134a2 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -149,21 +149,26 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - Verify GPUs reporting status "healthy" - Check alerts array for active warnings/critical -7. **Check model inference via LiteLLM** — Test one model on each GPU host: - - POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8) - - POST /v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) - - POST /v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) - - Use master key for auth +7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health + check runs on the **backend edge**, not the public edge, so these paths carry the + `/litellm/` prefix: + - POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8) + - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) + - POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) + - Auth uses the dedicated `monitor` agent key, read on CT 116 from + `/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference — + the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to this list. The RTX 5070 host now serves `gpu-vision`. - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so an agent key can return a different set than the master key. Always state which key a - model list was read with. Verified 2026-09-12 with the master key: + model list was read with. Verified 2026-09-12 on the backend surface + (`http://{{backend_host}}/litellm/v1/models`) with the `monitor` key: `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. 8. **Check agent keys**: - - GET /key/list with master key → verify all 6 agents have keys + - GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys 9. **Check Grafana**: - GET {{grafana_url}}/api/health → expect 200 diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index f961460..3768941 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -61,14 +61,14 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) | Host | IP | Hardware | Models Served | Engine | Context | Parallel | |------|-----|----------|---------------|--------|---------|----------| | llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 | -| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 | +| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gpu-vision | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 | | amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 | > Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced. ## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116) -`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20). +`model_name`s served (live registry, verified 2026-09-12 with the master key): `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. `gemma-4-12b`, `gpu-light`, and `crew-auto` are retired and absent from the registry. ### Context Cap Split (2026-08-20) @@ -78,19 +78,19 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) Preferred implementation: uncap shared pool, add capped alias for crew-only. -- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200). -- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively. -- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','gemma-4-12b','strix-moe','gpu-dense','gpu-light']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard 6 only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them. +- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.70, rpm 500, api_base .8) + strix-moe (0.20, rpm 60, api_base .15) + gpu-vision (0.10, rpm 200, api_base .110). +- `gpu-dense` is the high-rpm alias (rpm 500) onto qwen3.6-27B-code; `gpu-light` is retired — the RTX 5070 alias is now `gpu-vision`. +- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','strix-moe','gpu-dense']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard set only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them. `/v1/models` is key-scoped, so a model list is only meaningful with the key it was read with. - **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`. ## Model Fallback Chains (LiteLLM) -| Primary | Timeout | Fallback | Timeout | -|---------|---------|----------|---------| -| qwen3.6-27B-code | 300s | gemma-4-12b | 120s | -| gemma-4-12b | 120s | qwen3.6-27B-code | 300s | -| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — | -| syslog-auto (balanced) | 300s | qwen → gemma | — | +| Primary | Timeout | Fallback chain (live `router_settings.fallbacks`) | Timeout | +|---------|---------|--------------------------------------------------|---------| +| syslog-auto (balanced) | 300s | qwen3.6-27B-code → strix-moe → gpu-vision | 300s | +| qwen3.6-27B-code | 300s | strix-moe | 300s | +| strix-moe | 300s | qwen3.6-27B-code → gpu-vision | 300s | +| gpu-vision | 300s | — (leaf) | — | > Global: request_timeout=300s, nginx proxy_read_timeout=600s @@ -137,10 +137,19 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only. Run this first on every cycle. Results feed into remediation rules below. -### 1. Check public endpoints -- GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") # served directly by nginx -- GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") # served directly by nginx -- GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11) +### 1. Check the end-user surfaces +The public edge and the backend edge serve the same app under different paths; they are not +interchangeable, so every probe names the surface it targets. + +Public edge — `{{public_url}}` serves the app at the ROOT (the `/litellm/` prefix 404s): +- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard") +- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI") +- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge) + +Backend edge — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`: +- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") +- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") +- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11) ### 2. Check LiteLLM health (no-auth) - GET http://{{backend_host}}/litellm/health/liveliness → expect 200 @@ -158,14 +167,14 @@ Run this first on every cycle. Results feed into remediation rules below. - Verify GPUs reporting status "healthy" - Check alerts array for active warnings/critical -### 5. Check model inference via LiteLLM — test each model -- POST /v1/chat/completions model=gemma-4-12b → expect 200 -- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 -- POST /v1/chat/completions model=strix-moe → expect 200 -- Use master key for auth +### 5. Check model inference via LiteLLM — test one model per GPU host (backend surface) +- POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8) +- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) +- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) +- Auth uses the dedicated `monitor` agent key, read on CT 116 from `/etc/litellm-monitor.env` (root-only 0600) — never the master key ### 6. Check agent keys -- GET /key/list with master key → verify all 6 agents have keys +- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys ### 7. Check Grafana - GET {{grafana_url}}/api/health → expect 200