diff --git a/litellm-health.prose.md b/litellm-health.prose.md index ab502c8..c4e383f 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -82,19 +82,20 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) | Host | IP | Hardware | Models Served | Engine | Context | Parallel | |------|-----|----------|---------------|--------|---------|----------| | llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 | -| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 | +| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gpu-vision | llama-server systemd | **128K** | 2 | | amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 | ## Model Fallback Chains (LiteLLM) -| Primary | Timeout | Fallback | Timeout | -|---------|---------|----------|---------| -| qwen3.6-27B-code | 300s | gemma-4-12b | 120s | -| gemma-4-12b | 120s | qwen3.6-27B-code | 300s | -| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — | -| syslog-auto (balanced) | 300s | qwen → gemma | — | +| Primary | Timeout | Fallback chain (live `router_settings.fallbacks`) | Timeout | +|---------|---------|--------------------------------------------------|---------| +| syslog-auto (balanced) | 300s | qwen3.6-27B-code → strix-moe → gpu-vision | 300s | +| qwen3.6-27B-code | 300s | strix-moe | 300s | +| strix-moe | 300s | qwen3.6-27B-code → gpu-vision | 300s | +| gpu-vision | 300s | — (leaf) | — | -> Global: request_timeout=300s, nginx proxy_read_timeout=600s +> Global: request_timeout=300s, nginx proxy_read_timeout=600s. `gemma-4-12b` was retired +> and is NOT in the registry — do not re-add it to this table. ## Containers on CT 116 @@ -112,10 +113,22 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) 1. **Read parameters** — Use provided values or defaults -2. **Check public endpoints**: - - GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") - - GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") - - GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11) +2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME + app under DIFFERENT paths. They are not interchangeable, so every probe below must name + the surface it targets. Never point a check at a path that only resolves on the other + surface. + + **Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the + ROOT; the `/litellm/` prefix does not exist there and 404s: + - GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard") + - GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI") + - GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge) + + **Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`: + - GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") + - GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") + - GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, + added 2026-09-11) 3. **Check LiteLLM health (no-auth)**: - GET http://{{backend_host}}/litellm/health/liveliness → expect 200 @@ -136,11 +149,18 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - Verify GPUs reporting status "healthy" - Check alerts array for active warnings/critical -7. **Check model inference via LiteLLM** — Test each model: - - POST /v1/chat/completions model=gemma-4-12b → expect 200 - - POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 - - POST /v1/chat/completions model=strix-moe → expect 200 +7. **Check model inference via LiteLLM** — Test one model on each GPU host: + - POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8) + - POST /v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) + - POST /v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) - Use master key for auth + - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to + this list. The RTX 5070 host now serves `gpu-vision`. + - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so + an agent key can return a different set than the master key. Always state which key a + model list was read with. Verified 2026-09-12 with the master key: + `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, + `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. 8. **Check agent keys**: - GET /key/list with master key → verify all 6 agents have keys