docs(litellm-health): fix public-vs-backend probe surfaces and stale gemma model list

- Execution step 2 now documents the public edge and the backend edge as two
  distinct surfaces: public serves /ui/ and /docs (404 on the /litellm/ prefix),
  backend http://192.168.68.116 serves /litellm/ui/ and /litellm/docs (with /ui/
  and /docs as 301 helpers). Each probe names its surface.
- GPU topology: ocu-llm RTX 5070 now serves gpu-vision (gemma-4-12b retired).
- Fallback/timeout table rewritten to the live router_settings.fallbacks chains.
- Step 7 model list: gemma-4-12b -> gpu-vision, with a key-scoped /v1/models note
  and the 2026-09-12 master-key registry snapshot.
This commit is contained in:
abiba
2026-09-12 15:11:24 +00:00
parent e3752af162
commit 1a598d0fcb
+36 -16
View File
@@ -82,19 +82,20 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
| Host | IP | Hardware | Models Served | Engine | Context | Parallel |
|------|-----|----------|---------------|--------|---------|----------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gpu-vision | llama-server systemd | **128K** | 2 |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 |
## Model Fallback Chains (LiteLLM)
| Primary | Timeout | Fallback | Timeout |
|---------|---------|----------|---------|
| qwen3.6-27B-code | 300s | gemma-4-12b | 120s |
| gemma-4-12b | 120s | qwen3.6-27B-code | 300s |
| qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — |
| syslog-auto (balanced) | 300s | qwen → gemma | — |
| Primary | Timeout | Fallback chain (live `router_settings.fallbacks`) | Timeout |
|---------|---------|--------------------------------------------------|---------|
| syslog-auto (balanced) | 300s | qwen3.6-27B-code → strix-moe → gpu-vision | 300s |
| qwen3.6-27B-code | 300s | strix-moe | 300s |
| strix-moe | 300s | qwen3.6-27B-code → gpu-vision | 300s |
| gpu-vision | 300s | — (leaf) | — |
> Global: request_timeout=300s, nginx proxy_read_timeout=600s
> Global: request_timeout=300s, nginx proxy_read_timeout=600s. `gemma-4-12b` was retired
> and is NOT in the registry — do not re-add it to this table.
## Containers on CT 116
@@ -112,10 +113,22 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
1. **Read parameters** — Use provided values or defaults
2. **Check public endpoints**:
- GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
- GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11)
2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME
app under DIFFERENT paths. They are not interchangeable, so every probe below must name
the surface it targets. Never point a check at a path that only resolves on the other
surface.
**Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the
ROOT; the `/litellm/` prefix does not exist there and 404s:
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)
**Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`:
- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers,
added 2026-09-11)
3. **Check LiteLLM health (no-auth)**:
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
@@ -136,11 +149,18 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- Verify GPUs reporting status "healthy"
- Check alerts array for active warnings/critical
7. **Check model inference via LiteLLM** — Test each model:
- POST /v1/chat/completions model=gemma-4-12b → expect 200
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
- POST /v1/chat/completions model=strix-moe → expect 200
7. **Check model inference via LiteLLM** — Test one model on each GPU host:
- POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8)
- POST /v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
- POST /v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
- Use master key for auth
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
this list. The RTX 5070 host now serves `gpu-vision`.
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
an agent key can return a different set than the master key. Always state which key a
model list was read with. Verified 2026-09-12 with the master key:
`qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`,
`qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`.
8. **Check agent keys**:
- GET /key/list with master key → verify all 6 agents have keys