Files
prose-contracts/litellm-health.prose.md
T
root f99f7e1e34
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
fix: restore per-host probe coverage + sweep residual retired names
A. litellm-health step 7: restore .8 → gpu-dense probe (was duplicated
   to gpu-vision after 8210fd9). Add note documenting monitor key scope
   gap for gpu-dense.
B. Sweep remaining retired names presented as usable:
   - README.md:91: qwen3.6-27B-code → gpu-dense in runnable example
   - gpu-fleet.prose.md:14: qwen3.6-35B-udq4 → Carnice-Qwen3.6-MoE...
   - infrastructure-control.prose.md:226: qwen3.6-35B-udq4 → strix-moe
   - proxmox-monitor.prose.md:90: qwen3.6-35B-udq4 → strix-moe
C. Audit test: 10/10 passed (retired raw names now hard-fail)

Fix-forward from 8210fd9 (direct master push).
2026-09-12 21:56:24 +00:00

173 lines
8.9 KiB
Markdown

---
kind: function
name: litellm-health
status: active
description: >
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container,
image and config removed). GPU monitoring via Prometheus/Grafana and
fleet dashboard (gpu-monitor :9100).
Designed as a reusable contract for any Syslog agent.
Source of truth: gpu-fleet.prose.md
## Monitoring / Alerting (as-built 2026-08-09)
- LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]`
(failure_callback alone does NOT mount /metrics — verified 2026-08-09).
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
- Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100).
---
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
```
Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
│
Key validation
Fallback chains
Budget tracking
│
Prometheus ← metrics
│
Grafana :3001
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
and config removed). nginx routes /v1 → LiteLLM directly.
Router slot booking + circuit breakers replaced by
LiteLLM native fallbacks + timeouts.
```
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
- Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
## Parameters
- public_url: string — The public LiteLLM URL (default: "https://litellm.sysloggh.net")
- backend_host: string — Internal CT host for container checks (default: "192.168.68.116")
- auth_host: string — Authentik server for OIDC (default: "192.168.68.11")
- gpu_hosts: array — GPU inference hosts (default: ["192.168.68.8", "192.168.68.110", "192.168.68.15"])
- gpu_dashboard_url: string — Fleet dashboard (default: "http://192.168.68.24:9100")
- grafana_url: string — Grafana dashboards (default: "http://192.168.68.116:3001")
## Returns
- overall_status: "healthy" | "degraded" | "down"
- checks: array of { name: string, status: string, detail: string }
- timestamp: string — ISO timestamp of the check run
- duration_ms: number — How long the check took
## Requires
- SSH key access to backend_host for container checks
- Network access to public_url, auth_host, and gpu_dashboard_url
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
model inference checks — the master key must never be used for inference
## GPU Fleet Topology
| Host | IP | Hardware | Role |
|------|-----|----------|------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
`crew-auto`).
> Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and
> per-model timeouts in this contract is intentionally superseded — that state is config,
> and this contract points at the CT 116 config instead. Step 7 likewise authenticates with
> the dedicated `monitor` key, not the master key, which is admin-only.
## Containers on CT 116
| Container | Image | Port | Health Check |
|-----------|-------|------|-------------|
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
| harness-redis | redis:7-alpine | :6379 | PING |
| harness-dashboard | inference-harness-dashboard | :3000 | /health |
| harness-grafana | grafana/grafana | :3000→:3001 | /api/health |
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
## Execution
1. **Read parameters** — Use provided values or defaults
2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME
app under DIFFERENT paths. They are not interchangeable, so every probe below must name
the surface it targets. Never point a check at a path that only resolves on the other
surface.
**Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the
ROOT; the `/litellm/` prefix does not exist there and 404s:
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)
**Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`:
- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers,
added 2026-09-11)
3. **Check LiteLLM health (no-auth)**:
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
4. **Check backend container health**:
- SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
- Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter,
trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
- nginx `/health/unified` is now a `301` redirect to `/gpu/gpu-data` (same payload)
6. **Check GPU fleet health** (via fleet dashboard):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
- Verify GPUs reporting status "healthy"
- Check alerts array for active warnings/critical
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
check runs on the **backend edge**, not the public edge, so these paths carry the
`/litellm/` prefix:
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
- Auth uses the dedicated `monitor` agent key, read on CT 116 from
`/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference —
the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`).
- NOTE: the monitor key is scoped for gpu-vision, strix-moe, syslog-auto but NOT gpu-dense.
The gpu-dense probe above requires a key with gpu-dense access; if unavailable, the probe
should be run with a different key or the contract should note the gap rather than silently
dropping host coverage.
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
this list. The RTX 5070 host now serves `gpu-vision`.
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
the set depends on the key. Always state which key a model list was read with — a
snapshot without its key is not evidence. This probe uses the `monitor` key on the
backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model
registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather
than freezing a list here.
8. **Check agent keys**:
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
9. **Check Grafana**:
- GET {{grafana_url}}/api/health → expect 200
10. **Compile and report** — Determine overall_status from individual check results