9.0 KiB
kind, name, status, deprecated_on, replaced_by, note, description
| kind | name | status | deprecated_on | replaced_by | note | description |
|---|---|---|---|---|---|---|
| function | litellm-health | deprecated | 2026-07-09 | litellm-self-heal.prose.md | Consolidated into litellm-self-heal.prose.md to eliminate duplication of architecture diagrams, GPU topology, timeout tables, and container lists. Health check is now § Health Check within litellm-self-heal. This file is retained for reference only — use litellm-self-heal instead. | Verifies the LiteLLM inference stack health. Current architecture (2026-07-09): nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy. Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container, image and config removed). GPU monitoring via Prometheus/Grafana and fleet dashboard (gpu-monitor :9100). Designed as a reusable contract for any Syslog agent. Source of truth: gpu-fleet.prose.md ## Monitoring / Alerting (as-built 2026-08-09) - LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]` (failure_callback alone does NOT mount /metrics — verified 2026-08-09). Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/. - Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09. - Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100). |
Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
│
Key validation
Fallback chains
Budget tracking
│
Prometheus ← metrics
│
Grafana :3001
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
and config removed). nginx routes /v1 → LiteLLM directly.
Router slot booking + circuit breakers replaced by
LiteLLM native fallbacks + timeouts.
What changed (v3.2.0 → v4.0.0 — 2026-07-08):
- Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
- Timeouts and fallback chains are config state — read them from CT 116
/opt/inference-harness/litellm_config.yaml; they are not duplicated here.
Parameters
- public_url: string — The public LiteLLM URL (default: "https://litellm.sysloggh.net")
- backend_host: string — Internal CT host for container checks (default: "192.168.68.116")
- auth_host: string — Authentik server for OIDC (default: "192.168.68.11")
- gpu_hosts: array — GPU inference hosts (default: ["192.168.68.8", "192.168.68.110", "192.168.68.15"])
- gpu_dashboard_url: string — Fleet dashboard (default: "http://192.168.68.24:9100")
- grafana_url: string — Grafana dashboards (default: "http://192.168.68.116:3001")
Returns
- overall_status: "healthy" | "degraded" | "down"
- checks: array of { name: string, status: string, detail: string }
- timestamp: string — ISO timestamp of the check run
- duration_ms: number — How long the check took
Requires
- SSH key access to backend_host for container checks
- Network access to public_url, auth_host, and gpu_dashboard_url
- LiteLLM master key for admin endpoints only (
/key/list,/key/generate,/key/info) - The dedicated
monitoragent key on CT 116 at/etc/litellm-monitor.env(root-only 0600) for model inference checks — the master key must never be used for inference
GPU Fleet Topology
| Host | IP | Hardware | Role |
|---|---|---|---|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (gpu-dense) |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (gpu-vision) |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (strix-moe) |
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
CT 116 /opt/inference-harness/litellm_config.yaml. Do not duplicate those values in
contracts — read them there. Do not re-add retired names (gemma-4-12b, gpu-light,
crew-auto).
Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and per-model timeouts in this contract is intentionally superseded — that state is config, and this contract points at the CT 116 config instead. Step 7 likewise authenticates with the dedicated
monitorkey, not the master key, which is admin-only.
Containers on CT 116
| Container | Image | Port | Health Check |
|---|---|---|---|
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
| harness-redis | redis:7-alpine | :6379 | PING |
| harness-dashboard | inference-harness-dashboard | :3000 | /health |
| harness-grafana | grafana/grafana | :3000→:3001 | /api/health |
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
Execution
-
Read parameters — Use provided values or defaults
-
Check the end-user surfaces — the public edge and the backend edge serve the SAME app under DIFFERENT paths. They are not interchangeable, so every probe below must name the surface it targets. Never point a check at a path that only resolves on the other surface.
Public edge —
{{public_url}}(https://litellm.sysloggh.net) serves the app at the ROOT; the/litellm/prefix does not exist there and 404s:- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)
Backend edge —
http://{{backend_host}}(port 80) serves the app UNDER/litellm/:- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11)
-
Check LiteLLM health (no-auth):
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
-
Check backend container health:
- SSH to {{backend_host}} →
docker ps→ verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11) - Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter, trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
- SSH to {{backend_host}} →
-
Check GPU fleet health via gpu-monitor (router decommissioned 2026-09-11):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
- nginx
/health/unifiedis now a301redirect to/gpu/gpu-data(same payload)
-
Check GPU fleet health (via fleet dashboard):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
- Verify GPUs reporting status "healthy"
- Check alerts array for active warnings/critical
-
Check model inference via LiteLLM — Test one model on each GPU host. The health check runs on the backend edge, not the public edge, so these paths carry the
/litellm/prefix:- POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
- Auth uses the dedicated
monitoragent key, read on CT 116 from/etc/litellm-monitor.env(root-only 0600). Do NOT use the master key for inference — the master key is for admin endpoints only (/key/list,/key/generate,/key/info). gemma-4-12bwas retired and returns 400Invalid model name— do not re-add it to this list. The RTX 5070 host now servesgpu-vision./v1/modelsis key-scoped: a model is only visible to keys allowed to use it, so the set depends on the key. Always state which key a model list was read with — a snapshot without its key is not evidence. This probe uses themonitorkey on the backend surface (http://{{backend_host}}/litellm/v1/models). The authoritative model registry is CT 116/opt/inference-harness/litellm_config.yaml; read it there rather than freezing a list here.
-
Check agent keys:
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
-
Check Grafana:
- GET {{grafana_url}}/api/health → expect 200
-
Compile and report — Determine overall_status from individual check results