--- kind: function name: litellm-health status: active description: > Verifies the LiteLLM inference stack health. Current architecture (2026-07-09): nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy. Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container, image and config removed). GPU monitoring via Prometheus/Grafana and fleet dashboard (gpu-monitor :9100). Designed as a reusable contract for any Syslog agent. Source of truth: gpu-fleet.prose.md ## Monitoring / Alerting (as-built 2026-08-09) - LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]` (failure_callback alone does NOT mount /metrics — verified 2026-08-09). Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/. - Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09. - Prometheus node job covers 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). Note: .4:9100 is a DEAD target (no route, down for weeks, not a live node). --- ## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU) ``` Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) │ Key validation Fallback chains Budget tracking │ Prometheus ← metrics │ Grafana :3001 harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image and config removed). nginx routes /v1 → LiteLLM directly. Router slot booking + circuit breakers replaced by LiteLLM native fallbacks + timeouts. ``` **What changed (v3.2.0 → v4.0.0 — 2026-07-08)**: - Router REMOVED from request path — LiteLLM proxies directly to GPU - All GPUs at parallel 2 (was parallel 1) - NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12) - Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here. ## Parameters - public_url: string — The public LiteLLM URL (default: "https://litellm.sysloggh.net") - backend_host: string — Internal CT host for container checks (default: "192.168.68.116") - auth_host: string — Authentik server for OIDC (default: "192.168.68.11") - gpu_hosts: array — GPU inference hosts (default: ["192.168.68.8", "192.168.68.110", "192.168.68.15"]) - gpu_dashboard_url: string — Fleet dashboard (default: "http://192.168.68.24:9100") - grafana_url: string — Grafana dashboards (default: "http://192.168.68.116:3001") ## Returns - overall_status: "healthy" | "degraded" | "down" - checks: array of { name: string, status: string, detail: string } - timestamp: string — ISO timestamp of the check run - duration_ms: number — How long the check took ## Requires - SSH key access to backend_host for container checks - Network access to public_url, auth_host, and gpu_dashboard_url - LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`) - The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`, `strix-moe`, `syslog-auto`). Retrieve from the executor's host via: ``` monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2" master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY" ``` If credentials are missing or unreadable, the probe must report `credential-missing` (not bare 401 or "0 keys"). The master key must never be used for inference. ## GPU Fleet Topology | Host | IP | Hardware | Role | |------|-----|----------|------| | llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) | | ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) | | amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) | Single source of truth for models, aliases, rpm caps, weights and fallback chains: CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`, `crew-auto`). > Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and > per-model timeouts in this contract is intentionally superseded — that state is config, > and this contract points at the CT 116 config instead. Step 7 likewise authenticates with > the dedicated `monitor` key, not the master key, which is admin-only. ## Containers on CT 116 | Container | Image | Port | Health Check | |-----------|-------|------|-------------| | harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness | | harness-nginx | nginx:alpine | :80 | HTTP 200 on /health | | harness-postgres | postgres:16-alpine | :5432 | pg_isready | | harness-redis | redis:7-alpine | :6379 | PING | | harness-dashboard | inference-harness-dashboard | :3000 | /health | | harness-grafana | grafana/grafana | :3000→:3001 | /api/health | | harness-prometheus | prom/prometheus | :9090 | /-/healthy | (## Execution ) **Execution model**: This contract is executed by a host-scheduled cron job (see `scripts/contract-run.sh`). The cron job runs the monitoring script directly on the target host and appends the result to `/var/log/contract-runs/.log`. An agent-session acknowledgement (a `done:` line in the ops status log) is NOT execution — it only proves the agent read the result and reported it. The actual monitoring work happens in the host cron job. 1. **Read parameters** — Use provided values or defaults 2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME app under DIFFERENT paths. They are not interchangeable, so every probe below must name the surface it targets. Never point a check at a path that only resolves on the other surface. **Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the ROOT; the `/litellm/` prefix does not exist there and 404s: - GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard") - GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI") - GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge) **Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`: - GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") - GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") - GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11) 3. **Check LiteLLM health (no-auth)**: - GET http://{{backend_host}}/litellm/health/liveliness → expect 200 4. **Check backend container health**: - SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11) - Critical: harness-litellm, harness-nginx, harness-postgres - Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter, trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11) 5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11): - GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON - nginx `/health/unified` is now a `301` redirect to `/gpu/gpu-data` (same payload) 6. **Check GPU fleet health** (via fleet dashboard): - GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON - Verify GPUs reporting status "healthy" - Check alerts array for active warnings/critical 7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health check runs on the **backend edge**, not the public edge, so these paths carry the `/litellm/` prefix: - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8) - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) - POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) - Auth uses the dedicated `monitor` agent key. Retrieve via: `ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"` Do NOT use the master key for inference — the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via: `ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"` - KEY SCOPE: the `monitor` key MUST be scoped for the three probed aliases (`gpu-dense`, `gpu-vision`, `strix-moe`) plus the `syslog-auto` fallback pool, otherwise the probe returns 403 and the host is not covered. If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing alias) and re-run — never drop the host from the probe to make the check pass. - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to this list. The RTX 5070 host now serves `gpu-vision`. - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so the set depends on the key. Always state which key a model list was read with — a snapshot without its key is not evidence. This probe uses the `monitor` key on the backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather than freezing a list here. 8. **Check agent keys**: - GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys - **IMPORTANT**: Run the curl on the CT 116 HOST, not inside the container. The `harness-litellm` container has no curl/wget. Use: `ssh root@192.168.68.116 "curl -s -H 'Authorization: Bearer $MASTER_KEY' http://127.0.0.1:4000/key/list"` - If the response is empty or unparseable, report `admin-call-failed` (not "0 agent keys") 9. **Check Grafana**: - GET {{grafana_url}}/api/health → expect 200 10. **Compile and report** — Determine overall_status from individual check results ## Executor Script (2026-09-13) **Run `scripts/litellm-health-check.py` from the clone.** This script implements all 11 checks defined above and reports results in a standardized format. Paste its output in the status line. - Hand-rolled probes are **not** an acceptable substitute for the script. - Backend-edge checks (steps 2–8) must use `http://192.168.68.116` (internal IP), **not** the public URL `https://litellm.sysloggh.net` (which returns 401 for those paths). - Docker Stats (step 10) must be fetched from the CT 116 host itself (`127.0.0.1:9324/metrics`) because the `harness-docker-stats` container binds to localhost on CT 116. - Admin Key List (step 8) requires the master key expanded locally before SSH, then embedded in the remote curl command with proper quoting. Expected output on a healthy fleet: 11/11 passing checks.