Add 'Executor Script' section: - Run scripts/litellm-health-check.py from the clone - Hand-rolled probes not acceptable substitute - Backend-edge checks use internal IP 192.168.68.116, not public URL - Docker Stats fetched from CT 116 host (127.0.0.1:9324/metrics) - Admin Key List requires proper quoting for SSH commands
11 KiB
kind, name, status, description
| kind | name | status | description |
|---|---|---|---|
| function | litellm-health | active | Verifies the LiteLLM inference stack health. Current architecture (2026-07-09): nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy. Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container, image and config removed). GPU monitoring via Prometheus/Grafana and fleet dashboard (gpu-monitor :9100). Designed as a reusable contract for any Syslog agent. Source of truth: gpu-fleet.prose.md ## Monitoring / Alerting (as-built 2026-08-09) - LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]` (failure_callback alone does NOT mount /metrics — verified 2026-08-09). Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/. - Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09. - Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100). |
Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
│
Key validation
Fallback chains
Budget tracking
│
Prometheus ← metrics
│
Grafana :3001
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
and config removed). nginx routes /v1 → LiteLLM directly.
Router slot booking + circuit breakers replaced by
LiteLLM native fallbacks + timeouts.
What changed (v3.2.0 → v4.0.0 — 2026-07-08):
- Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12)
- Timeouts and fallback chains are config state — read them from CT 116
/opt/inference-harness/litellm_config.yaml; they are not duplicated here.
Parameters
- public_url: string — The public LiteLLM URL (default: "https://litellm.sysloggh.net")
- backend_host: string — Internal CT host for container checks (default: "192.168.68.116")
- auth_host: string — Authentik server for OIDC (default: "192.168.68.11")
- gpu_hosts: array — GPU inference hosts (default: ["192.168.68.8", "192.168.68.110", "192.168.68.15"])
- gpu_dashboard_url: string — Fleet dashboard (default: "http://192.168.68.24:9100")
- grafana_url: string — Grafana dashboards (default: "http://192.168.68.116:3001")
Returns
- overall_status: "healthy" | "degraded" | "down"
- checks: array of { name: string, status: string, detail: string }
- timestamp: string — ISO timestamp of the check run
- duration_ms: number — How long the check took
Requires
-
SSH key access to backend_host for container checks
-
Network access to public_url, auth_host, and gpu_dashboard_url
-
LiteLLM master key for admin endpoints only (
/key/list,/key/generate,/key/info) -
The dedicated
monitoragent key on CT 116 at/etc/litellm-monitor.env(root-only 0600) for model inference checks, scoped for every alias step 7 probes (gpu-dense,gpu-vision,strix-moe,syslog-auto). Retrieve from the executor's host via:monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2" master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"If credentials are missing or unreadable, the probe must report
credential-missing(not bare 401 or "0 keys"). The master key must never be used for inference.
GPU Fleet Topology
| Host | IP | Hardware | Role |
|---|---|---|---|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (gpu-dense) |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (gpu-vision) |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (strix-moe) |
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
CT 116 /opt/inference-harness/litellm_config.yaml. Do not duplicate those values in
contracts — read them there. Do not re-add retired names (gemma-4-12b, gpu-light,
crew-auto).
Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and per-model timeouts in this contract is intentionally superseded — that state is config, and this contract points at the CT 116 config instead. Step 7 likewise authenticates with the dedicated
monitorkey, not the master key, which is admin-only.
Containers on CT 116
| Container | Image | Port | Health Check |
|---|---|---|---|
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
| harness-redis | redis:7-alpine | :6379 | PING |
| harness-dashboard | inference-harness-dashboard | :3000 | /health |
| harness-grafana | grafana/grafana | :3000→:3001 | /api/health |
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
Execution
-
Read parameters — Use provided values or defaults
-
Check the end-user surfaces — the public edge and the backend edge serve the SAME app under DIFFERENT paths. They are not interchangeable, so every probe below must name the surface it targets. Never point a check at a path that only resolves on the other surface.
Public edge —
{{public_url}}(https://litellm.sysloggh.net) serves the app at the ROOT; the/litellm/prefix does not exist there and 404s:- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)
Backend edge —
http://{{backend_host}}(port 80) serves the app UNDER/litellm/:- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11)
-
Check LiteLLM health (no-auth):
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
-
Check backend container health:
- SSH to {{backend_host}} →
docker ps→ verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11) - Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter, trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
- SSH to {{backend_host}} →
-
Check GPU fleet health via gpu-monitor (router decommissioned 2026-09-11):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
- nginx
/health/unifiedis now a301redirect to/gpu/gpu-data(same payload)
-
Check GPU fleet health (via fleet dashboard):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
- Verify GPUs reporting status "healthy"
- Check alerts array for active warnings/critical
-
Check model inference via LiteLLM — Test one model on each GPU host. The health check runs on the backend edge, not the public edge, so these paths carry the
/litellm/prefix:- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
- Auth uses the dedicated
monitoragent key. Retrieve via:ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"Do NOT use the master key for inference — the master key is for admin endpoints only (/key/list,/key/generate,/key/info). Retrieve master key via:ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY" - KEY SCOPE: the
monitorkey MUST be scoped for the three probed aliases (gpu-dense,gpu-vision,strix-moe) plus thesyslog-autofallback pool, otherwise the probe returns 403 and the host is not covered. If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing alias) and re-run — never drop the host from the probe to make the check pass. gemma-4-12bwas retired and returns 400Invalid model name— do not re-add it to this list. The RTX 5070 host now servesgpu-vision./v1/modelsis key-scoped: a model is only visible to keys allowed to use it, so the set depends on the key. Always state which key a model list was read with — a snapshot without its key is not evidence. This probe uses themonitorkey on the backend surface (http://{{backend_host}}/litellm/v1/models). The authoritative model registry is CT 116/opt/inference-harness/litellm_config.yaml; read it there rather than freezing a list here.
-
Check agent keys:
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
- IMPORTANT: Run the curl on the CT 116 HOST, not inside the container. The
harness-litellmcontainer has no curl/wget. Use:ssh root@192.168.68.116 "curl -s -H 'Authorization: Bearer $MASTER_KEY' http://127.0.0.1:4000/key/list" - If the response is empty or unparseable, report
admin-call-failed(not "0 agent keys")
-
Check Grafana:
- GET {{grafana_url}}/api/health → expect 200
-
Compile and report — Determine overall_status from individual check results
Executor Script (2026-09-13)
Run scripts/litellm-health-check.py from the clone. This script implements all 11
checks defined above and reports results in a standardized format. Paste its output in
the status line.
- Hand-rolled probes are not an acceptable substitute for the script.
- Backend-edge checks (steps 2–8) must use
http://192.168.68.116(internal IP), not the public URLhttps://litellm.sysloggh.net(which returns 401 for those paths). - Docker Stats (step 10) must be fetched from the CT 116 host itself (
127.0.0.1:9324/metrics) because theharness-docker-statscontainer binds to localhost on CT 116. - Admin Key List (step 8) requires the master key expanded locally before SSH, then embedded in the remote curl command with proper quoting.
Expected output on a healthy fleet: 11/11 passing checks.