Files
prose-contracts/litellm-health.prose.md
root bc7a55122f
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
fix: pm2 contract corrections
pm2-self-heal.prose.md:
- Add AS-BUILT note: gpu-monitor is systemd-managed, NOT PM2
- gpu-watchdog is decommissioned and folded into gpu-monitor.service
- gitea-runner is KEPT; abiba-zulip is KEPT (online for days)
- spoton-service was deleted; live PM2 set is 4 processes
- Preserve historical context for crash-loop guard

litellm-health.prose.md:
- Correct Prometheus node coverage: 6 nodes (.4/.5/.6/.9/.12/.15:9100)
- Note .4:9100 is DEAD target (no route, down for weeks)
- Clarify this does not read as 6 healthy nodes
2026-09-15 13:05:32 +00:00

11 KiB
Raw Permalink Blame History

kind, name, status, description
kind name status description
function litellm-health active Verifies the LiteLLM inference stack health. Current architecture (2026-07-09): nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy. Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container, image and config removed). GPU monitoring via Prometheus/Grafana and fleet dashboard (gpu-monitor :9100). Designed as a reusable contract for any Syslog agent. Source of truth: gpu-fleet.prose.md ## Monitoring / Alerting (as-built 2026-08-09) - LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]` (failure_callback alone does NOT mount /metrics — verified 2026-08-09). Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/. - Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09. - Prometheus node job covers 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). Note: .4:9100 is a DEAD target (no route, down for weeks, not a live node).

Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)

Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
                        │
                  Key validation
                  Fallback chains
                  Budget tracking
                        │
                  Prometheus ← metrics
                        │
                    Grafana :3001

  harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
  and config removed). nginx routes /v1 → LiteLLM directly.
  Router slot booking + circuit breakers replaced by
  LiteLLM native fallbacks + timeouts.

What changed (v3.2.0 → v4.0.0 — 2026-07-08):

  • Router REMOVED from request path — LiteLLM proxies directly to GPU
  • All GPUs at parallel 2 (was parallel 1)
  • NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12)
  • Timeouts and fallback chains are config state — read them from CT 116 /opt/inference-harness/litellm_config.yaml; they are not duplicated here.

Parameters

  • public_url: string — The public LiteLLM URL (default: "https://litellm.sysloggh.net")
  • backend_host: string — Internal CT host for container checks (default: "192.168.68.116")
  • auth_host: string — Authentik server for OIDC (default: "192.168.68.11")
  • gpu_hosts: array — GPU inference hosts (default: ["192.168.68.8", "192.168.68.110", "192.168.68.15"])
  • gpu_dashboard_url: string — Fleet dashboard (default: "http://192.168.68.24:9100")
  • grafana_url: string — Grafana dashboards (default: "http://192.168.68.116:3001")

Returns

  • overall_status: "healthy" | "degraded" | "down"
  • checks: array of { name: string, status: string, detail: string }
  • timestamp: string — ISO timestamp of the check run
  • duration_ms: number — How long the check took

Requires

  • SSH key access to backend_host for container checks

  • Network access to public_url, auth_host, and gpu_dashboard_url

  • LiteLLM master key for admin endpoints only (/key/list, /key/generate, /key/info)

  • The dedicated monitor agent key on CT 116 at /etc/litellm-monitor.env (root-only 0600) for model inference checks, scoped for every alias step 7 probes (gpu-dense, gpu-vision, strix-moe, syslog-auto). Retrieve from the executor's host via:

    monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"
    master key:  ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"
    

    If credentials are missing or unreadable, the probe must report credential-missing (not bare 401 or "0 keys"). The master key must never be used for inference.

GPU Fleet Topology

Host IP Hardware Role
llm-gpu 192.168.68.8 NVIDIA RTX 3090 (24 GB) heavy reasoning (gpu-dense)
ocu-llm 192.168.68.110 NVIDIA RTX 5070 (12 GB) vision / web extract / light tasks (gpu-vision)
amdpve 192.168.68.15 AMD Strix Halo 64GB UMA compression (strix-moe)

Single source of truth for models, aliases, rpm caps, weights and fallback chains: CT 116 /opt/inference-harness/litellm_config.yaml. Do not duplicate those values in contracts — read them there. Do not re-add retired names (gemma-4-12b, gpu-light, crew-auto).

Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and per-model timeouts in this contract is intentionally superseded — that state is config, and this contract points at the CT 116 config instead. Step 7 likewise authenticates with the dedicated monitor key, not the master key, which is admin-only.

Containers on CT 116

Container Image Port Health Check
harness-litellm ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 :4000→:4000 /health/liveliness
harness-nginx nginx:alpine :80 HTTP 200 on /health
harness-postgres postgres:16-alpine :5432 pg_isready
harness-redis redis:7-alpine :6379 PING
harness-dashboard inference-harness-dashboard :3000 /health
harness-grafana grafana/grafana :3000→:3001 /api/health
harness-prometheus prom/prometheus :9090 /-/healthy

Execution

  1. Read parameters — Use provided values or defaults

  2. Check the end-user surfaces — the public edge and the backend edge serve the SAME app under DIFFERENT paths. They are not interchangeable, so every probe below must name the surface it targets. Never point a check at a path that only resolves on the other surface.

    Public edge — {{public_url}} (https://litellm.sysloggh.net) serves the app at the ROOT; the /litellm/ prefix does not exist there and 404s:

    • GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
    • GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
    • GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)

    Backend edge — http://{{backend_host}} (port 80) serves the app UNDER /litellm/:

  3. Check LiteLLM health (no-auth):

  4. Check backend container health:

    • SSH to {{backend_host}} → docker ps → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
    • Critical: harness-litellm, harness-nginx, harness-postgres
    • Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter, trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
  5. Check GPU fleet health via gpu-monitor (router decommissioned 2026-09-11):

    • GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
    • nginx /health/unified is now a 301 redirect to /gpu/gpu-data (same payload)
  6. Check GPU fleet health (via fleet dashboard):

    • GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
    • Verify GPUs reporting status "healthy"
    • Check alerts array for active warnings/critical
  7. Check model inference via LiteLLM — Test one model on each GPU host. The health check runs on the backend edge, not the public edge, so these paths carry the /litellm/ prefix:

    • POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
    • POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
    • POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
    • Auth uses the dedicated monitor agent key. Retrieve via: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2" Do NOT use the master key for inference — the master key is for admin endpoints only (/key/list, /key/generate, /key/info). Retrieve master key via: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"
    • KEY SCOPE: the monitor key MUST be scoped for the three probed aliases (gpu-dense, gpu-vision, strix-moe) plus the syslog-auto fallback pool, otherwise the probe returns 403 and the host is not covered. If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing alias) and re-run — never drop the host from the probe to make the check pass.
    • gemma-4-12b was retired and returns 400 Invalid model name — do not re-add it to this list. The RTX 5070 host now serves gpu-vision.
    • /v1/models is key-scoped: a model is only visible to keys allowed to use it, so the set depends on the key. Always state which key a model list was read with — a snapshot without its key is not evidence. This probe uses the monitor key on the backend surface (http://{{backend_host}}/litellm/v1/models). The authoritative model registry is CT 116 /opt/inference-harness/litellm_config.yaml; read it there rather than freezing a list here.
  8. Check agent keys:

    • GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
    • IMPORTANT: Run the curl on the CT 116 HOST, not inside the container. The harness-litellm container has no curl/wget. Use: ssh root@192.168.68.116 "curl -s -H 'Authorization: Bearer $MASTER_KEY' http://127.0.0.1:4000/key/list"
    • If the response is empty or unparseable, report admin-call-failed (not "0 agent keys")
  9. Check Grafana:

    • GET {{grafana_url}}/api/health → expect 200
  10. Compile and report — Determine overall_status from individual check results

Executor Script (2026-09-13)

Run scripts/litellm-health-check.py from the clone. This script implements all 11 checks defined above and reports results in a standardized format. Paste its output in the status line.

  • Hand-rolled probes are not an acceptable substitute for the script.
  • Backend-edge checks (steps 2–8) must use http://192.168.68.116 (internal IP), not the public URL https://litellm.sysloggh.net (which returns 401 for those paths).
  • Docker Stats (step 10) must be fetched from the CT 116 host itself (127.0.0.1:9324/metrics) because the harness-docker-stats container binds to localhost on CT 116.
  • Admin Key List (step 8) requires the master key expanded locally before SSH, then embedded in the remote curl command with proper quoting.

Expected output on a healthy fleet: 11/11 passing checks.