Files
prose-contracts/litellm-health.prose.md
T
root 38d7e8b064
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Update litellm-health contract to mandate executor script
Add 'Executor Script' section:
- Run scripts/litellm-health-check.py from the clone
- Hand-rolled probes not acceptable substitute
- Backend-edge checks use internal IP 192.168.68.116, not public URL
- Docker Stats fetched from CT 116 host (127.0.0.1:9324/metrics)
- Admin Key List requires proper quoting for SSH commands
2026-09-14 03:15:42 +00:00

204 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
kind: function
name: litellm-health
status: active
description: >
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container,
image and config removed). GPU monitoring via Prometheus/Grafana and
fleet dashboard (gpu-monitor :9100).
Designed as a reusable contract for any Syslog agent.
Source of truth: gpu-fleet.prose.md
## Monitoring / Alerting (as-built 2026-08-09)
- LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]`
(failure_callback alone does NOT mount /metrics — verified 2026-08-09).
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
- Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100).
---
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
```
Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
│
Key validation
Fallback chains
Budget tracking
│
Prometheus ← metrics
│
Grafana :3001
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
and config removed). nginx routes /v1 → LiteLLM directly.
Router slot booking + circuit breakers replaced by
LiteLLM native fallbacks + timeouts.
```
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
- Router REMOVED from request path — LiteLLM proxies directly to GPU
- All GPUs at parallel 2 (was parallel 1)
- NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12)
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
## Parameters
- public_url: string — The public LiteLLM URL (default: "https://litellm.sysloggh.net")
- backend_host: string — Internal CT host for container checks (default: "192.168.68.116")
- auth_host: string — Authentik server for OIDC (default: "192.168.68.11")
- gpu_hosts: array — GPU inference hosts (default: ["192.168.68.8", "192.168.68.110", "192.168.68.15"])
- gpu_dashboard_url: string — Fleet dashboard (default: "http://192.168.68.24:9100")
- grafana_url: string — Grafana dashboards (default: "http://192.168.68.116:3001")
## Returns
- overall_status: "healthy" | "degraded" | "down"
- checks: array of { name: string, status: string, detail: string }
- timestamp: string — ISO timestamp of the check run
- duration_ms: number — How long the check took
## Requires
- SSH key access to backend_host for container checks
- Network access to public_url, auth_host, and gpu_dashboard_url
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`,
`strix-moe`, `syslog-auto`). Retrieve from the executor's host via:
```
monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"
master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"
```
If credentials are missing or unreadable, the probe must report `credential-missing` (not bare 401 or "0 keys").
The master key must never be used for inference.
## GPU Fleet Topology
| Host | IP | Hardware | Role |
|------|-----|----------|------|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
`crew-auto`).
> Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and
> per-model timeouts in this contract is intentionally superseded — that state is config,
> and this contract points at the CT 116 config instead. Step 7 likewise authenticates with
> the dedicated `monitor` key, not the master key, which is admin-only.
## Containers on CT 116
| Container | Image | Port | Health Check |
|-----------|-------|------|-------------|
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
| harness-redis | redis:7-alpine | :6379 | PING |
| harness-dashboard | inference-harness-dashboard | :3000 | /health |
| harness-grafana | grafana/grafana | :3000→:3001 | /api/health |
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
## Execution
1. **Read parameters** — Use provided values or defaults
2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME
app under DIFFERENT paths. They are not interchangeable, so every probe below must name
the surface it targets. Never point a check at a path that only resolves on the other
surface.
**Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the
ROOT; the `/litellm/` prefix does not exist there and 404s:
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)
**Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`:
- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers,
added 2026-09-11)
3. **Check LiteLLM health (no-auth)**:
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
4. **Check backend container health**:
- SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
- Critical: harness-litellm, harness-nginx, harness-postgres
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter,
trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
- nginx `/health/unified` is now a `301` redirect to `/gpu/gpu-data` (same payload)
6. **Check GPU fleet health** (via fleet dashboard):
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
- Verify GPUs reporting status "healthy"
- Check alerts array for active warnings/critical
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
check runs on the **backend edge**, not the public edge, so these paths carry the
`/litellm/` prefix:
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
- Auth uses the dedicated `monitor` agent key. Retrieve via:
`ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"`
Do NOT use the master key for inference — the master key is for admin endpoints only
(`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via:
`ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"`
- KEY SCOPE: the `monitor` key MUST be scoped for the three probed aliases (`gpu-dense`,
`gpu-vision`, `strix-moe`) plus the `syslog-auto` fallback pool, otherwise the probe
returns 403 and the host is not covered.
If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing
alias) and re-run — never drop the host from the probe to make the check pass.
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
this list. The RTX 5070 host now serves `gpu-vision`.
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
the set depends on the key. Always state which key a model list was read with — a
snapshot without its key is not evidence. This probe uses the `monitor` key on the
backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model
registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather
than freezing a list here.
8. **Check agent keys**:
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
- **IMPORTANT**: Run the curl on the CT 116 HOST, not inside the container. The `harness-litellm` container has no curl/wget. Use:
`ssh root@192.168.68.116 "curl -s -H 'Authorization: Bearer $MASTER_KEY' http://127.0.0.1:4000/key/list"`
- If the response is empty or unparseable, report `admin-call-failed` (not "0 agent keys")
9. **Check Grafana**:
- GET {{grafana_url}}/api/health → expect 200
10. **Compile and report** — Determine overall_status from individual check results
## Executor Script (2026-09-13)
**Run `scripts/litellm-health-check.py` from the clone.** This script implements all 11
checks defined above and reports results in a standardized format. Paste its output in
the status line.
- Hand-rolled probes are **not** an acceptable substitute for the script.
- Backend-edge checks (steps 2–8) must use `http://192.168.68.116` (internal IP),
**not** the public URL `https://litellm.sysloggh.net` (which returns 401 for those paths).
- Docker Stats (step 10) must be fetched from the CT 116 host itself (`127.0.0.1:9324/metrics`)
because the `harness-docker-stats` container binds to localhost on CT 116.
- Admin Key List (step 8) requires the master key expanded locally before SSH, then embedded
in the remote curl command with proper quoting.
Expected output on a healthy fleet: 11/11 passing checks.