169 lines
8.6 KiB
Markdown
169 lines
8.6 KiB
Markdown
---
|
|
kind: function
|
|
name: litellm-health
|
|
status: active
|
|
description: >
|
|
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
|
|
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
|
|
Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container,
|
|
image and config removed). GPU monitoring via Prometheus/Grafana and
|
|
fleet dashboard (gpu-monitor :9100).
|
|
Designed as a reusable contract for any Syslog agent.
|
|
|
|
Source of truth: gpu-fleet.prose.md
|
|
|
|
## Monitoring / Alerting (as-built 2026-08-09)
|
|
|
|
- LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]`
|
|
(failure_callback alone does NOT mount /metrics — verified 2026-08-09).
|
|
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
|
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
|
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
|
- Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100).
|
|
---
|
|
|
|
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
|
|
|
```
|
|
Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
|
│
|
|
Key validation
|
|
Fallback chains
|
|
Budget tracking
|
|
│
|
|
Prometheus ← metrics
|
|
│
|
|
Grafana :3001
|
|
|
|
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
|
|
and config removed). nginx routes /v1 → LiteLLM directly.
|
|
Router slot booking + circuit breakers replaced by
|
|
LiteLLM native fallbacks + timeouts.
|
|
```
|
|
|
|
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
|
|
- Router REMOVED from request path — LiteLLM proxies directly to GPU
|
|
- All GPUs at parallel 2 (was parallel 1)
|
|
- NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
|
|
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
|
|
|
|
## Parameters
|
|
|
|
- public_url: string — The public LiteLLM URL (default: "https://litellm.sysloggh.net")
|
|
- backend_host: string — Internal CT host for container checks (default: "192.168.68.116")
|
|
- auth_host: string — Authentik server for OIDC (default: "192.168.68.11")
|
|
- gpu_hosts: array — GPU inference hosts (default: ["192.168.68.8", "192.168.68.110", "192.168.68.15"])
|
|
- gpu_dashboard_url: string — Fleet dashboard (default: "http://192.168.68.24:9100")
|
|
- grafana_url: string — Grafana dashboards (default: "http://192.168.68.116:3001")
|
|
|
|
## Returns
|
|
|
|
- overall_status: "healthy" | "degraded" | "down"
|
|
- checks: array of { name: string, status: string, detail: string }
|
|
- timestamp: string — ISO timestamp of the check run
|
|
- duration_ms: number — How long the check took
|
|
|
|
## Requires
|
|
|
|
- SSH key access to backend_host for container checks
|
|
- Network access to public_url, auth_host, and gpu_dashboard_url
|
|
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
|
|
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
|
|
model inference checks — the master key must never be used for inference
|
|
|
|
## GPU Fleet Topology
|
|
|
|
| Host | IP | Hardware | Role |
|
|
|------|-----|----------|------|
|
|
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
|
|
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
|
|
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
|
|
|
|
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
|
|
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
|
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
|
|
`crew-auto`).
|
|
|
|
> Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and
|
|
> per-model timeouts in this contract is intentionally superseded — that state is config,
|
|
> and this contract points at the CT 116 config instead. Step 7 likewise authenticates with
|
|
> the dedicated `monitor` key, not the master key, which is admin-only.
|
|
|
|
## Containers on CT 116
|
|
|
|
| Container | Image | Port | Health Check |
|
|
|-----------|-------|------|-------------|
|
|
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
|
|
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
|
|
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
|
|
| harness-redis | redis:7-alpine | :6379 | PING |
|
|
| harness-dashboard | inference-harness-dashboard | :3000 | /health |
|
|
| harness-grafana | grafana/grafana | :3000→:3001 | /api/health |
|
|
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
|
|
|
|
## Execution
|
|
|
|
1. **Read parameters** — Use provided values or defaults
|
|
|
|
2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME
|
|
app under DIFFERENT paths. They are not interchangeable, so every probe below must name
|
|
the surface it targets. Never point a check at a path that only resolves on the other
|
|
surface.
|
|
|
|
**Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the
|
|
ROOT; the `/litellm/` prefix does not exist there and 404s:
|
|
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
|
|
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
|
|
- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)
|
|
|
|
**Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`:
|
|
- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
|
|
- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
|
|
- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers,
|
|
added 2026-09-11)
|
|
|
|
3. **Check LiteLLM health (no-auth)**:
|
|
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
|
|
|
4. **Check backend container health**:
|
|
- SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
|
|
- Critical: harness-litellm, harness-nginx, harness-postgres
|
|
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
|
|
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter,
|
|
trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
|
|
|
|
5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11):
|
|
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
|
- nginx `/health/unified` is now a `301` redirect to `/gpu/gpu-data` (same payload)
|
|
|
|
6. **Check GPU fleet health** (via fleet dashboard):
|
|
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
|
- Verify GPUs reporting status "healthy"
|
|
- Check alerts array for active warnings/critical
|
|
|
|
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
|
|
check runs on the **backend edge**, not the public edge, so these paths carry the
|
|
`/litellm/` prefix:
|
|
- POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8)
|
|
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
|
|
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
|
|
- Auth uses the dedicated `monitor` agent key, read on CT 116 from
|
|
`/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference —
|
|
the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`).
|
|
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
|
|
this list. The RTX 5070 host now serves `gpu-vision`.
|
|
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
|
|
the set depends on the key. Always state which key a model list was read with — a
|
|
snapshot without its key is not evidence. This probe uses the `monitor` key on the
|
|
backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model
|
|
registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather
|
|
than freezing a list here.
|
|
|
|
8. **Check agent keys**:
|
|
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
|
|
|
|
9. **Check Grafana**:
|
|
- GET {{grafana_url}}/api/health → expect 200
|
|
|
|
10. **Compile and report** — Determine overall_status from individual check results
|