PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
pm2-self-heal.prose.md: - Add AS-BUILT note: gpu-monitor is systemd-managed, NOT PM2 - gpu-watchdog is decommissioned and folded into gpu-monitor.service - gitea-runner is KEPT; abiba-zulip is KEPT (online for days) - spoton-service was deleted; live PM2 set is 4 processes - Preserve historical context for crash-loop guard litellm-health.prose.md: - Correct Prometheus node coverage: 6 nodes (.4/.5/.6/.9/.12/.15:9100) - Note .4:9100 is DEAD target (no route, down for weeks) - Clarify this does not read as 6 healthy nodes
204 lines
11 KiB
Markdown
204 lines
11 KiB
Markdown
---
|
||
kind: function
|
||
name: litellm-health
|
||
status: active
|
||
description: >
|
||
Verifies the LiteLLM inference stack health. Current architecture (2026-07-09):
|
||
nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy.
|
||
Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container,
|
||
image and config removed). GPU monitoring via Prometheus/Grafana and
|
||
fleet dashboard (gpu-monitor :9100).
|
||
Designed as a reusable contract for any Syslog agent.
|
||
|
||
Source of truth: gpu-fleet.prose.md
|
||
|
||
## Monitoring / Alerting (as-built 2026-08-09)
|
||
|
||
- LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]`
|
||
(failure_callback alone does NOT mount /metrics — verified 2026-08-09).
|
||
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
||
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
||
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
||
- Prometheus node job covers 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). Note: .4:9100 is a DEAD target (no route, down for weeks, not a live node).
|
||
---
|
||
|
||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||
|
||
```
|
||
Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
|
||
│
|
||
Key validation
|
||
Fallback chains
|
||
Budget tracking
|
||
│
|
||
Prometheus ← metrics
|
||
│
|
||
Grafana :3001
|
||
|
||
harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
|
||
and config removed). nginx routes /v1 → LiteLLM directly.
|
||
Router slot booking + circuit breakers replaced by
|
||
LiteLLM native fallbacks + timeouts.
|
||
```
|
||
|
||
**What changed (v3.2.0 → v4.0.0 — 2026-07-08)**:
|
||
- Router REMOVED from request path — LiteLLM proxies directly to GPU
|
||
- All GPUs at parallel 2 (was parallel 1)
|
||
- NVIDIA context reduced 256K→128K to free VRAM — the stable NVIDIA ceiling (2026-07-17); Strix Halo runs 256K (2026-09-12)
|
||
- Timeouts and fallback chains are config state — read them from CT 116 `/opt/inference-harness/litellm_config.yaml`; they are not duplicated here.
|
||
|
||
## Parameters
|
||
|
||
- public_url: string — The public LiteLLM URL (default: "https://litellm.sysloggh.net")
|
||
- backend_host: string — Internal CT host for container checks (default: "192.168.68.116")
|
||
- auth_host: string — Authentik server for OIDC (default: "192.168.68.11")
|
||
- gpu_hosts: array — GPU inference hosts (default: ["192.168.68.8", "192.168.68.110", "192.168.68.15"])
|
||
- gpu_dashboard_url: string — Fleet dashboard (default: "http://192.168.68.24:9100")
|
||
- grafana_url: string — Grafana dashboards (default: "http://192.168.68.116:3001")
|
||
|
||
## Returns
|
||
|
||
- overall_status: "healthy" | "degraded" | "down"
|
||
- checks: array of { name: string, status: string, detail: string }
|
||
- timestamp: string — ISO timestamp of the check run
|
||
- duration_ms: number — How long the check took
|
||
|
||
## Requires
|
||
|
||
- SSH key access to backend_host for container checks
|
||
- Network access to public_url, auth_host, and gpu_dashboard_url
|
||
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
|
||
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
|
||
model inference checks, scoped for every alias step 7 probes (`gpu-dense`, `gpu-vision`,
|
||
`strix-moe`, `syslog-auto`). Retrieve from the executor's host via:
|
||
|
||
```
|
||
monitor key: ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"
|
||
master key: ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"
|
||
```
|
||
|
||
If credentials are missing or unreadable, the probe must report `credential-missing` (not bare 401 or "0 keys").
|
||
The master key must never be used for inference.
|
||
|
||
## GPU Fleet Topology
|
||
|
||
| Host | IP | Hardware | Role |
|
||
|------|-----|----------|------|
|
||
| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) |
|
||
| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) |
|
||
| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) |
|
||
|
||
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
|
||
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
|
||
contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`,
|
||
`crew-auto`).
|
||
|
||
> Re-scope note (2026-09-12): the earlier plan to restate the live fallback chains and
|
||
> per-model timeouts in this contract is intentionally superseded — that state is config,
|
||
> and this contract points at the CT 116 config instead. Step 7 likewise authenticates with
|
||
> the dedicated `monitor` key, not the master key, which is admin-only.
|
||
|
||
## Containers on CT 116
|
||
|
||
| Container | Image | Port | Health Check |
|
||
|-----------|-------|------|-------------|
|
||
| harness-litellm | ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 | :4000→:4000 | /health/liveliness |
|
||
| harness-nginx | nginx:alpine | :80 | HTTP 200 on /health |
|
||
| harness-postgres | postgres:16-alpine | :5432 | pg_isready |
|
||
| harness-redis | redis:7-alpine | :6379 | PING |
|
||
| harness-dashboard | inference-harness-dashboard | :3000 | /health |
|
||
| harness-grafana | grafana/grafana | :3000→:3001 | /api/health |
|
||
| harness-prometheus | prom/prometheus | :9090 | /-/healthy |
|
||
|
||
## Execution
|
||
|
||
1. **Read parameters** — Use provided values or defaults
|
||
|
||
2. **Check the end-user surfaces** — the public edge and the backend edge serve the SAME
|
||
app under DIFFERENT paths. They are not interchangeable, so every probe below must name
|
||
the surface it targets. Never point a check at a path that only resolves on the other
|
||
surface.
|
||
|
||
**Public edge** — `{{public_url}}` (https://litellm.sysloggh.net) serves the app at the
|
||
ROOT; the `/litellm/` prefix does not exist there and 404s:
|
||
- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard")
|
||
- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI")
|
||
- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge)
|
||
|
||
**Backend edge** — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`:
|
||
- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
|
||
- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
|
||
- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers,
|
||
added 2026-09-11)
|
||
|
||
3. **Check LiteLLM health (no-auth)**:
|
||
- GET http://{{backend_host}}/litellm/health/liveliness → expect 200
|
||
|
||
4. **Check backend container health**:
|
||
- SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
|
||
- Critical: harness-litellm, harness-nginx, harness-postgres
|
||
- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus,
|
||
harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter,
|
||
trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
|
||
|
||
5. **Check GPU fleet health via gpu-monitor** (router decommissioned 2026-09-11):
|
||
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
||
- nginx `/health/unified` is now a `301` redirect to `/gpu/gpu-data` (same payload)
|
||
|
||
6. **Check GPU fleet health** (via fleet dashboard):
|
||
- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
|
||
- Verify GPUs reporting status "healthy"
|
||
- Check alerts array for active warnings/critical
|
||
|
||
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
|
||
check runs on the **backend edge**, not the public edge, so these paths carry the
|
||
`/litellm/` prefix:
|
||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
|
||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
|
||
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
|
||
- Auth uses the dedicated `monitor` agent key. Retrieve via:
|
||
`ssh root@192.168.68.116 "grep LITELLM_MONITOR_KEY /etc/litellm-monitor.env | cut -d= -f2"`
|
||
Do NOT use the master key for inference — the master key is for admin endpoints only
|
||
(`/key/list`, `/key/generate`, `/key/info`). Retrieve master key via:
|
||
`ssh root@192.168.68.116 "docker exec harness-litellm printenv LITELLM_MASTER_KEY"`
|
||
- KEY SCOPE: the `monitor` key MUST be scoped for the three probed aliases (`gpu-dense`,
|
||
`gpu-vision`, `strix-moe`) plus the `syslog-auto` fallback pool, otherwise the probe
|
||
returns 403 and the host is not covered.
|
||
If a probe returns 403, widen the monitor key's model list on CT 116 (add the missing
|
||
alias) and re-run — never drop the host from the probe to make the check pass.
|
||
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
|
||
this list. The RTX 5070 host now serves `gpu-vision`.
|
||
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
|
||
the set depends on the key. Always state which key a model list was read with — a
|
||
snapshot without its key is not evidence. This probe uses the `monitor` key on the
|
||
backend surface (`http://{{backend_host}}/litellm/v1/models`). The authoritative model
|
||
registry is CT 116 `/opt/inference-harness/litellm_config.yaml`; read it there rather
|
||
than freezing a list here.
|
||
|
||
8. **Check agent keys**:
|
||
- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys
|
||
- **IMPORTANT**: Run the curl on the CT 116 HOST, not inside the container. The `harness-litellm` container has no curl/wget. Use:
|
||
`ssh root@192.168.68.116 "curl -s -H 'Authorization: Bearer $MASTER_KEY' http://127.0.0.1:4000/key/list"`
|
||
- If the response is empty or unparseable, report `admin-call-failed` (not "0 agent keys")
|
||
|
||
9. **Check Grafana**:
|
||
- GET {{grafana_url}}/api/health → expect 200
|
||
|
||
10. **Compile and report** — Determine overall_status from individual check results
|
||
|
||
## Executor Script (2026-09-13)
|
||
|
||
**Run `scripts/litellm-health-check.py` from the clone.** This script implements all 11
|
||
checks defined above and reports results in a standardized format. Paste its output in
|
||
the status line.
|
||
|
||
- Hand-rolled probes are **not** an acceptable substitute for the script.
|
||
- Backend-edge checks (steps 2–8) must use `http://192.168.68.116` (internal IP),
|
||
**not** the public URL `https://litellm.sysloggh.net` (which returns 401 for those paths).
|
||
- Docker Stats (step 10) must be fetched from the CT 116 host itself (`127.0.0.1:9324/metrics`)
|
||
because the `harness-docker-stats` container binds to localhost on CT 116.
|
||
- Admin Key List (step 8) requires the master key expanded locally before SSH, then embedded
|
||
in the remote curl command with proper quoting.
|
||
|
||
Expected output on a healthy fleet: 11/11 passing checks.
|