Files
prose-contracts/litellm-health.prose.md
T
agent-zero 7730cc7c03
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
feat(litellm): 1.99.1 update coverage, trove agent, nginx /ui /docs path fixes
2026-09-11 15:32:20 -04:00

6.9 KiB

kind, name, status, deprecated_on, replaced_by, note, description
kind name status deprecated_on replaced_by note description
function litellm-health deprecated 2026-07-09 litellm-self-heal.prose.md Consolidated into litellm-self-heal.prose.md to eliminate duplication of architecture diagrams, GPU topology, timeout tables, and container lists. Health check is now § Health Check within litellm-self-heal. This file is retained for reference only — use litellm-self-heal instead. Verifies the LiteLLM inference stack health. Current architecture (2026-07-09): nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy. Router (harness-router :9000) was DECOMMISSIONED 2026-09-11 (container, image and config removed). GPU monitoring via Prometheus/Grafana and fleet dashboard (gpu-monitor :9100). Designed as a reusable contract for any Syslog agent. Source of truth: gpu-fleet.prose.md ## Monitoring / Alerting (as-built 2026-08-09) - LiteLLM /metrics scrape job: requires `litellm_settings.success_callback: [prometheus]` (failure_callback alone does NOT mount /metrics — verified 2026-08-09). Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/. - Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09. - Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100).

Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)

Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
                        │
                  Key validation
                  Fallback chains
                  Budget tracking
                        │
                  Prometheus ← metrics
                        │
                    Grafana :3001

  harness-router :9000 — DECOMMISSIONED 2026-09-11 (container, image
  and config removed). nginx routes /v1 → LiteLLM directly.
  Router slot booking + circuit breakers replaced by
  LiteLLM native fallbacks + timeouts.

What changed (v3.2.0 → v4.0.0 — 2026-07-08):

  • Router REMOVED from request path — LiteLLM proxies directly to GPU
  • All GPUs at parallel 2 (was parallel 1)
  • NVIDIA context reduced 256K→128K to free VRAM — now the stable ceiling across all GPUs (2026-07-17)
  • LiteLLM timeouts tuned: gemma 25→120s, qwen 40→90s (SUPERSEDED 2026-07-16: qwen 300s, gemma 120s, strix 300s — see litellm-self-heal)
  • nginx proxy_read_timeout: 600s, LiteLLM request_timeout: 300s

Parameters

  • public_url: string — The public LiteLLM URL (default: "https://litellm.sysloggh.net")
  • backend_host: string — Internal CT host for container checks (default: "192.168.68.116")
  • auth_host: string — Authentik server for OIDC (default: "192.168.68.11")
  • gpu_hosts: array — GPU inference hosts (default: ["192.168.68.8", "192.168.68.110", "192.168.68.15"])
  • gpu_dashboard_url: string — Fleet dashboard (default: "http://192.168.68.24:9100")
  • grafana_url: string — Grafana dashboards (default: "http://192.168.68.116:3001")

Returns

  • overall_status: "healthy" | "degraded" | "down"
  • checks: array of { name: string, status: string, detail: string }
  • timestamp: string — ISO timestamp of the check run
  • duration_ms: number — How long the check took

Requires

  • SSH key access to backend_host for container checks
  • Network access to public_url, auth_host, and gpu_dashboard_url
  • LiteLLM master key for key management endpoints

GPU Fleet Topology

Host IP Hardware Models Served Engine Context Parallel
llm-gpu 192.168.68.8 NVIDIA RTX 3090 (24 GB) qwen3.6-27B-code llama-server systemd 128K 2
ocu-llm 192.168.68.110 NVIDIA RTX 5070 (12 GB) gemma-4-12b llama-server systemd 128K 2
amdpve 192.168.68.15 AMD Strix Halo 64GB UMA qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) llama-server systemd (Vulkan) 128K 2

Model Fallback Chains (LiteLLM)

Primary Timeout Fallback Timeout
qwen3.6-27B-code 300s gemma-4-12b 120s
gemma-4-12b 120s qwen3.6-27B-code 300s
qwen3.6-35B-udq4 / strix-moe 300s qwen → gemma —
syslog-auto (balanced) 300s qwen → gemma —

Global: request_timeout=300s, nginx proxy_read_timeout=600s

Containers on CT 116

Container Image Port Health Check
harness-litellm ghcr.io/docker.litellm.ai/berriai/litellm:1.99.1 :4000→:4000 /health/liveliness
harness-nginx nginx:alpine :80 HTTP 200 on /health
harness-postgres postgres:16-alpine :5432 pg_isready
harness-redis redis:7-alpine :6379 PING
harness-dashboard inference-harness-dashboard :3000 /health
harness-grafana grafana/grafana :3000→:3001 /api/health
harness-prometheus prom/prometheus :9090 /-/healthy

Execution

  1. Read parameters — Use provided values or defaults

  2. Check public endpoints:

    • GET {{public_url}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard")
    • GET {{public_url}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI")
    • GET {{public_url}}/ui/ and {{public_url}}/docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11)
  3. Check LiteLLM health (no-auth):

  4. Check backend container health:

    • SSH to {{backend_host}} → docker ps → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11)
    • Critical: harness-litellm, harness-nginx, harness-postgres
    • Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter, trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11)
  5. Check GPU fleet health via gpu-monitor (router decommissioned 2026-09-11):

    • GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
    • nginx /health/unified is now a 301 redirect to /gpu/gpu-data (same payload)
  6. Check GPU fleet health (via fleet dashboard):

    • GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON
    • Verify GPUs reporting status "healthy"
    • Check alerts array for active warnings/critical
  7. Check model inference via LiteLLM — Test each model:

    • POST /v1/chat/completions model=gemma-4-12b → expect 200
    • POST /v1/chat/completions model=qwen3.6-27B-code → expect 200
    • POST /v1/chat/completions model=strix-moe → expect 200
    • Use master key for auth
  8. Check agent keys:

    • GET /key/list with master key → verify all 6 agents have keys
  9. Check Grafana:

    • GET {{grafana_url}}/api/health → expect 200
  10. Compile and report — Determine overall_status from individual check results