From bb17c2120fa6a23360225a64fc130da7458bbe03 Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 16:00:56 +0000 Subject: [PATCH] no-mistakes(document): Make litellm-health the live owner; dedupe probes --- README.md | 4 +-- contract-registry.yaml | 4 +-- litellm-health.prose.md | 9 +---- litellm-self-heal.prose.md | 68 ++++++-------------------------------- 4 files changed, 15 insertions(+), 70 deletions(-) diff --git a/README.md b/README.md index ac61914..c3860c2 100644 --- a/README.md +++ b/README.md @@ -116,7 +116,7 @@ Run on trigger or schedule. Maintain persistent world-model state across runs. | `zulip-health` | Zulip | Checks Zulip connectivity, message flow, and bot responsiveness. | | `zulip-mention-reliability` | Zulip | Diagnoses and fixes @mention detection issues in Zulip. | | `zulip-approval-fix` | Zulip | Fixes broken /approve and /deny slash commands for Hermes agents. | -| `litellm-self-heal` | LiteLLM | Consolidated health check + self-healing for the full nginx → LiteLLM → GPU chain. Verifies 8 containers, 3 GPUs, model inference, and agent keys. Applies 9 remediation rules. (litellm-health merged into this contract 2026-07-09.) | +| `litellm-self-heal` | LiteLLM | Applies remediation rules for LiteLLM stack failures detected by `litellm-health` (full nginx → LiteLLM → GPU chain). 9 remediation rules. | | `gpu-fleet` | GPU | Manages the GPU inference fleet: model deployment, registration, health checks, LiteLLM sync. | | `gpu-monitor` | GPU | Comprehensive GPU fleet monitor — polls sidecars, router, LiteLLM every 15s, renders SSE dashboard. | | `proxmox-monitor` | Infra | Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack on CT 116. | @@ -149,7 +149,7 @@ Called on-demand as single-render tools. | Contract | Description | |---|---| | `litellm-api-keys` | Manages LiteLLM API keys for agent identity. Create, rotate, verify, and list agent keys. References gpu-fleet for current key inventory. | -| `litellm-health` | ⚠️ **DEPRECATED** — consolidated into `litellm-self-heal` (2026-07-09). Retained for reference only. | +| `litellm-health` | LiteLLM health check: public vs backend surfaces, CT 116 containers, GPU fleet, one model per GPU host, and agent keys. Owner of the probes; `litellm-self-heal` owns remediation. | | `infrastructure-monitoring` | Target-state for Prometheus + GPU exporters + Grafana. Core stack deployed, GPU exporters NOT live. | | `stirling-pdf-agent-access` | Documents the Stirling-PDF API access pattern for agents — global API key, 12 operations, curl examples. Agents use the `stirling-pdf-api` shared skill for templates. | | `hello-world` | Minimal test contract — verifies the OpenProse execution pipeline works. | diff --git a/contract-registry.yaml b/contract-registry.yaml index 93f398a..019be2b 100644 --- a/contract-registry.yaml +++ b/contract-registry.yaml @@ -693,8 +693,8 @@ contracts: version: 1.0.0 trigger: type: scheduled - cadence: '*/10 * * * *' - description: "Every 10 minutes \u2014 LiteLLM proxy health" + cadence: '5 3,7,11,15,19,23 * * *' + description: "4-hourly staggered dispatch via fm-send (run contract litellm-health)" cron_job_id: null execution: agent: abiba diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 471eb54..4a430ca 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -1,14 +1,7 @@ --- kind: function name: litellm-health -status: deprecated -deprecated_on: 2026-07-09 -replaced_by: litellm-self-heal.prose.md -note: > - Consolidated into litellm-self-heal.prose.md to eliminate duplication - of architecture diagrams, GPU topology, timeout tables, and container - lists. Health check is now § Health Check within litellm-self-heal. - This file is retained for reference only — use litellm-self-heal instead. +status: active description: > Verifies the LiteLLM inference stack health. Current architecture (2026-07-09): nginx:80 → LiteLLM:4000 → GPU(llama-server) via direct proxy. diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 28ca553..189cdb0 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -11,22 +11,22 @@ note: > Script: `/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116 (cron `0 */6 * * *`). Reports to /var/log/litellm/health-*.json and Gitea (SyslogSolution/health-logs). GPU monitoring integrated from gpu-monitor on .24:9100. - - Consolidated from litellm-health + litellm-self-heal on 2026-07-09 to eliminate - duplication of architecture diagrams, GPU topology, timeout tables, and container - lists. Health check is now § Health Check within this contract. - + + Health probes are owned by litellm-health.prose.md (dispatched as + `run contract: litellm-health`). This contract owns remediation only — it does not + re-specify the probes. + Source of truth for GPU topology and keys: gpu-fleet.prose.md Last verified: 2026-07-12 description: > - LiteLLM inference stack health monitoring + self-healing. Verifies the full - nginx → LiteLLM → GPU chain, 12 containers on CT 116, 3 GPU hosts, model - inference, and agent keys. Applies remediation rules for common failures. + LiteLLM inference stack remediation. Applies remediation rules for failures detected + by litellm-health.prose.md (nginx → LiteLLM → GPU chain, CT 116 containers, GPU hosts, + model inference, and agent keys). Reports every action via Zulip DM and Gitea (SyslogSolution/health-logs). --- --- -# LiteLLM Operations — Health Check + Self-Heal +# LiteLLM Operations — Self-Heal (Remediation) ## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU) @@ -138,55 +138,7 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- ## Health Check -Run this first on every cycle. Results feed into remediation rules below. - -### 1. Check the end-user surfaces -The public edge and the backend edge serve the same app under different paths; they are not -interchangeable, so every probe names the surface it targets. - -Public edge — `{{public_url}}` serves the app at the ROOT (the `/litellm/` prefix 404s): -- GET {{public_url}}/ui/ → expect 200 ("LiteLLM Dashboard") -- GET {{public_url}}/docs → expect 200 ("LiteLLM API - Swagger UI") -- GET {{public_url}}/litellm/ui/ and {{public_url}}/litellm/docs → expect 404 (not served on this edge) - -Backend edge — `http://{{backend_host}}` (port 80) serves the app UNDER `/litellm/`: -- GET http://{{backend_host}}/litellm/ui/ → expect 200 ("LiteLLM Dashboard") -- GET http://{{backend_host}}/litellm/docs → expect 200 ("LiteLLM API - Swagger UI") -- GET http://{{backend_host}}/ui/ and /docs → expect 301 → /litellm/... → 200 (one-hop redirect helpers, added 2026-09-11) - -### 2. Check LiteLLM health (no-auth) -- GET http://{{backend_host}}/litellm/health/liveliness → expect 200 - -### 3. Check backend container health -- SSH to {{backend_host}} → `docker ps` → verify 12 containers healthy (11 harness + trove-agent-docker, added 2026-09-11) -- Critical: harness-litellm, harness-nginx, harness-postgres -- Monitoring: harness-redis, harness-dashboard, harness-grafana, harness-prometheus, - harness-alertmanager, harness-zulip-bridge, harness-docker-stats, harness-pve-exporter, - trove-agent-docker (ghcr.io/techdox/trove-agent-docker, added 2026-09-11) -- Decommissioned 2026-09-11: harness-router (container, image and config removed) - -### 4. Check GPU fleet health (via fleet dashboard) -- GET {{gpu_dashboard_url}}/gpu-data → expect 200 with GPU metrics JSON -- Verify GPUs reporting status "healthy" -- Check alerts array for active warnings/critical - -### 5. Check model inference via LiteLLM — test one model per GPU host (backend surface) -- POST http://{{backend_host}}/litellm/v1/chat/completions model=qwen3.6-27B-code → expect 200 (RTX 3090, .8) -- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) -- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) -- Auth uses the dedicated `monitor` agent key, read on CT 116 from `/etc/litellm-monitor.env` (root-only 0600) — never the master key - -### 6. Check agent keys -- GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys - -### 7. Check Grafana -- GET {{grafana_url}}/api/health → expect 200 - -### 8. Compile overall status -Determine overall_status from individual check results: -- "healthy" — all checks pass -- "degraded" — 1-2 non-critical checks fail -- "down" — critical checks fail +Health probes are owned by `litellm-health.prose.md` (dispatched as `run contract: litellm-health`). This contract owns remediation only — it does not re-specify the probes. --- ---