diff --git a/agent-health-check.prose.md b/agent-health-check.prose.md new file mode 100644 index 0000000..cdc95cd --- /dev/null +++ b/agent-health-check.prose.md @@ -0,0 +1,104 @@ +--- +kind: responsibility +name: agent-health-check +description: > + Consolidated agent health verification for the LiteLLM + GPU + Zulip + + gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 4 hours + via cron and on-demand via "run contract: agent-health-check". Verifies: + LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway + liveness, CT liveness, gateway log health, config YAML integrity, + wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything. +title: Agent Health Check — Consolidated +version: 1.0.0 +runtime_contract: 2 +agent: abiba +--- + +# Agent Health Check + +Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet. +Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18,22 * * *`) and on-demand. Never restarts anything — detects and reports only. + +**Cadence rationale** (2026-08-28 decision): monitoring dispatches moved from hourly to every 4 hours to reduce probe load on GPU hosts while keeping detection latency acceptable (up to 4 hours). + +## Requires + +- **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`) +- **SSH access** to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24) +- **Python 3** for script execution +- **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints + +## Maintains + +- last_check: timestamp — When the last full diagnostic ran +- overall_severity: "healthy" | "degraded" | "critical" +- liteLLM_keys: map of agent → key validity +- gpu_ports: map of host → port conflict status +- agents: map of agent → streaming health + gateway liveness +- ct_liveness: map of CT → active status +- config_integrity: map of config file → valid/invalid + +## Execution + +### check-health + +**RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool +calls; never repeat a prior report unless a live probe fails.** + +```bash +# Run the consolidated health check script +python3 /root/scripts/agent-health-check.py --json +``` + +**Report format**: Begin every report with the **absolute path the script +executed from** so a stale-consumer report is distinguishable from a real fault +at read time. Summarize actual results from each check. Apply the standing probe +rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed. + +**Expected output**: JSON with `overall_severity` field. If `healthy`, report +"Agent health check: OK". If `degraded` or `critical`, report the specific +failures and their severity. + +### Probe Shape (per standing rules from 1150.msg) + +1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the + service answered — report the code, never "down". A redirect is not a failure. + Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. +2. **A failed probe is never a service verdict.** Print + `probe-failed: ` naming the exact URL/host/port and the failure + kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only + then report. +3. **Say which probe produced each number.** "Grafana: 000" is unusable; + "Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s + (retried at 25s: also timeout)" is actionable. + +## Strategies + +### When LiteLLM keys are invalid +Report the specific agent + key name. Do not attempt to fix — credential +rotation is a separate operation. + +### When GPU port conflicts are detected +Report the conflicting ports and processes. Do not kill processes — that's a +destructive action requiring captain approval. + +### When gateway liveness is degraded +Report the specific CT + gateway status. Do not restart unless the restart +debounce window has passed. + +### When CT liveness is down +Report the specific CT. Do not restart — that's a destructive action. + +### When config YAML is invalid +Report the specific file + parse error. Do not fix — that's a config change. + +### When gateway log health is degraded +Report the specific gateway + log health status (error patterns, stale connections, connectivity issues). Do not restart — that's a destructive action. + +Note: the script may perform additional diagnostics beyond the seven contract checks listed under Execution. + +## Continuity + +- **Every 4 hours** (2, 6, 10, 14, 18, 22 UTC at :35): Scheduled cron check while Abiba is running +- **On `agent-health` command**: Run on-demand and report to user +- **On critical alert**: Escalate to relay message immediately