From 8a2ea2d0d7913d92315bb887da6af89c7fd2df07 Mon Sep 17 00:00:00 2001 From: root Date: Wed, 16 Sep 2026 04:36:43 +0000 Subject: [PATCH] docs: add agent-health-check.prose.md contract MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The consolidated agent health check contract (wraps scripts/agent-health-check.py v4). Created during earlier work but never committed — was a stray untracked file in the execution clone, making the home look dirty to the fleet update path. --- agent-health-check.prose.md | 98 +++++++++++++++++++++++++++++++++++++ 1 file changed, 98 insertions(+) create mode 100644 agent-health-check.prose.md diff --git a/agent-health-check.prose.md b/agent-health-check.prose.md new file mode 100644 index 0000000..9e845de --- /dev/null +++ b/agent-health-check.prose.md @@ -0,0 +1,98 @@ +--- +kind: responsibility +name: agent-health-check +description: > + Consolidated agent health verification for the LiteLLM + GPU + Zulip + + gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every + 10 minutes via cron and on-demand via "run contract: agent-health-check". + Verifies: LiteLLM key validity, GPU port conflicts, agent Zulip streaming, + gateway liveness, CT liveness, config YAML integrity, wrapper/CLI integrity, + vault secret non-emptiness. NEVER restarts anything. +title: Agent Health Check — Consolidated +version: 1.0.0 +runtime_contract: 2 +agent: abiba +--- + +# Agent Health Check + +Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet. +Runs every 10 minutes via cron (`*/10 * * * *`) and on-demand. Never restarts +anything — detects and reports only. + +## Requires + +- **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`) +- **SSH access** to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24) +- **Python 3** for script execution +- **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints + +## Maintains + +- last_check: timestamp — When the last full diagnostic ran +- overall_severity: "healthy" | "degraded" | "critical" +- liteLLM_keys: map of agent → key validity +- gpu_ports: map of host → port conflict status +- agents: map of agent → streaming health + gateway liveness +- ct_liveness: map of CT → active status +- config_integrity: map of config file → valid/invalid + +## Execution + +### check-health + +**RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool +calls; never repeat a prior report unless a live probe fails.** + +```bash +# Run the consolidated health check script +python3 /root/scripts/agent-health-check.py --json +``` + +**Report format**: Begin every report with the **absolute path the script +executed from** so a stale-consumer report is distinguishable from a real fault +at read time. Summarize actual results from each check. Apply the standing probe +rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed. + +**Expected output**: JSON with `overall_severity` field. If `healthy`, report +"Agent health check: OK". If `degraded` or `critical`, report the specific +failures and their severity. + +### Probe Shape (per standing rules from 1150.msg) + +1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the + service answered — report the code, never "down". A redirect is not a failure. + Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. +2. **A failed probe is never a service verdict.** Print + `probe-failed: ` naming the exact URL/host/port and the failure + kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only + then report. +3. **Say which probe produced each number.** "Grafana: 000" is unusable; + "Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s + (retried at 25s: also timeout)" is actionable. + +## Strategies + +### When LiteLLM keys are invalid +Report the specific agent + key name. Do not attempt to fix — credential +rotation is a separate operation. + +### When GPU port conflicts are detected +Report the conflicting ports and processes. Do not kill processes — that's a +destructive action requiring captain approval. + +### When gateway liveness is degraded +Report the specific CT + gateway status. Do not restart unless the restart +debounce window has passed. + +### When CT liveness is down +Report the specific CT. Do not restart — that's a destructive action. + +### When config YAML is invalid +Report the specific file + parse error. Do not fix — that's a config change. + +## Continuity + +- **Every 10 minutes**: Scheduled cron check while Abiba is running +- **On `agent-health` command**: Run on-demand and report to user +- **On critical alert**: Escalate to relay message immediately