--- kind: responsibility name: agent-health-check description: > Consolidated agent health verification for the LiteLLM + GPU + Zulip + gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 4 hours via cron and on-demand via "run contract: agent-health-check". Verifies: LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway liveness, CT liveness, gateway log health, config YAML integrity, wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything. title: Agent Health Check — Consolidated version: 1.0.0 runtime_contract: 2 agent: abiba --- # Agent Health Check Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet. Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18,22 * * *`) and on-demand. Never restarts anything — detects and reports only. **Cadence rationale** (2026-08-28 decision): monitoring dispatches moved from hourly to every 4 hours to reduce probe load on GPU hosts while keeping detection latency acceptable (up to 4 hours). ## Requires - **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`) - **SSH access** to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24) - **Python 3** for script execution - **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints ## Maintains - last_check: timestamp — When the last full diagnostic ran - overall_severity: "healthy" | "degraded" | "critical" - liteLLM_keys: map of agent → key validity - gpu_ports: map of host → port conflict status - agents: map of agent → streaming health + gateway liveness - ct_liveness: map of CT → active status - config_integrity: map of config file → valid/invalid ## Execution ### check-health **RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool calls; never repeat a prior report unless a live probe fails.** ```bash # Run the consolidated health check script python3 /root/scripts/agent-health-check.py --json ``` **Report format**: Begin every report with the **absolute path the script executed from** so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual results from each check. Apply the standing probe rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed. **Expected output**: JSON with `overall_severity` field. If `healthy`, report "Agent health check: OK". If `degraded` or `critical`, report the specific failures and their severity. **Mandatory report legs** (2026-09-17 decision, 1295.msg): Every report line MUST include one clause per check leg, even when a leg is skipped or fails. A missing leg must never look the same as a healthy leg. Required legs and their templates in every state (healthy, skipped, partial): - `LiteLLM keys: 4/4 (tanko, abiba, koby, koonimo) valid` - partial: `LiteLLM keys: 3/4 (tanko valid; koby invalid; koonimo valid; abiba probe-failed: 192.168.68.116:4000 timeout)` - `GPU ports: 3/3 (rtx3090, rtx5070, strixhalo) healthy` - skipped: `GPU ports: SKIPPED (no SSH access to GPU hosts)` - partial: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 probe-failed: 192.168.68.110:9400 timeout; strixhalo healthy)` - The GPU leg degrades when a probe fails OR when the port is not listening (`gpu-no-port`) OR when the port is owned by a process other than the service's (`gpu-ghost`) — in every case the failing host and reason must be named. - `CTs: 4/4 running (tanko, abiba, koby, koonimo)` - partial: `CTs: 3/4 (tanko running; abiba running; koby probe-failed: ssh root@192.168.68.129 timeout; koonimo running)` - `Vault secrets: 3/3 present` - partial: `Vault secrets: 2/3 (tanko present; koby present; koonimo missing)` The compact form in the summary line is acceptable (e.g. `rtx5070 timeout`) as long as the host is identifiable from context; the full `probe-failed: ` form is required when a leg reports a failure in the detail section. ### Probe Shape (per standing rules from 1150.msg) 1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe. 2. **A failed probe is never a service verdict.** Print `probe-failed: ` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report. 3. **Say which probe produced each number.** "Grafana: 000" is unusable; "Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s (retried at 25s: also timeout)" is actionable. ## Strategies ### When LiteLLM keys are invalid Report the specific agent + key name. Do not attempt to fix — credential rotation is a separate operation. ### When GPU port conflicts are detected Report the conflicting ports and processes. Do not kill processes — that's a destructive action requiring captain approval. ### When gateway liveness is degraded Report the specific CT + gateway status. Do not restart unless the restart debounce window has passed. ### When CT liveness is down Report the specific CT. Do not restart — that's a destructive action. ### When config YAML is invalid Report the specific file + parse error. Do not fix — that's a config change. ### When gateway log health is degraded Report the specific gateway + log health status (error patterns, stale connections, connectivity issues). Do not restart — that's a destructive action. Note: the script may perform additional diagnostics beyond the seven contract checks listed under Execution. ## Continuity - **Every 4 hours** (2, 6, 10, 14, 18, 22 UTC at :35): Scheduled cron check while Abiba is running - **On `agent-health` command**: Run on-demand and report to user - **On critical alert**: Escalate to relay message immediately