Files
prose-contracts/agent-health-check.prose.md
T
root 8a2ea2d0d7
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
docs: add agent-health-check.prose.md contract
The consolidated agent health check contract (wraps scripts/agent-health-check.py v4).
Created during earlier work but never committed — was a stray untracked file in
the execution clone, making the home look dirty to the fleet update path.
2026-09-16 04:36:43 +00:00

3.8 KiB

kind, name, description, title, version, runtime_contract, agent
kind name description title version runtime_contract agent
responsibility agent-health-check Consolidated agent health verification for the LiteLLM + GPU + Zulip + gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 10 minutes via cron and on-demand via "run contract: agent-health-check". Verifies: LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway liveness, CT liveness, config YAML integrity, wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything. Agent Health Check — Consolidated 1.0.0 2 abiba

Agent Health Check

Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet. Runs every 10 minutes via cron (*/10 * * * *) and on-demand. Never restarts anything — detects and reports only.

Requires

  • LiteLLM admin key for key validation (retrieved from /root/.pi/agent/env.sh)
  • SSH access to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24)
  • Python 3 for script execution
  • Network access to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints

Maintains

  • last_check: timestamp — When the last full diagnostic ran
  • overall_severity: "healthy" | "degraded" | "critical"
  • liteLLM_keys: map of agent → key validity
  • gpu_ports: map of host → port conflict status
  • agents: map of agent → streaming health + gateway liveness
  • ct_liveness: map of CT → active status
  • config_integrity: map of config file → valid/invalid

Execution

check-health

RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool calls; never repeat a prior report unless a live probe fails.

# Run the consolidated health check script
python3 /root/scripts/agent-health-check.py --json

Report format: Begin every report with the absolute path the script executed from so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual results from each check. Apply the standing probe rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.

Expected output: JSON with overall_severity field. If healthy, report "Agent health check: OK". If degraded or critical, report the specific failures and their severity.

Probe Shape (per standing rules from 1150.msg)

  1. Any HTTP status means ALIVE. 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
  2. A failed probe is never a service verdict. Print probe-failed: <target> <kind> naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
  3. Say which probe produced each number. "Grafana: 000" is unusable; "Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.

Strategies

When LiteLLM keys are invalid

Report the specific agent + key name. Do not attempt to fix — credential rotation is a separate operation.

When GPU port conflicts are detected

Report the conflicting ports and processes. Do not kill processes — that's a destructive action requiring captain approval.

When gateway liveness is degraded

Report the specific CT + gateway status. Do not restart unless the restart debounce window has passed.

When CT liveness is down

Report the specific CT. Do not restart — that's a destructive action.

When config YAML is invalid

Report the specific file + parse error. Do not fix — that's a config change.

Continuity

  • Every 10 minutes: Scheduled cron check while Abiba is running
  • On agent-health command: Run on-demand and report to user
  • On critical alert: Escalate to relay message immediately