Files
prose-contracts/agent-health-check.prose.md
T
root 5c1c8d7c19
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
fix: PR #107 review fixes — cron cadence + gateway log health check
FAIL 1: Cron cadence was */10 * * * * (every 10 min) but the real crontab
on CT 100 is 35 2,6,10,14,18,22 * * * (every 4 hours at :35). Fixed in
frontmatter, body, and Continuity section. Added 4-hour rationale note.

FAIL 2: Added gateway log health to the list of checks (frontmatter +
Strategies section). Added note that script may perform additional
diagnostics beyond the seven contract checks.
2026-09-16 04:50:50 +00:00

4.4 KiB

kind, name, description, title, version, runtime_contract, agent
kind name description title version runtime_contract agent
responsibility agent-health-check Consolidated agent health verification for the LiteLLM + GPU + Zulip + gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 4 hours via cron and on-demand via "run contract: agent-health-check". Verifies: LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway liveness, CT liveness, gateway log health, config YAML integrity, wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything. Agent Health Check — Consolidated 1.0.0 2 abiba

Agent Health Check

Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet. Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (35 2,6,10,14,18,22 * * *) and on-demand. Never restarts anything — detects and reports only.

Cadence rationale (2026-08-28 decision): monitoring dispatches moved from hourly to every 4 hours to reduce probe load on GPU hosts while keeping detection latency acceptable (up to 4 hours).

Requires

  • LiteLLM admin key for key validation (retrieved from /root/.pi/agent/env.sh)
  • SSH access to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24)
  • Python 3 for script execution
  • Network access to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints

Maintains

  • last_check: timestamp — When the last full diagnostic ran
  • overall_severity: "healthy" | "degraded" | "critical"
  • liteLLM_keys: map of agent → key validity
  • gpu_ports: map of host → port conflict status
  • agents: map of agent → streaming health + gateway liveness
  • ct_liveness: map of CT → active status
  • config_integrity: map of config file → valid/invalid

Execution

check-health

RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool calls; never repeat a prior report unless a live probe fails.

# Run the consolidated health check script
python3 /root/scripts/agent-health-check.py --json

Report format: Begin every report with the absolute path the script executed from so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual results from each check. Apply the standing probe rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.

Expected output: JSON with overall_severity field. If healthy, report "Agent health check: OK". If degraded or critical, report the specific failures and their severity.

Probe Shape (per standing rules from 1150.msg)

  1. Any HTTP status means ALIVE. 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
  2. A failed probe is never a service verdict. Print probe-failed: <target> <kind> naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report.
  3. Say which probe produced each number. "Grafana: 000" is unusable; "Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.

Strategies

When LiteLLM keys are invalid

Report the specific agent + key name. Do not attempt to fix — credential rotation is a separate operation.

When GPU port conflicts are detected

Report the conflicting ports and processes. Do not kill processes — that's a destructive action requiring captain approval.

When gateway liveness is degraded

Report the specific CT + gateway status. Do not restart unless the restart debounce window has passed.

When CT liveness is down

Report the specific CT. Do not restart — that's a destructive action.

When config YAML is invalid

Report the specific file + parse error. Do not fix — that's a config change.

When gateway log health is degraded

Report the specific gateway + log health status (error patterns, stale connections, connectivity issues). Do not restart — that's a destructive action.

Note: the script may perform additional diagnostics beyond the seven contract checks listed under Execution.

Continuity

  • Every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35): Scheduled cron check while Abiba is running
  • On agent-health command: Run on-demand and report to user
  • On critical alert: Escalate to relay message immediately