Merge pull request 'docs(contracts): add the missing agent-health-check contract' (#107) from fix/agent-health-check-contract-20260916 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 25s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s

This commit was merged in pull request #107.
This commit is contained in:
2026-09-16 05:03:26 +00:00
+104
View File
@@ -0,0 +1,104 @@
---
kind: responsibility
name: agent-health-check
description: >
Consolidated agent health verification for the LiteLLM + GPU + Zulip +
gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 4 hours
via cron and on-demand via "run contract: agent-health-check". Verifies:
LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway
liveness, CT liveness, gateway log health, config YAML integrity,
wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything.
title: Agent Health Check — Consolidated
version: 1.0.0
runtime_contract: 2
agent: abiba
---
# Agent Health Check
Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet.
Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18,22 * * *`) and on-demand. Never restarts anything — detects and reports only.
**Cadence rationale** (2026-08-28 decision): monitoring dispatches moved from hourly to every 4 hours to reduce probe load on GPU hosts while keeping detection latency acceptable (up to 4 hours).
## Requires
- **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`)
- **SSH access** to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24)
- **Python 3** for script execution
- **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints
## Maintains
- last_check: timestamp — When the last full diagnostic ran
- overall_severity: "healthy" | "degraded" | "critical"
- liteLLM_keys: map of agent → key validity
- gpu_ports: map of host → port conflict status
- agents: map of agent → streaming health + gateway liveness
- ct_liveness: map of CT → active status
- config_integrity: map of config file → valid/invalid
## Execution
### check-health
**RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool
calls; never repeat a prior report unless a live probe fails.**
```bash
# Run the consolidated health check script
python3 /root/scripts/agent-health-check.py --json
```
**Report format**: Begin every report with the **absolute path the script
executed from** so a stale-consumer report is distinguishable from a real fault
at read time. Summarize actual results from each check. Apply the standing probe
rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.
**Expected output**: JSON with `overall_severity` field. If `healthy`, report
"Agent health check: OK". If `degraded` or `critical`, report the specific
failures and their severity.
### Probe Shape (per standing rules from 1150.msg)
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
service answered — report the code, never "down". A redirect is not a failure.
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
2. **A failed probe is never a service verdict.** Print
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
then report.
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
(retried at 25s: also timeout)" is actionable.
## Strategies
### When LiteLLM keys are invalid
Report the specific agent + key name. Do not attempt to fix — credential
rotation is a separate operation.
### When GPU port conflicts are detected
Report the conflicting ports and processes. Do not kill processes — that's a
destructive action requiring captain approval.
### When gateway liveness is degraded
Report the specific CT + gateway status. Do not restart unless the restart
debounce window has passed.
### When CT liveness is down
Report the specific CT. Do not restart — that's a destructive action.
### When config YAML is invalid
Report the specific file + parse error. Do not fix — that's a config change.
### When gateway log health is degraded
Report the specific gateway + log health status (error patterns, stale connections, connectivity issues). Do not restart — that's a destructive action.
Note: the script may perform additional diagnostics beyond the seven contract checks listed under Execution.
## Continuity
- **Every 4 hours** (2, 6, 10, 14, 18, 22 UTC at :35): Scheduled cron check while Abiba is running
- **On `agent-health` command**: Run on-demand and report to user
- **On critical alert**: Escalate to relay message immediately