PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
1. Replace '(## Execution' + ')' with '## Execution' in 6 contract files 2. Make LOG_DIR honor CONTRACT_RUN_LOG_DIR env var (default: /var/log/contract-runs) 3. Fix Zulip alert: correct URL (https://chat.sysloggh.net/api/v1), user (abiba-bot@chat.sysloggh.net), take ZULIP_API_KEY from environment, and write alert failures to run log
143 lines
6.9 KiB
Markdown
143 lines
6.9 KiB
Markdown
---
|
|
kind: responsibility
|
|
name: agent-health-check
|
|
description: >
|
|
Consolidated agent health verification for the LiteLLM + GPU + Zulip +
|
|
gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 4 hours
|
|
via cron and on-demand via "run contract: agent-health-check". Verifies:
|
|
LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway
|
|
liveness, CT liveness, gateway log health, config YAML integrity,
|
|
wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything.
|
|
title: Agent Health Check — Consolidated
|
|
version: 1.0.0
|
|
runtime_contract: 2
|
|
agent: abiba
|
|
---
|
|
|
|
# Agent Health Check
|
|
|
|
Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet.
|
|
Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (`35 2,6,10,14,18,22 * * *`) and on-demand. Never restarts anything — detects and reports only.
|
|
|
|
**Cadence rationale** (2026-08-28 decision): monitoring dispatches moved from hourly to every 4 hours to reduce probe load on GPU hosts while keeping detection latency acceptable (up to 4 hours).
|
|
|
|
## Requires
|
|
|
|
- **LiteLLM admin key** for key validation (retrieved from `/root/.pi/agent/env.sh`)
|
|
- **SSH access** to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24)
|
|
- **Python 3** for script execution
|
|
- **Network access** to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints
|
|
|
|
## Maintains
|
|
|
|
- last_check: timestamp — When the last full diagnostic ran
|
|
- overall_severity: "healthy" | "degraded" | "critical"
|
|
- liteLLM_keys: map of agent → key validity
|
|
- gpu_ports: map of host → port conflict status
|
|
- agents: map of agent → streaming health + gateway liveness
|
|
- ct_liveness: map of CT → active status
|
|
- config_integrity: map of config file → valid/invalid
|
|
|
|
## Execution
|
|
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
|
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
|
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
|
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
|
execution — it only proves the agent read the result and reported it. The actual
|
|
monitoring work happens in the host cron job.
|
|
|
|
|
|
### check-health
|
|
|
|
**RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool
|
|
calls; never repeat a prior report unless a live probe fails.**
|
|
|
|
```bash
|
|
# Run the consolidated health check script
|
|
python3 /root/scripts/agent-health-check.py --json
|
|
```
|
|
|
|
**Report format**: Begin every report with the **absolute path the script
|
|
executed from** so a stale-consumer report is distinguishable from a real fault
|
|
at read time. Summarize actual results from each check. Apply the standing probe
|
|
rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.
|
|
|
|
**Expected output**: JSON with `overall_severity` field. If `healthy`, report
|
|
"Agent health check: OK". If `degraded` or `critical`, report the specific
|
|
failures and their severity.
|
|
|
|
**Mandatory report legs** (2026-09-17 decision, 1295.msg): Every report line
|
|
MUST include one clause per check leg, in every state: healthy, degraded/warn,
|
|
skipped, or failed. A missing leg must never look the same as a healthy leg.
|
|
Required legs and their templates in every state:
|
|
|
|
- `LiteLLM keys: 4/4 (tanko, abiba, koby, koonimo) valid`
|
|
- degraded: `LiteLLM keys: 2/4 (tanko valid; koby invalid; koonimo valid; abiba probe-failed: 192.168.68.116:4000 timeout)`
|
|
- skipped: `LiteLLM keys: SKIPPED (LiteLLM router unreachable)`
|
|
- `GPU ports: 3/3 (rtx3090, rtx5070, strixhalo) healthy`
|
|
- skipped: `GPU ports: SKIPPED (no SSH access to GPU hosts)`
|
|
- degraded/warn: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 degraded: svc=inactive, port owned by 1234; strixhalo healthy)`
|
|
- failed: `GPU ports: 2/3 (rtx3090 healthy; rtx5070 probe-failed: 192.168.68.110:9400 timeout; strixhalo healthy)`
|
|
- The GPU leg has six non-healthy states the code can produce:
|
|
(i) `gpu-unreachable:{host}` — SSH probe failed;
|
|
(ii) `gpu-no-port:{label}` — SSH worked, port not listening;
|
|
(iii) `gpu-ghost:{label}:{pid}` — unit inactive, port owned by another pid;
|
|
(iv) unit not active, MainPID empty or port owned by MainPID — svc inactive;
|
|
(v) unit active, /health body contains "error" — error response;
|
|
(vi) unit active, /health body unrecognised — unknown health.
|
|
In every case the failing host and reason must be named.
|
|
- `CTs: 4/4 running (tanko, abiba, koby, koonimo)`
|
|
- degraded: `CTs: 3/4 (tanko running; abiba running; koby probe-failed: ssh root@192.168.68.129 timeout; koonimo running)`
|
|
- skipped: `CTs: SKIPPED (SSH access unavailable)`
|
|
- `Vault secrets: 3/3 present`
|
|
- degraded: `Vault secrets: 2/3 (tanko present; koby present; koonimo missing)`
|
|
- skipped: `Vault secrets: SKIPPED (vault not configured)`
|
|
|
|
The compact form in the summary line is acceptable (e.g. `rtx5070 timeout`) as
|
|
long as the host is identifiable from context; the full `probe-failed: <target>
|
|
<kind>` form is required when a leg reports a failure in the detail section.
|
|
|
|
### Probe Shape (per standing rules from 1150.msg)
|
|
|
|
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
|
service answered — report the code, never "down". A redirect is not a failure.
|
|
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
|
|
2. **A failed probe is never a service verdict.** Print
|
|
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
|
|
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
|
|
then report.
|
|
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
|
|
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
|
|
(retried at 25s: also timeout)" is actionable.
|
|
|
|
## Strategies
|
|
|
|
### When LiteLLM keys are invalid
|
|
Report the specific agent + key name. Do not attempt to fix — credential
|
|
rotation is a separate operation.
|
|
|
|
### When GPU port conflicts are detected
|
|
Report the conflicting ports and processes. Do not kill processes — that's a
|
|
destructive action requiring captain approval.
|
|
|
|
### When gateway liveness is degraded
|
|
Report the specific CT + gateway status. Do not restart unless the restart
|
|
debounce window has passed.
|
|
|
|
### When CT liveness is down
|
|
Report the specific CT. Do not restart — that's a destructive action.
|
|
|
|
### When config YAML is invalid
|
|
Report the specific file + parse error. Do not fix — that's a config change.
|
|
|
|
### When gateway log health is degraded
|
|
Report the specific gateway + log health status (error patterns, stale connections, connectivity issues). Do not restart — that's a destructive action.
|
|
|
|
Note: the script may perform additional diagnostics beyond the seven contract checks listed under Execution.
|
|
|
|
## Continuity
|
|
|
|
- **Every 4 hours** (2, 6, 10, 14, 18, 22 UTC at :35): Scheduled cron check while Abiba is running
|
|
- **On `agent-health` command**: Run on-demand and report to user
|
|
- **On critical alert**: Escalate to relay message immediately
|