- Implement four-state PBS GC logic in proxmox-monitor.sh: - probe-failed: unparseable JSON, store not found, or empty body → FAIL - running: collection in progress (last-run-endtime absent, upid present) → DO NOT FAIL - stale: no completed run within 48h → FAIL, naming last completed run age - healthy: completed within 48h → PASS, naming endtime and pending bytes - Add tests/test_pbs_gc_states.sh covering all four states - Proves the test bites on the pre-fix version (5/6 tests fail) - All 6 tests pass against the fixed version - Update contract-run.sh to map disk-gc-threat-response -> scripts/disk-gc-scan.py - Update Execution sections of host-scheduled contracts: infrastructure-monitoring, zulip-health, litellm-health, agent-health-check, disk-gc-threat-response, pm2-self-heal Adding note that execution is host-scheduled via cron, not agent session ack.
6.9 KiB
kind, name, description, title, version, runtime_contract, agent
| kind | name | description | title | version | runtime_contract | agent |
|---|---|---|---|---|---|---|
| responsibility | agent-health-check | Consolidated agent health verification for the LiteLLM + GPU + Zulip + gateway fleet. Wraps scripts/agent-health-check.py (v4). Runs every 4 hours via cron and on-demand via "run contract: agent-health-check". Verifies: LiteLLM key validity, GPU port conflicts, agent Zulip streaming, gateway liveness, CT liveness, gateway log health, config YAML integrity, wrapper/CLI integrity, vault secret non-emptiness. NEVER restarts anything. | Agent Health Check — Consolidated | 1.0.0 | 2 | abiba |
Agent Health Check
Consolidated health verification for the LiteLLM + GPU + Zulip + gateway fleet.
Runs every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35) via cron (35 2,6,10,14,18,22 * * *) and on-demand. Never restarts anything — detects and reports only.
Cadence rationale (2026-08-28 decision): monitoring dispatches moved from hourly to every 4 hours to reduce probe load on GPU hosts while keeping detection latency acceptable (up to 4 hours).
Requires
- LiteLLM admin key for key validation (retrieved from
/root/.pi/agent/env.sh) - SSH access to GPU hosts (.8, .110, .15) and agent CTs (.122, .129, .114, .24)
- Python 3 for script execution
- Network access to LiteLLM (:4000), GPU exporters (:9400), and gateway endpoints
Maintains
- last_check: timestamp — When the last full diagnostic ran
- overall_severity: "healthy" | "degraded" | "critical"
- liteLLM_keys: map of agent → key validity
- gpu_ports: map of host → port conflict status
- agents: map of agent → streaming health + gateway liveness
- ct_liveness: map of CT → active status
- config_integrity: map of config file → valid/invalid
(## Execution
)
Execution model: This contract is executed by a host-scheduled cron job (see
scripts/contract-run.sh). The cron job runs the monitoring script directly on the
target host and appends the result to /var/log/contract-runs/<contract>.log. An
agent-session acknowledgement (a done: line in the ops status log) is NOT
execution — it only proves the agent read the result and reported it. The actual
monitoring work happens in the host cron job.
check-health
RUN LIVE, NEVER ECHO — every dispatch must execute the script with real tool calls; never repeat a prior report unless a live probe fails.
# Run the consolidated health check script
python3 /root/scripts/agent-health-check.py --json
Report format: Begin every report with the absolute path the script executed from so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual results from each check. Apply the standing probe rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed.
Expected output: JSON with overall_severity field. If healthy, report
"Agent health check: OK". If degraded or critical, report the specific
failures and their severity.
Mandatory report legs (2026-09-17 decision, 1295.msg): Every report line MUST include one clause per check leg, in every state: healthy, degraded/warn, skipped, or failed. A missing leg must never look the same as a healthy leg. Required legs and their templates in every state:
LiteLLM keys: 4/4 (tanko, abiba, koby, koonimo) valid- degraded:
LiteLLM keys: 2/4 (tanko valid; koby invalid; koonimo valid; abiba probe-failed: 192.168.68.116:4000 timeout) - skipped:
LiteLLM keys: SKIPPED (LiteLLM router unreachable)
- degraded:
GPU ports: 3/3 (rtx3090, rtx5070, strixhalo) healthy- skipped:
GPU ports: SKIPPED (no SSH access to GPU hosts) - degraded/warn:
GPU ports: 2/3 (rtx3090 healthy; rtx5070 degraded: svc=inactive, port owned by 1234; strixhalo healthy) - failed:
GPU ports: 2/3 (rtx3090 healthy; rtx5070 probe-failed: 192.168.68.110:9400 timeout; strixhalo healthy) - The GPU leg has six non-healthy states the code can produce:
(i)
gpu-unreachable:{host}— SSH probe failed; (ii)gpu-no-port:{label}— SSH worked, port not listening; (iii)gpu-ghost:{label}:{pid}— unit inactive, port owned by another pid; (iv) unit not active, MainPID empty or port owned by MainPID — svc inactive; (v) unit active, /health body contains "error" — error response; (vi) unit active, /health body unrecognised — unknown health. In every case the failing host and reason must be named.
- skipped:
CTs: 4/4 running (tanko, abiba, koby, koonimo)- degraded:
CTs: 3/4 (tanko running; abiba running; koby probe-failed: ssh root@192.168.68.129 timeout; koonimo running) - skipped:
CTs: SKIPPED (SSH access unavailable)
- degraded:
Vault secrets: 3/3 present- degraded:
Vault secrets: 2/3 (tanko present; koby present; koonimo missing) - skipped:
Vault secrets: SKIPPED (vault not configured)
- degraded:
The compact form in the summary line is acceptable (e.g. rtx5070 timeout) as
long as the host is identifiable from context; the full probe-failed: <target> <kind> form is required when a leg reports a failure in the detail section.
Probe Shape (per standing rules from 1150.msg)
- Any HTTP status means ALIVE. 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe.
- A failed probe is never a service verdict. Print
probe-failed: <target> <kind>naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report. - Say which probe produced each number. "Grafana: 000" is unusable; "Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
Strategies
When LiteLLM keys are invalid
Report the specific agent + key name. Do not attempt to fix — credential rotation is a separate operation.
When GPU port conflicts are detected
Report the conflicting ports and processes. Do not kill processes — that's a destructive action requiring captain approval.
When gateway liveness is degraded
Report the specific CT + gateway status. Do not restart unless the restart debounce window has passed.
When CT liveness is down
Report the specific CT. Do not restart — that's a destructive action.
When config YAML is invalid
Report the specific file + parse error. Do not fix — that's a config change.
When gateway log health is degraded
Report the specific gateway + log health status (error patterns, stale connections, connectivity issues). Do not restart — that's a destructive action.
Note: the script may perform additional diagnostics beyond the seven contract checks listed under Execution.
Continuity
- Every 4 hours (2, 6, 10, 14, 18, 22 UTC at :35): Scheduled cron check while Abiba is running
- On
agent-healthcommand: Run on-demand and report to user - On critical alert: Escalate to relay message immediately