PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
- Implement four-state PBS GC logic in proxmox-monitor.sh: - probe-failed: unparseable JSON, store not found, or empty body → FAIL - running: collection in progress (last-run-endtime absent, upid present) → DO NOT FAIL - stale: no completed run within 48h → FAIL, naming last completed run age - healthy: completed within 48h → PASS, naming endtime and pending bytes - Add tests/test_pbs_gc_states.sh covering all four states - Proves the test bites on the pre-fix version (5/6 tests fail) - All 6 tests pass against the fixed version - Update contract-run.sh to map disk-gc-threat-response -> scripts/disk-gc-scan.py - Update Execution sections of host-scheduled contracts: infrastructure-monitoring, zulip-health, litellm-health, agent-health-check, disk-gc-threat-response, pm2-self-heal Adding note that execution is host-scheduled via cron, not agent session ack.
5.0 KiB
5.0 KiB
kind, name, description
| kind | name | description |
|---|---|---|
| responsibility | pm2-self-heal | Monitors critical PM2 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog) and auto-restarts any that are stopped or errored. Logs every action to Gitea health-logs and alerts the owner via Telegram (primary) or Zulip DM (secondary). Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure. |
Maintains
- abiba-telegram: { status: "online", uptime: string, restarts: number }
- abiba-zulip: { status: "online", uptime: string, restarts: number }
- gitea-runner: { status: "online", uptime: string, restarts: number }
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
- last_check: timestamp
Status (2026-08-03):
abiba-zulipfully restored — live Zulip bridge, heartbeating, monitored.gpu-monitorruns via systemd only (gpu-monitor.service); NOT PM2-tracked.gpu-watchdogretired from PM2 (folded into gpu-monitor.service).zulip-watchdogremains live and PM2-managed.
Continuity
- Self-driven: check every 300 seconds (5 min)
- Also wakes on user request (
prose run pm2-self-heal)
Remediation Rules
Rule 1: Process Stopped or Errored
- Detect:
pm2 statusshows "stopped" or "errored" for a named process - Fix:
pm2 restart <name> - Verify: Re-check status after 5 seconds
- Escalate: If still failed after 2 retries, send Zulip DM to owner
Rule 2: Process Restarting Too Often (crash-loop guard)
- Detect:
pm2 statusshows restarts > 30 (cumulative lifetime counter) or a process reporting a restart count > 1000 while showing "online" (a crash-loop mask) - Note: PM2 counter never decrements; only full delete+re-add resets it
- Fix:
pm2 delete <name> && pm2 start <ecosystem> --only <name> - Script guard (2026-08-16):
scripts/pm2-self-heal.shnow restartsabiba-telegramwhenTEL_RESTARTS > 1000even if the process reports "online" — catches a quiet crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts spoton incident). Alerts include the restart count. - AS-BUILT (2026-09-15): spoton-service was deleted with its app; the live PM2 set is four processes (abiba-telegram, abiba-zulip, gitea-runner, gpu-monitor). The spoton reference above is historical context for the crash-loop guard, not a live process.
- Escalate: Only when restarts > 30 — alerts to Zulip DM
- Historical fix: Previous cycles were caused by abiba-zulip extension's stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2). This was the root cause of the 18-restart accumulation. Threshold raised and PM2 counter reset on 2026-06-28.
(## Execution
)
Execution model: This contract is executed by a host-scheduled cron job (see
scripts/contract-run.sh). The cron job runs the monitoring script directly on the
target host and appends the result to /var/log/contract-runs/<contract>.log. An
agent-session acknowledgement (a done: line in the ops status log) is NOT
execution — it only proves the agent read the result and reported it. The actual
monitoring work happens in the host cron job.
- Check PM2 status — Run
pm2 status --no-colorand parse the table (5th data column = PID, 8th = restarts, 9th = status) - Check abiba-telegram (safe to auto-restart):
- If status is "online" → pass
- If status is "stopped" or "errored" → apply Rule 1
- If restarts > 1000 → apply Rule 2 (crash-loop guard)
- Check abiba-zulip (live Zulip bridge, heartbeating):
- If status is "online" → pass, log restarts count
- If status is "stopped" or "errored" → restart (
pm2 restart abiba-zulip— fully restored) - If restarts > 5 in last hour → alert owner with full diagnostics
- Check gitea-runner:
- If status is "online" → pass
- If status is "stopped" or "errored" → apply Rule 1
- If restarts > 5 → alert owner
- Check zulip-watchdog:
- If status is "online" → pass
- If status is "stopped" or "errored" → apply Rule 1
- If restarts > 5 → alert owner
- Log results — Append to
SyslogSolution/health-logs/pm2/{timestamp}.mdin Gitea (not knowledge graph — hard rule) - Alert — Send Telegram (primary) or Zulip DM (secondary) to owner if escalation needed (do NOT run pm2 commands during alerting)
- Wait 5 min → repeat from step 1
Example Output (when healthy)
{
"run_id": "pm2-heal-001",
"checked_at": "2026-06-26T15:00:00Z",
"processes": {
"abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 },
"abiba-telegram": { "status": "online", "uptime": "48h", "restarts": 0 }
},
"actions_taken": [],
"overall": "healthy"
}
Example Output (when fixed)
{
"processes": {
"abiba-telegram": { "status": "errored → restarted → online" },
"abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 }
},
"actions_taken": [
{ "target": "abiba-telegram", "action": "pm2 restart", "result": "online" }
],
"overall": "fixed"
}