--- kind: responsibility name: pm2-self-heal description: > Monitors critical PM2 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog) and auto-restarts any that are stopped or errored. Logs every action to Gitea health-logs and alerts the owner via Telegram (primary) or Zulip DM (secondary). Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure. --- ## Maintains - abiba-telegram: { status: "online", uptime: string, restarts: number } - abiba-zulip: { status: "online", uptime: string, restarts: number } - gitea-runner: { status: "online", uptime: string, restarts: number } - zulip-watchdog: { status: "online", uptime: string, restarts: number } - last_check: timestamp > **Status (2026-08-03):** `abiba-zulip` fully restored — live Zulip bridge, heartbeating, monitored. `gpu-monitor` runs via systemd only (gpu-monitor.service); NOT PM2-tracked. `gpu-watchdog` retired from PM2 (folded into gpu-monitor.service). `zulip-watchdog` remains live and PM2-managed. ## Continuity - Self-driven: check every 300 seconds (5 min) - Also wakes on user request (`prose run pm2-self-heal`) ## Remediation Rules ### Rule 1: Process Stopped or Errored - **Detect**: `pm2 status` shows "stopped" or "errored" for a named process - **Fix**: `pm2 restart ` - **Verify**: Re-check status after 5 seconds - **Escalate**: If still failed after 2 retries, send Zulip DM to owner ### Rule 2: Process Restarting Too Often (crash-loop guard) - **Detect**: `pm2 status` shows restarts > 30 (cumulative lifetime counter) **or** a process reporting a restart count > 1000 while showing "online" (a crash-loop mask) - **Note**: PM2 counter never decrements; only full delete+re-add resets it - **Fix**: `pm2 delete && pm2 start --only ` - **Script guard (2026-08-16)**: `scripts/pm2-self-heal.sh` now restarts `abiba-telegram` when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts spoton incident). Alerts include the restart count. - **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is four processes (abiba-telegram, abiba-zulip, gitea-runner, gpu-monitor). The spoton reference above is historical context for the crash-loop guard, not a live process. - **Escalate**: Only when restarts > 30 — alerts to Zulip DM - **Historical fix**: Previous cycles were caused by abiba-zulip extension's stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2). This was the root cause of the 18-restart accumulation. Threshold raised and PM2 counter reset on 2026-06-28. ## Execution **Execution model**: This contract is executed by a host-scheduled cron job (see `scripts/contract-run.sh`). The cron job runs the monitoring script directly on the target host and appends the result to `/var/log/contract-runs/.log`. An agent-session acknowledgement (a `done:` line in the ops status log) is NOT execution — it only proves the agent read the result and reported it. The actual monitoring work happens in the host cron job. 1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status) 2. **Check abiba-telegram** (safe to auto-restart): - If status is "online" → pass - If status is "stopped" or "errored" → apply Rule 1 - If restarts > 1000 → apply Rule 2 (crash-loop guard) 3. **Check abiba-zulip** (live Zulip bridge, heartbeating): - If status is "online" → pass, log restarts count - If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored) - If restarts > 5 in last hour → alert owner with full diagnostics 4. **Check gitea-runner**: - If status is "online" → pass - If status is "stopped" or "errored" → apply Rule 1 - If restarts > 5 → alert owner 5. **Check zulip-watchdog**: - If status is "online" → pass - If status is "stopped" or "errored" → apply Rule 1 - If restarts > 5 → alert owner 6. **Log results** — the durable per-run record is `/var/log/contract-runs/pm2-self-heal-.log` on CT 100, written by `scripts/contract-run.sh` from `/etc/cron.d/contract-runner` every 4 hours, with a firstmate inbox note raised on any non-zero exit. **CORRECTED 2026-09-26 (relay-785):** this step previously required appending to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea and called it a "hard rule". That posting was **never implemented** — `scripts/pm2-self-heal.sh` contains no Gitea or git-push code — so the requirement was a coverage claim the executor did not honour, and `health-logs/pm2/` has held only its init commit since 2026-07-28. The claim is retired rather than implemented: the contract-runner's per-run logs plus its failure note already give a durable record and a working alarm, and a second posting path would add work without adding a signal. `health-logs/pm2/` is left as historical evidence, not as a live obligation. 7. **Alert** — Send Telegram (primary) or Zulip DM (secondary) to owner if escalation needed (do NOT run pm2 commands during alerting) 8. **Wait 5 min** → repeat from step 1 ## Example Output (when healthy) ```json { "run_id": "pm2-heal-001", "checked_at": "2026-06-26T15:00:00Z", "processes": { "abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 }, "abiba-telegram": { "status": "online", "uptime": "48h", "restarts": 0 } }, "actions_taken": [], "overall": "healthy" } ``` ## Example Output (when fixed) ```json { "processes": { "abiba-telegram": { "status": "errored → restarted → online" }, "abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 } }, "actions_taken": [ { "target": "abiba-telegram", "action": "pm2 restart", "result": "online" } ], "overall": "fixed" } ```