--- kind: responsibility name: pm2-self-heal description: > Monitors critical PM2 processes on CT100/.24 (the container this agent runs in) and auto-restarts any that are stopped or errored. Logs every action to the knowledge graph and alerts the owner via Zulip DM on failures. Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure. **Execution host: CT100/.24 local pm2 — NOT CT116, NOT via SSH** (CT116 runs systemd+docker, no PM2). **Process scope (5/5, 2026-08-09 ruling, ecosystem authoritative):** abiba-zulip, abiba-telegram, gitea-runner, spoton-service, zulip-watchdog report_only_agents: - koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129 --- ## Maintains - abiba-telegram: { status: "online", uptime: string, restarts: number } - abiba-zulip: { status: "online", uptime: string, restarts: number } - gpu-monitor: { status: "online", uptime: string, restarts: number } (systemd-managed, PM2-tracked) - gitea-runner: { status: "online", uptime: string, restarts: number } - last_check: timestamp > **Status (2026-08-03):** `abiba-zulip` fully restored — live Zulip bridge, heartbeating, monitored. `gpu-monitor` runs via systemd (stable); PM2 tracks it for status reporting only. `gpu-watchdog` retired from PM2. ## Continuity - Self-driven: check every 300 seconds (5 min) - Also wakes on user request (`prose run pm2-self-heal`) ## Remediation Rules ### Rule 1: Process Stopped or Errored - **Detect**: `pm2 status` shows "stopped" or "errored" for a named process - **Fix**: `pm2 restart ` - **Verify**: Re-check status after 5 seconds - **Escalate**: If still failed after 2 retries, send Zulip DM to owner ### Rule 2: Process Restarting Too Often - **Detect**: `pm2 status` shows restarts > 30 (cumulative lifetime counter) - **Note**: PM2 counter never decrements; only full delete+re-add resets it - **Fix**: `pm2 delete && pm2 start --only ` - **Escalate**: Only when restarts > 30 — alerts to Zulip DM - **Historical fix**: Previous cycles were caused by abiba-zulip extension's stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2). This was the root cause of the 18-restart accumulation. Threshold raised and PM2 counter reset on 2026-06-28. ## Execution **Execution host: CT100/.24 — run `pm2` commands LOCAL to this container. Never SSH to CT116** (CT116 runs systemd+docker, no PM2 set). **Process scope (5/5, 2026-08-09 ruling, ecosystem authoritative):** 1. abiba-zulip (live Zulip bridge, heartbeating) 2. abiba-telegram 3. gitea-runner 4. spoton-service 5. zulip-watchdog 1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status) 2. **Check all 5 processes** (abiba-zulip, abiba-telegram, gitea-runner, spoton-service, zulip-watchdog): - If status is "online" → pass, log restarts count - If status is "stopped" or "errored" → apply Rule 1 (restart) - If restarts > 5 → alert owner 3. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule) 4. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting) 5. **Wait 5 min** → repeat from step 1 ## Example Output (when healthy) ```json { "run_id": "pm2-heal-001", "checked_at": "2026-06-26T15:00:00Z", "processes": { "abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 }, "abiba-telegram": { "status": "online", "uptime": "48h", "restarts": 0 } }, "actions_taken": [], "overall": "healthy" } ``` ## Example Output (when fixed) ```json { "processes": { "abiba-telegram": { "status": "errored → restarted → online" }, "abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 } }, "actions_taken": [ { "target": "abiba-telegram", "action": "pm2 restart", "result": "online" } ], "overall": "fixed" } ```