diff --git a/pm2-self-heal.prose.md b/pm2-self-heal.prose.md index 1215fb7..d828470 100644 --- a/pm2-self-heal.prose.md +++ b/pm2-self-heal.prose.md @@ -2,10 +2,13 @@ kind: responsibility name: pm2-self-heal description: > - Monitors critical PM2 processes (abiba-zulip, abiba-telegram) and + Monitors critical PM2 processes on CT100/.24 (the container this agent runs in) and auto-restarts any that are stopped or errored. Logs every action to the knowledge graph and alerts the owner via Zulip DM on failures. Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure. + **Execution host: CT100/.24 local pm2 — NOT CT116, NOT via SSH** (CT116 runs systemd+docker, no PM2). + **Process scope (5/5, 2026-08-09 ruling, ecosystem authoritative):** + abiba-zulip, abiba-telegram, gitea-runner, spoton-service, zulip-watchdog report_only_agents: - koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129 --- @@ -45,18 +48,23 @@ report_only_agents: ## Execution +**Execution host: CT100/.24 — run `pm2` commands LOCAL to this container. Never SSH to CT116** (CT116 runs systemd+docker, no PM2 set). + +**Process scope (5/5, 2026-08-09 ruling, ecosystem authoritative):** +1. abiba-zulip (live Zulip bridge, heartbeating) +2. abiba-telegram +3. gitea-runner +4. spoton-service +5. zulip-watchdog + 1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status) -2. **Check abiba-telegram**: - - If status is "online" → pass - - If status is "stopped" or "errored" → apply Rule 1 - - If restarts > 5 → alert owner -3. **Check abiba-zulip** (live Zulip bridge, heartbeating): +2. **Check all 5 processes** (abiba-zulip, abiba-telegram, gitea-runner, spoton-service, zulip-watchdog): - If status is "online" → pass, log restarts count - - If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored) - - If restarts > 5 in last hour → alert owner with full diagnostics -4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule) -5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting) -6. **Wait 5 min** → repeat from step 1 + - If status is "stopped" or "errored" → apply Rule 1 (restart) + - If restarts > 5 → alert owner +3. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule) +4. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting) +5. **Wait 5 min** → repeat from step 1 ## Example Output (when healthy)