The PM2 set lives on CT100/.24 (local, not CT116). CT116 runs systemd+docker, no PM2. Process scope (5/5, 2026-08-09 ruling, ecosystem authoritative): abiba-zulip, abiba-telegram, gitea-runner, spoton-service, zulip-watchdog This prevents future dispatches from drifting to the wrong host.
3.9 KiB
kind, name, description, report_only_agents
| kind | name | description | report_only_agents | |
|---|---|---|---|---|
| responsibility | pm2-self-heal | Monitors critical PM2 processes on CT100/.24 (the container this agent runs in) and auto-restarts any that are stopped or errored. Logs every action to the knowledge graph and alerts the owner via Zulip DM on failures. Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure. **Execution host: CT100/.24 local pm2 — NOT CT116, NOT via SSH** (CT116 runs systemd+docker, no PM2). **Process scope (5/5, 2026-08-09 ruling, ecosystem authoritative):** abiba-zulip, abiba-telegram, gitea-runner, spoton-service, zulip-watchdog |
|
Maintains
- abiba-telegram: { status: "online", uptime: string, restarts: number }
- abiba-zulip: { status: "online", uptime: string, restarts: number }
- gpu-monitor: { status: "online", uptime: string, restarts: number } (systemd-managed, PM2-tracked)
- gitea-runner: { status: "online", uptime: string, restarts: number }
- last_check: timestamp
Status (2026-08-03):
abiba-zulipfully restored — live Zulip bridge, heartbeating, monitored.gpu-monitorruns via systemd (stable); PM2 tracks it for status reporting only.gpu-watchdogretired from PM2.
Continuity
- Self-driven: check every 300 seconds (5 min)
- Also wakes on user request (
prose run pm2-self-heal)
Remediation Rules
Rule 1: Process Stopped or Errored
- Detect:
pm2 statusshows "stopped" or "errored" for a named process - Fix:
pm2 restart <name> - Verify: Re-check status after 5 seconds
- Escalate: If still failed after 2 retries, send Zulip DM to owner
Rule 2: Process Restarting Too Often
- Detect:
pm2 statusshows restarts > 30 (cumulative lifetime counter) - Note: PM2 counter never decrements; only full delete+re-add resets it
- Fix:
pm2 delete <name> && pm2 start <ecosystem> --only <name> - Escalate: Only when restarts > 30 — alerts to Zulip DM
- Historical fix: Previous cycles were caused by abiba-zulip extension's stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2). This was the root cause of the 18-restart accumulation. Threshold raised and PM2 counter reset on 2026-06-28.
Execution
Execution host: CT100/.24 — run pm2 commands LOCAL to this container. Never SSH to CT116 (CT116 runs systemd+docker, no PM2 set).
Process scope (5/5, 2026-08-09 ruling, ecosystem authoritative):
-
abiba-zulip (live Zulip bridge, heartbeating)
-
abiba-telegram
-
gitea-runner
-
spoton-service
-
zulip-watchdog
-
Check PM2 status — Run
pm2 status --no-colorand parse the table (5th data column = PID, 8th = restarts, 9th = status) -
Check all 5 processes (abiba-zulip, abiba-telegram, gitea-runner, spoton-service, zulip-watchdog):
- If status is "online" → pass, log restarts count
- If status is "stopped" or "errored" → apply Rule 1 (restart)
- If restarts > 5 → alert owner
-
Log results — Append to
SyslogSolution/health-logs/pm2/{timestamp}.mdin Gitea (not knowledge graph — hard rule) -
Alert — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
-
Wait 5 min → repeat from step 1
Example Output (when healthy)
{
"run_id": "pm2-heal-001",
"checked_at": "2026-06-26T15:00:00Z",
"processes": {
"abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 },
"abiba-telegram": { "status": "online", "uptime": "48h", "restarts": 0 }
},
"actions_taken": [],
"overall": "healthy"
}
Example Output (when fixed)
{
"processes": {
"abiba-telegram": { "status": "errored → restarted → online" },
"abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 }
},
"actions_taken": [
{ "target": "abiba-telegram", "action": "pm2 restart", "result": "online" }
],
"overall": "fixed"
}