Files
prose-contracts/pm2-self-heal.prose.md
root 9e87927444
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
pm2-self-heal: reconcile contract with live PM2 set and alert channels
- Add zulip-watchdog to Maintains (it's running, infrastructure-monitoring expects it)
- Remove gpu-monitor from PM2 Maintains (it's systemd-only, not PM2-tracked)
- Add Execution steps for all monitored processes (gitea-runner, zulip-watchdog)
- Update alert channel: Telegram is primary, Zulip DM is secondary
- Update script to check all 4 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog)
- Add restart count thresholds for all processes
- Update log output to include all process statuses
2026-09-22 11:25:45 +00:00

4.5 KiB

kind, name, description
kind name description
responsibility pm2-self-heal Monitors critical PM2 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog) and auto-restarts any that are stopped or errored. Logs every action to Gitea health-logs and alerts the owner via Telegram (primary) or Zulip DM (secondary). Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure.

Maintains

  • abiba-telegram: { status: "online", uptime: string, restarts: number }
  • abiba-zulip: { status: "online", uptime: string, restarts: number }
  • gitea-runner: { status: "online", uptime: string, restarts: number }
  • zulip-watchdog: { status: "online", uptime: string, restarts: number }
  • last_check: timestamp

Status (2026-08-03): abiba-zulip fully restored — live Zulip bridge, heartbeating, monitored. gpu-monitor runs via systemd only (gpu-monitor.service); NOT PM2-tracked. gpu-watchdog retired from PM2 (folded into gpu-monitor.service). zulip-watchdog remains live and PM2-managed.

Continuity

  • Self-driven: check every 300 seconds (5 min)
  • Also wakes on user request (prose run pm2-self-heal)

Remediation Rules

Rule 1: Process Stopped or Errored

  • Detect: pm2 status shows "stopped" or "errored" for a named process
  • Fix: pm2 restart <name>
  • Verify: Re-check status after 5 seconds
  • Escalate: If still failed after 2 retries, send Zulip DM to owner

Rule 2: Process Restarting Too Often (crash-loop guard)

  • Detect: pm2 status shows restarts > 30 (cumulative lifetime counter) or a process reporting a restart count > 1000 while showing "online" (a crash-loop mask)
  • Note: PM2 counter never decrements; only full delete+re-add resets it
  • Fix: pm2 delete <name> && pm2 start <ecosystem> --only <name>
  • Script guard (2026-08-16): scripts/pm2-self-heal.sh now restarts abiba-telegram when TEL_RESTARTS > 1000 even if the process reports "online" — catches a quiet crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts spoton incident). Alerts include the restart count.
  • AS-BUILT (2026-09-15): spoton-service was deleted with its app; the live PM2 set is four processes (abiba-telegram, abiba-zulip, gitea-runner, gpu-monitor). The spoton reference above is historical context for the crash-loop guard, not a live process.
  • Escalate: Only when restarts > 30 — alerts to Zulip DM
  • Historical fix: Previous cycles were caused by abiba-zulip extension's stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2). This was the root cause of the 18-restart accumulation. Threshold raised and PM2 counter reset on 2026-06-28.

Execution

  1. Check PM2 status — Run pm2 status --no-color and parse the table (5th data column = PID, 8th = restarts, 9th = status)
  2. Check abiba-telegram (safe to auto-restart):
    • If status is "online" → pass
    • If status is "stopped" or "errored" → apply Rule 1
    • If restarts > 1000 → apply Rule 2 (crash-loop guard)
  3. Check abiba-zulip (live Zulip bridge, heartbeating):
    • If status is "online" → pass, log restarts count
    • If status is "stopped" or "errored" → restart (pm2 restart abiba-zulip — fully restored)
    • If restarts > 5 in last hour → alert owner with full diagnostics
  4. Check gitea-runner:
    • If status is "online" → pass
    • If status is "stopped" or "errored" → apply Rule 1
    • If restarts > 5 → alert owner
  5. Check zulip-watchdog:
    • If status is "online" → pass
    • If status is "stopped" or "errored" → apply Rule 1
    • If restarts > 5 → alert owner
  6. Log results — Append to SyslogSolution/health-logs/pm2/{timestamp}.md in Gitea (not knowledge graph — hard rule)
  7. Alert — Send Telegram (primary) or Zulip DM (secondary) to owner if escalation needed (do NOT run pm2 commands during alerting)
  8. Wait 5 min → repeat from step 1

Example Output (when healthy)

{
  "run_id": "pm2-heal-001",
  "checked_at": "2026-06-26T15:00:00Z",
  "processes": {
    "abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 },
    "abiba-telegram": { "status": "online", "uptime": "48h", "restarts": 0 }
  },
  "actions_taken": [],
  "overall": "healthy"
}

Example Output (when fixed)

{
  "processes": {
    "abiba-telegram": { "status": "errored → restarted → online" },
    "abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 }
  },
  "actions_taken": [
    { "target": "abiba-telegram", "action": "pm2 restart", "result": "online" }
  ],
  "overall": "fixed"
}