PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- pm2-self-heal.prose.md: description said 'logs every action to the knowledge graph', contradicting its own execution step (line ~60) that says '(not knowledge graph - hard rule)'. Align description to Gitea. - zulip-self-heal.prose.md: Reporting section said 'every cycle produces a knowledge graph node'. Contract is RETIRED; remove graph-node directive. - Both now consistently log to SyslogSolution/health-logs, never the graph.
98 lines
3.8 KiB
Markdown
98 lines
3.8 KiB
Markdown
---
|
|
kind: responsibility
|
|
name: pm2-self-heal
|
|
description: >
|
|
Monitors critical PM2 processes (abiba-zulip, abiba-telegram, gitea-runner,
|
|
spoton-service, zulip-watchdog) and auto-restarts any that are stopped or
|
|
errored. Logs every action to Gitea (SyslogSolution/health-logs — not
|
|
knowledge graph, hard rule) and alerts the owner via
|
|
Zulip DM on failures.
|
|
CRITICAL: Never restart abiba-zulip — it runs this contract.
|
|
AS-BUILT 2026-08-09 (captain ruling, ecosystem is authoritative):
|
|
gpu-monitor is systemd-managed (gpu-monitor.service) — NOT PM2;
|
|
gpu-watchdog decommissioned (function folded into gpu-monitor.service);
|
|
gitea-runner KEPT (online in PM2); abiba-zulip KEPT (online 4d+, the
|
|
2026-07-04 'removed/decommissioned' note was stale and is removed).
|
|
---
|
|
|
|
## Maintains
|
|
|
|
- abiba-telegram: { status: "online", uptime: string, restarts: number }
|
|
- abiba-zulip: { status: "online", uptime: string, restarts: number }
|
|
- gitea-runner: { status: "online", uptime: string, restarts: number }
|
|
- spoton-service: { status: "online", uptime: string, restarts: number }
|
|
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
|
- last_check: timestamp
|
|
|
|
> **Note (2026-07-04, SUPERSEDED 2026-08-09):** `abiba-zulip` remains ONLINE and
|
|
> is monitored — the decommission note was stale (process re-added; do not treat
|
|
> it as removed).
|
|
|
|
## Continuity
|
|
|
|
- Self-driven: check every 300 seconds (5 min)
|
|
- Also wakes on user request (`prose run pm2-self-heal`)
|
|
|
|
## Remediation Rules
|
|
|
|
### Rule 1: Process Stopped or Errored
|
|
- **Detect**: `pm2 status` shows "stopped" or "errored" for a named process
|
|
- **Fix**: `pm2 restart <name>`
|
|
- **Verify**: Re-check status after 5 seconds
|
|
- **Escalate**: If still failed after 2 retries, send Zulip DM to owner
|
|
|
|
### Rule 2: Process Restarting Too Often
|
|
- **Detect**: `pm2 status` shows restarts > 30 (cumulative lifetime counter)
|
|
- **Note**: PM2 counter never decrements; only full delete+re-add resets it
|
|
- **Fix**: `pm2 delete <name> && pm2 start <ecosystem> --only <name>`
|
|
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
|
|
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
|
|
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
|
|
This was the root cause of the 18-restart accumulation. Threshold raised
|
|
and PM2 counter reset on 2026-06-28.
|
|
|
|
## Execution
|
|
|
|
1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status)
|
|
2. **Check abiba-telegram**:
|
|
- If status is "online" → pass
|
|
- If status is "stopped" or "errored" → apply Rule 1
|
|
- If restarts > 5 → alert owner
|
|
3. **Check abiba-zulip** (self-process, read-only):
|
|
- If status is "online" → pass, log restarts count
|
|
- If status is "stopped" or "errored" → **DO NOT RESTART** — alert owner immediately
|
|
- If restarts > 5 in last hour → alert owner with full diagnostics
|
|
4. **Log results** — Append to `SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea (not knowledge graph — hard rule)
|
|
5. **Alert** — Send Zulip DM to owner if escalation needed (do NOT run pm2 commands during alerting)
|
|
6. **Wait 5 min** → repeat from step 1
|
|
|
|
## Example Output (when healthy)
|
|
|
|
```json
|
|
{
|
|
"run_id": "pm2-heal-001",
|
|
"checked_at": "2026-06-26T15:00:00Z",
|
|
"processes": {
|
|
"abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 },
|
|
"abiba-telegram": { "status": "online", "uptime": "48h", "restarts": 0 }
|
|
},
|
|
"actions_taken": [],
|
|
"overall": "healthy"
|
|
}
|
|
```
|
|
|
|
## Example Output (when fixed)
|
|
|
|
```json
|
|
{
|
|
"processes": {
|
|
"abiba-telegram": { "status": "errored → restarted → online" },
|
|
"abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 }
|
|
},
|
|
"actions_taken": [
|
|
{ "target": "abiba-telegram", "action": "pm2 restart", "result": "online" }
|
|
],
|
|
"overall": "fixed"
|
|
}
|
|
```
|