PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
relay-785 / self-heal-contract-stale-script-pointers-20260924. gpu/ had been silent since 2026-09-14T18:02:03Z (12 days) and pm2/ had never posted a run. FINDING - the gpu-self-heal executor was never lost, only its schedule was. The brief concluded the mechanism was gone. It is not: on CT 116 /opt/inference-harness/scripts/gpu-self-heal.py exists (mtime 2026-09-21 11:43), is self-posting via gitea-logger.sh, and runs cleanly - I ran it and it posted health-logs/gpu/gpu-self-heal-20260926-144618.json. What was missing was the cron entry, dropped during a CT 116 /etc/cron.d rework on 2026-09-21. So this is a schedule restoration, not a resurrection. DECISIONS * gpu-self-heal -> RESTORE. Schedule re-added as /etc/cron.d/gpu-self-heal on CT 116, matching the observed historical cadence (:02 past 0/6/12/18). Proven with a REAL unattended cron fire (temporary */1 entry, removed after): gpu-self-heal-20260926-144802.json committed 14:48:03Z. The contract now names the executor, schedule, log and posting, which it previously did not. * pm2 health-logs posting -> RETIRE the claim. pm2-self-heal.prose.md required appending to health-logs/pm2/ as a 'hard rule', but scripts/pm2-self-heal.sh contains no Gitea or push code and never did; the directory has held only its init commit since 2026-07-28. Corrected to point at the contract-runner's durable per-run logs and failure note instead of adding a second, redundant posting path. DEAD-MAN'S-SWITCH - scripts/health-log-freshness.py + contract. Absence of logs must raise an alarm, and the producer cannot raise it, so this runs on CT 100, a different host from the producers, and fails when the newest health-logs/gpu entry is older than 12h (litellm 18h). Verified it would have caught the real gap: evaluated at 2026-09-20 the newest gpu/ entry was 126h old against a 12h limit -> STALE. Bugs found and fixed while testing, each caught by a test that bit: * the documented HEALTH_LOG_MAX_AGE_* override was never implemented; * a Gitea password held under GITEA_TOKEN was sent as an API token -> HTTP 401; now every candidate auth is tried and the first that works is used; * the ~/.git-credentials fallback filtered on GITEA_URL's host, which missed whenever GITEA_URL pointed at the internal IP -> zero candidates. Exit codes: 0 fresh, 1 stale-or-unreadable, 2 cannot run. A directory that cannot be read is a failure, never a skip. Evidence: live PASS; stale override names the directory and producer; no credential -> exit 2; whole thing runs green through contract-run.sh. prose-lint: PASSED.
122 lines
5.8 KiB
Markdown
122 lines
5.8 KiB
Markdown
---
|
|
kind: responsibility
|
|
name: pm2-self-heal
|
|
description: >
|
|
Monitors critical PM2 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog)
|
|
and auto-restarts any that are stopped or errored. Logs every action to
|
|
Gitea health-logs and alerts the owner via Telegram (primary) or Zulip DM (secondary).
|
|
Abiba-zulip is the live Zulip bridge and may be restarted; alert owner on failure.
|
|
---
|
|
|
|
## Maintains
|
|
|
|
- abiba-telegram: { status: "online", uptime: string, restarts: number }
|
|
- abiba-zulip: { status: "online", uptime: string, restarts: number }
|
|
- gitea-runner: { status: "online", uptime: string, restarts: number }
|
|
- zulip-watchdog: { status: "online", uptime: string, restarts: number }
|
|
- last_check: timestamp
|
|
|
|
> **Status (2026-08-03):** `abiba-zulip` fully restored — live Zulip bridge, heartbeating, monitored. `gpu-monitor` runs via systemd only (gpu-monitor.service); NOT PM2-tracked. `gpu-watchdog` retired from PM2 (folded into gpu-monitor.service). `zulip-watchdog` remains live and PM2-managed.
|
|
|
|
|
|
## Continuity
|
|
|
|
- Self-driven: check every 300 seconds (5 min)
|
|
- Also wakes on user request (`prose run pm2-self-heal`)
|
|
|
|
## Remediation Rules
|
|
|
|
### Rule 1: Process Stopped or Errored
|
|
- **Detect**: `pm2 status` shows "stopped" or "errored" for a named process
|
|
- **Fix**: `pm2 restart <name>`
|
|
- **Verify**: Re-check status after 5 seconds
|
|
- **Escalate**: If still failed after 2 retries, send Zulip DM to owner
|
|
|
|
### Rule 2: Process Restarting Too Often (crash-loop guard)
|
|
- **Detect**: `pm2 status` shows restarts > 30 (cumulative lifetime counter) **or** a
|
|
process reporting a restart count > 1000 while showing "online" (a crash-loop mask)
|
|
- **Note**: PM2 counter never decrements; only full delete+re-add resets it
|
|
- **Fix**: `pm2 delete <name> && pm2 start <ecosystem> --only <name>`
|
|
- **Script guard (2026-08-16)**: `scripts/pm2-self-heal.sh` now restarts `abiba-telegram`
|
|
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
|
|
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
|
|
spoton incident). Alerts include the restart count.
|
|
- **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is
|
|
four processes (abiba-telegram, abiba-zulip, gitea-runner, gpu-monitor). The spoton
|
|
reference above is historical context for the crash-loop guard, not a live process.
|
|
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
|
|
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
|
|
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
|
|
This was the root cause of the 18-restart accumulation. Threshold raised
|
|
and PM2 counter reset on 2026-06-28.
|
|
|
|
## Execution
|
|
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
|
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
|
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
|
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
|
execution — it only proves the agent read the result and reported it. The actual
|
|
monitoring work happens in the host cron job.
|
|
|
|
|
|
1. **Check PM2 status** — Run `pm2 status --no-color` and parse the table (5th data column = PID, 8th = restarts, 9th = status)
|
|
2. **Check abiba-telegram** (safe to auto-restart):
|
|
- If status is "online" → pass
|
|
- If status is "stopped" or "errored" → apply Rule 1
|
|
- If restarts > 1000 → apply Rule 2 (crash-loop guard)
|
|
3. **Check abiba-zulip** (live Zulip bridge, heartbeating):
|
|
- If status is "online" → pass, log restarts count
|
|
- If status is "stopped" or "errored" → restart (`pm2 restart abiba-zulip` — fully restored)
|
|
- If restarts > 5 in last hour → alert owner with full diagnostics
|
|
4. **Check gitea-runner**:
|
|
- If status is "online" → pass
|
|
- If status is "stopped" or "errored" → apply Rule 1
|
|
- If restarts > 5 → alert owner
|
|
5. **Check zulip-watchdog**:
|
|
- If status is "online" → pass
|
|
- If status is "stopped" or "errored" → apply Rule 1
|
|
- If restarts > 5 → alert owner
|
|
6. **Log results** — the durable per-run record is `/var/log/contract-runs/pm2-self-heal-<UTCstamp>.log` on CT 100, written by `scripts/contract-run.sh` from `/etc/cron.d/contract-runner` every 4 hours, with a firstmate inbox note raised on any non-zero exit.
|
|
|
|
**CORRECTED 2026-09-26 (relay-785):** this step previously required appending to
|
|
`SyslogSolution/health-logs/pm2/{timestamp}.md` in Gitea and called it a "hard rule".
|
|
That posting was **never implemented** — `scripts/pm2-self-heal.sh` contains no
|
|
Gitea or git-push code — so the requirement was a coverage claim the executor did
|
|
not honour, and `health-logs/pm2/` has held only its init commit since 2026-07-28.
|
|
The claim is retired rather than implemented: the contract-runner's per-run logs
|
|
plus its failure note already give a durable record and a working alarm, and a
|
|
second posting path would add work without adding a signal. `health-logs/pm2/`
|
|
is left as historical evidence, not as a live obligation.
|
|
7. **Alert** — Send Telegram (primary) or Zulip DM (secondary) to owner if escalation needed (do NOT run pm2 commands during alerting)
|
|
8. **Wait 5 min** → repeat from step 1
|
|
|
|
## Example Output (when healthy)
|
|
|
|
```json
|
|
{
|
|
"run_id": "pm2-heal-001",
|
|
"checked_at": "2026-06-26T15:00:00Z",
|
|
"processes": {
|
|
"abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 },
|
|
"abiba-telegram": { "status": "online", "uptime": "48h", "restarts": 0 }
|
|
},
|
|
"actions_taken": [],
|
|
"overall": "healthy"
|
|
}
|
|
```
|
|
|
|
## Example Output (when fixed)
|
|
|
|
```json
|
|
{
|
|
"processes": {
|
|
"abiba-telegram": { "status": "errored → restarted → online" },
|
|
"abiba-zulip": { "status": "online", "uptime": "2h", "restarts": 0 }
|
|
},
|
|
"actions_taken": [
|
|
{ "target": "abiba-telegram", "action": "pm2 restart", "result": "online" }
|
|
],
|
|
"overall": "fixed"
|
|
}
|
|
```
|