fix: pm2 contract corrections
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
pm2-self-heal.prose.md: - Add AS-BUILT note: gpu-monitor is systemd-managed, NOT PM2 - gpu-watchdog is decommissioned and folded into gpu-monitor.service - gitea-runner is KEPT; abiba-zulip is KEPT (online for days) - spoton-service was deleted; live PM2 set is 4 processes - Preserve historical context for crash-loop guard litellm-health.prose.md: - Correct Prometheus node coverage: 6 nodes (.4/.5/.6/.9/.12/.15:9100) - Note .4:9100 is DEAD target (no route, down for weeks) - Clarify this does not read as 6 healthy nodes
This commit is contained in:
@@ -19,7 +19,7 @@ description: >
|
|||||||
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/.
|
||||||
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
- Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver
|
||||||
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09.
|
||||||
- Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100).
|
- Prometheus node job covers 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). Note: .4:9100 is a DEAD target (no route, down for weeks, not a live node).
|
||||||
---
|
---
|
||||||
|
|
||||||
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU)
|
||||||
|
|||||||
@@ -2,6 +2,10 @@
|
|||||||
kind: responsibility
|
kind: responsibility
|
||||||
name: pm2-self-heal
|
name: pm2-self-heal
|
||||||
description: >
|
description: >
|
||||||
|
PM2 process health check for abiba-telegram, abiba-zulip, gitea-runner, and zulip-watchdog.
|
||||||
|
gpu-monitor is systemd-managed (gpu-monitor.service), NOT PM2.
|
||||||
|
gpu-watchdog is decommissioned and folded into gpu-monitor.service.
|
||||||
|
gitea-runner is KEPT. abiba-zulip is KEPT (online for days).
|
||||||
---
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -35,6 +39,9 @@ description: >
|
|||||||
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
|
when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet
|
||||||
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
|
crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts
|
||||||
spoton incident). Alerts include the restart count.
|
spoton incident). Alerts include the restart count.
|
||||||
|
- **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is
|
||||||
|
four processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog). The spoton
|
||||||
|
reference above is historical context for the crash-loop guard, not a live process.
|
||||||
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
|
- **Escalate**: Only when restarts > 30 — alerts to Zulip DM
|
||||||
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
|
- **Historical fix**: Previous cycles were caused by abiba-zulip extension's
|
||||||
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
|
stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2).
|
||||||
|
|||||||
Reference in New Issue
Block a user