From bc7a55122ff981ea6f64eb6ee4f8ea348561440d Mon Sep 17 00:00:00 2001 From: root Date: Tue, 15 Sep 2026 13:05:32 +0000 Subject: [PATCH] fix: pm2 contract corrections pm2-self-heal.prose.md: - Add AS-BUILT note: gpu-monitor is systemd-managed, NOT PM2 - gpu-watchdog is decommissioned and folded into gpu-monitor.service - gitea-runner is KEPT; abiba-zulip is KEPT (online for days) - spoton-service was deleted; live PM2 set is 4 processes - Preserve historical context for crash-loop guard litellm-health.prose.md: - Correct Prometheus node coverage: 6 nodes (.4/.5/.6/.9/.12/.15:9100) - Note .4:9100 is DEAD target (no route, down for weeks) - Clarify this does not read as 6 healthy nodes --- litellm-health.prose.md | 2 +- pm2-self-heal.prose.md | 7 +++++++ 2 files changed, 8 insertions(+), 1 deletion(-) diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 6449828..0c95c1b 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -19,7 +19,7 @@ description: > Scraped by Prometheus with Bearer master key; endpoint returns 307 → /metrics/. - Alertmanager (harness-alertmanager :9093) + zulip-bridge (:9102) deliver firing alerts to #agent-hub > alerts-infra via abiba-bot. Added 2026-08-09. - - Prometheus node job covers ALL 5 PVE nodes (.5/.6/.9/.12/.15:9100). + - Prometheus node job covers 6 PVE nodes (.4/.5/.6/.9/.12/.15:9100). Note: .4:9100 is a DEAD target (no route, down for weeks, not a live node). --- ## Architecture (v4.0.0 — Direct: nginx → LiteLLM → GPU) diff --git a/pm2-self-heal.prose.md b/pm2-self-heal.prose.md index e3df0e6..77abbd0 100644 --- a/pm2-self-heal.prose.md +++ b/pm2-self-heal.prose.md @@ -2,6 +2,10 @@ kind: responsibility name: pm2-self-heal description: > + PM2 process health check for abiba-telegram, abiba-zulip, gitea-runner, and zulip-watchdog. + gpu-monitor is systemd-managed (gpu-monitor.service), NOT PM2. + gpu-watchdog is decommissioned and folded into gpu-monitor.service. + gitea-runner is KEPT. abiba-zulip is KEPT (online for days). --- ## Maintains @@ -35,6 +39,9 @@ description: > when `TEL_RESTARTS > 1000` even if the process reports "online" — catches a quiet crash-loop that never toggles status to "stopped"/"errored" (e.g. the 10k-restarts spoton incident). Alerts include the restart count. +- **AS-BUILT (2026-09-15)**: spoton-service was deleted with its app; the live PM2 set is + four processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog). The spoton + reference above is historical context for the crash-loop guard, not a live process. - **Escalate**: Only when restarts > 30 — alerts to Zulip DM - **Historical fix**: Previous cycles were caused by abiba-zulip extension's stuck detection (STUCK_THRESHOLD_MS was 30min, raised to 4h in v2). -- 2.54.0