fix(monitoring): retire the Mumuni leg — stop probing her decommissioned .24 deployment
Captain ruling 2026-09-10: Mumuni moved off this host onto her own container (kagentz CT 105 on minipve, 192.168.68.14, dedicated `hermes` user) and is monitored from her side. This host must not monitor anything Mumuni. The stale probes fired false alerts repeatedly: * scripts/zulip-monitor.sh ssh'd to root@192.168.68.24 for the decommissioned deployment's ~/.hermes/gateway_state.json, read "unknown" on every run, and posted a 🔴 "Mumuni (Hermes) Zulip state: unknown" DM + #agent-hub stream alert each cycle. * scripts/daily-infra-report.py published a matching "mumuni:unknown" row in every digest. Changes: * zulip-monitor.sh: delete the "Platform B: Hermes (Mumuni)" leg and its notify; keep the Zulip-server, Platform A pi/Abiba (bridge), Platform B Tanko and Platform C Agent Zero legs. A comment records why the leg is retired so it is not re-added. Also fixes SC2155 so shellcheck is clean. * daily-infra-report.py: delete the .24 ~/.hermes/gateway_state.json agent probe, its hermes --version probe, and the now-dead mumuni render branch. Abiba (CT 100, its own .24 address) and Tanko legs unchanged; the PVE API token name is untouched. * zulip-health.prose.md (v3.1.0): drop the Mumuni-only B4 gateway-process, B5 heartbeat and B6 response-delivery steps and the stale 192.168.68.24 references; state explicitly that Mumuni is not monitored from this host. Tanko/Agent-Zero/bridge steps retained. * agent-health-check.py: correct the v2 changelog roster comment that still placed mumuni at .24/CT100. No behavior change — the mumuni probe was already absent from the AGENTS dict; v5 changelog notes the correction. Tests: tests/test_mumuni_monitor_removal.py pins the removal structurally and behaviorally — the shipped zulip-monitor.sh is run in a sandbox (only its LOG constant rewritten) with stub ssh/curl on PATH; the ssh stub records every target host, so "never reaches .24" and "no Mumuni notify even on the alert path" are asserted from observed behavior. A mutation check (re-inject the old leg) fails the suite, so the guarantee is not vacuous.
This commit is contained in:
@@ -17,8 +17,9 @@ Changelog:
|
||||
v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity,
|
||||
vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key
|
||||
name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}).
|
||||
Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114),
|
||||
abiba (.24).
|
||||
Fleet roster: tanko (.122), koby (.129), koonimo (.114), abiba (.24).
|
||||
(v2 also carried a mumuni probe; see v5 — mumuni is no longer probed: she
|
||||
moved to her own container, kagentz CT 105 / .14, and is monitored there.)
|
||||
v3 (2026-09-08): GPU unit repoint verified live (.8 llama-chat-api.service,
|
||||
.110 llama-server.service, .15 strix-server.service) — .8 was probing a stale
|
||||
llama-server unit that reads inactive, producing false UNREACHABLE legs.
|
||||
@@ -44,6 +45,11 @@ Changelog:
|
||||
separate from `failures`. Every run prints absolute execution provenance
|
||||
(script + cwd) in the header, in the cron ALERT line, and in --json output so
|
||||
a stale-consumer report is distinguishable from a fault at read time.
|
||||
v5 (2026-09-10): roster correction only, no behavior change. mumuni was removed
|
||||
from the AGENTS dict when she moved off this host onto her own container
|
||||
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
|
||||
from her side. This script must not probe mumuni or .24 — the v2 changelog
|
||||
roster line was the last reference still placing her at .24 / CT100.
|
||||
"""
|
||||
|
||||
import subprocess, json, sys, os, time, re, io, contextlib
|
||||
|
||||
Reference in New Issue
Block a user