Per defect report 1150.msg:
- Add standing probe rules section (2026-09-14)
- Step 1 (Zulip API): retry once at 25s on 000, print target + code
- Step 2 (Platform A): retry once at 25s on 000, print target + code, note must run on Abiba host
- Apply same shape as infrastructure-monitoring: any HTTP status = ALIVE; only 000/timeout = probe-failed
Verified: all probes now return real HTTP codes (Zulip API 200, Platform A 200, Tanko 401, Agent Zero 401)
Correct the dsh-web authentication fix after parent correction:
- Remove the unauthenticated :8081 endpoint (0.0.0.0 bind with no
auth_request = full Authentik bypass for the LAN). The script now removes
/etc/nginx/sites-enabled/dsh.token automatically if it reappears.
- Put the login path inside the Authentik-gated :80 server block as
location = /dsh-web-login; proxy to dsh-web with Host =
tankodhs.sysloggh.net so the 30-day cookie is bound to the public
authority, never to 127.0.0.1:3080.
- Isolate the rotating token in a generated include
/etc/dsh-web/nginx-login.conf; reload nginx only when it changes.
- Replace the disruptive capture (systemctl stop/start dsh-web) with a
non-disruptive read of the running service's journal, scoped to the
current systemd invocation so a restarted process's stale token is never
reused while the new banner is still pending.
- Keep x-dsh-task-board-proxy-token and Host $ak_origin_host intact in '/'.
- Document the corrected design (B4) in zulip-health.prose.md, v3.2.0.
Live-verified 2026-09-11: no auth bypass (302), :8081 refused (000), a
cookie minted before two dsh-web restarts still returns 200, the refreshed
token mints a fresh cookie, and the systemd ExecStartPost/timer refreshes
the token automatically without touching dsh-web.
- Add capture-dsh-token.sh script that captures the dsh-web launch token
- Add login endpoint (/dsh-web-login on :8081) that mints 30-day auth cookie
- Document the authentication flow in zulip-health.prose.md (Platform B4)
- Cookie is authority-bound to 127.0.0.1:3080 with 30-day expiry
- After first login, subsequent requests use the cookie — no token required
Captain ruling 2026-09-10: Mumuni moved off this host onto her own container
(kagentz CT 105 on minipve, 192.168.68.14, dedicated `hermes` user) and is
monitored from her side. This host must not monitor anything Mumuni.
The stale probes fired false alerts repeatedly:
* scripts/zulip-monitor.sh ssh'd to root@192.168.68.24 for the
decommissioned deployment's ~/.hermes/gateway_state.json, read "unknown"
on every run, and posted a 🔴 "Mumuni (Hermes) Zulip state: unknown" DM +
#agent-hub stream alert each cycle.
* scripts/daily-infra-report.py published a matching "mumuni:unknown" row in
every digest.
Changes:
* zulip-monitor.sh: delete the "Platform B: Hermes (Mumuni)" leg and its
notify; keep the Zulip-server, Platform A pi/Abiba (bridge), Platform B
Tanko and Platform C Agent Zero legs. A comment records why the leg is
retired so it is not re-added. Also fixes SC2155 so shellcheck is clean.
* daily-infra-report.py: delete the .24 ~/.hermes/gateway_state.json agent
probe, its hermes --version probe, and the now-dead mumuni render branch.
Abiba (CT 100, its own .24 address) and Tanko legs unchanged; the PVE API
token name is untouched.
* zulip-health.prose.md (v3.1.0): drop the Mumuni-only B4 gateway-process,
B5 heartbeat and B6 response-delivery steps and the stale 192.168.68.24
references; state explicitly that Mumuni is not monitored from this host.
Tanko/Agent-Zero/bridge steps retained.
* agent-health-check.py: correct the v2 changelog roster comment that still
placed mumuni at .24/CT100. No behavior change — the mumuni probe was
already absent from the AGENTS dict; v5 changelog notes the correction.
Tests: tests/test_mumuni_monitor_removal.py pins the removal structurally and
behaviorally — the shipped zulip-monitor.sh is run in a sandbox (only its LOG
constant rewritten) with stub ssh/curl on PATH; the ssh stub records every
target host, so "never reaches .24" and "no Mumuni notify even on the alert
path" are asserted from observed behavior. A mutation check (re-inject the old
leg) fails the suite, so the guarantee is not vacuous.
Post-merge verification of PR #54 found:
1. Lifecycle commands in zulip-health/zulip-self-heal used
'sudo systemctl' for hermes-gateway - sudo is NOT installed
on kagentz (no sudoers, no polkit user rules). Correct path:
root SSH invocation (root@192.168.68.14), matching how the Proxmox
host actually manages the unit.
2. hermes-zulip-restore.prose.md line 5 + zulip-resilience-v3 line 423
still said 'Mumuni CT100' in prose - repointed.
Found via independent review + live runtime check (whoami, which sudo,
journalctl, systemctl show). Refs PR #54, relay #738/#739.
Mumuni migrated off Abiba CT100 (192.168.68.24, purged) to dedicated
CT kagentz (192.168.68.14) on minipve. Verified live on kagentz:
- Gateway: systemd unit hermes-gateway.service (User=hermes), active
- Hermes home: /home/hermes/.hermes
- gateway_state.json present at ~/.hermes/gateway_state.json
Updates Mumuni rows/references in 9 contracts. Abiba-owned refs
(.24 pi agent, GPU monitor, infrastructure-control CRITICAL file)
and historical runs/ logs intentionally left unchanged.
Per relay #738 follow-ups. Refs #735-#738.