fix(monitoring): retire the Mumuni leg — stop probing her decommissioned .24 deployment

Captain ruling 2026-09-10: Mumuni moved off this host onto her own container
(kagentz CT 105 on minipve, 192.168.68.14, dedicated `hermes` user) and is
monitored from her side. This host must not monitor anything Mumuni.

The stale probes fired false alerts repeatedly:
  * scripts/zulip-monitor.sh ssh'd to root@192.168.68.24 for the
    decommissioned deployment's ~/.hermes/gateway_state.json, read "unknown"
    on every run, and posted a 🔴 "Mumuni (Hermes) Zulip state: unknown" DM +
    #agent-hub stream alert each cycle.
  * scripts/daily-infra-report.py published a matching "mumuni:unknown" row in
    every digest.

Changes:
  * zulip-monitor.sh: delete the "Platform B: Hermes (Mumuni)" leg and its
    notify; keep the Zulip-server, Platform A pi/Abiba (bridge), Platform B
    Tanko and Platform C Agent Zero legs. A comment records why the leg is
    retired so it is not re-added. Also fixes SC2155 so shellcheck is clean.
  * daily-infra-report.py: delete the .24 ~/.hermes/gateway_state.json agent
    probe, its hermes --version probe, and the now-dead mumuni render branch.
    Abiba (CT 100, its own .24 address) and Tanko legs unchanged; the PVE API
    token name is untouched.
  * zulip-health.prose.md (v3.1.0): drop the Mumuni-only B4 gateway-process,
    B5 heartbeat and B6 response-delivery steps and the stale 192.168.68.24
    references; state explicitly that Mumuni is not monitored from this host.
    Tanko/Agent-Zero/bridge steps retained.
  * agent-health-check.py: correct the v2 changelog roster comment that still
    placed mumuni at .24/CT100. No behavior change — the mumuni probe was
    already absent from the AGENTS dict; v5 changelog notes the correction.

Tests: tests/test_mumuni_monitor_removal.py pins the removal structurally and
behaviorally — the shipped zulip-monitor.sh is run in a sandbox (only its LOG
constant rewritten) with stub ssh/curl on PATH; the ssh stub records every
target host, so "never reaches .24" and "no Mumuni notify even on the alert
path" are asserted from observed behavior. A mutation check (re-inject the old
leg) fails the suite, so the guarantee is not vacuous.
This commit is contained in:
abiba
2026-09-10 10:11:19 +00:00
parent b1462f3e79
commit 952fca9c92
5 changed files with 338 additions and 90 deletions
+8 -2
View File
@@ -17,8 +17,9 @@ Changelog:
v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity, v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity,
vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key
name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}). name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}).
Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), Fleet roster: tanko (.122), koby (.129), koonimo (.114), abiba (.24).
abiba (.24). (v2 also carried a mumuni probe; see v5 — mumuni is no longer probed: she
moved to her own container, kagentz CT 105 / .14, and is monitored there.)
v3 (2026-09-08): GPU unit repoint verified live (.8 llama-chat-api.service, v3 (2026-09-08): GPU unit repoint verified live (.8 llama-chat-api.service,
.110 llama-server.service, .15 strix-server.service) — .8 was probing a stale .110 llama-server.service, .15 strix-server.service) — .8 was probing a stale
llama-server unit that reads inactive, producing false UNREACHABLE legs. llama-server unit that reads inactive, producing false UNREACHABLE legs.
@@ -44,6 +45,11 @@ Changelog:
separate from `failures`. Every run prints absolute execution provenance separate from `failures`. Every run prints absolute execution provenance
(script + cwd) in the header, in the cron ALERT line, and in --json output so (script + cwd) in the header, in the cron ALERT line, and in --json output so
a stale-consumer report is distinguishable from a fault at read time. a stale-consumer report is distinguishable from a fault at read time.
v5 (2026-09-10): roster correction only, no behavior change. mumuni was removed
from the AGENTS dict when she moved off this host onto her own container
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
from her side. This script must not probe mumuni or .24 — the v2 changelog
roster line was the last reference still placing her at .24 / CT100.
""" """
import subprocess, json, sys, os, time, re, io, contextlib import subprocess, json, sys, os, time, re, io, contextlib
+7 -26
View File
@@ -308,26 +308,13 @@ def collect():
"updated_at": "", "updated_at": "",
} }
# Mumuni (CT 100, IP 192.168.68.24) # Mumuni is deliberately absent from this digest: captain ruling 2026-09-10.
mumuni_state = ssh("192.168.68.24", "cat ~/.hermes/gateway_state.json 2>/dev/null") # She moved off this host onto her own container (kagentz CT 105 on minipve,
mumuni_data = {} # 192.168.68.14, dedicated `hermes` user) and is monitored from her side. The
try: # former probe ssh'd to 192.168.68.24 for the decommissioned deployment's
mumuni_data = json.loads(mumuni_state) if mumuni_state else {} # ~/.hermes/gateway_state.json, always read "unknown", and published a false
except: # "mumuni:unknown" line in the agent table and the gateway-unknown issue
mumuni_data = {} # count of every digest. Do NOT re-add an .24 / gateway_state probe.
mumuni_platforms = mumuni_data.get("platforms", {})
report["agents"]["mumuni"] = {
"platform": "hermes", "ct": 100, "ip": "192.168.68.24",
"gateway_state": mumuni_data.get("gateway_state", "unknown"),
"telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"),
"zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"),
"email_state": mumuni_platforms.get("email", {}).get("state", "unknown"),
"hermes_version": "",
}
# Get Hermes version
ver = ssh("192.168.68.24", "hermes --version 2>/dev/null | head -1")
if ver:
report["agents"]["mumuni"]["hermes_version"] = ver.split("·")[0].replace("Hermes Agent ","").strip()
return report return report
@@ -539,12 +526,6 @@ th {{ color: #8b949e; font-weight: normal; }}
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜") zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
gateway = agent.get("gateway_state", "?") gateway = agent.get("gateway_state", "?")
processed = "DSH" processed = "DSH"
elif name == "mumuni":
zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜")
gateway = agent.get("gateway_state", "?")
tg = "✅" if agent.get("telegram_state") == "connected" else "❌"
ver = agent.get("hermes_version", "")
processed = f"TG:{tg} v{ver}"
else: else:
zulip_state = "⬜" zulip_state = "⬜"
gateway = agent.get("gateway_state", "?") gateway = agent.get("gateway_state", "?")
+14 -19
View File
@@ -1,7 +1,10 @@
#!/bin/bash #!/bin/bash
# /root/scripts/zulip-monitor.sh — Zulip Mesh Health Monitor # /root/scripts/zulip-monitor.sh — Zulip Mesh Health Monitor
# Implements zulip-health.prose.md v2 # Implements zulip-health.prose.md v3
# Runs every 15 min via cron. Alerts: Zulip private DM to the owner plus a stream post to #agent-hub on topic 'zulip-health'. # Runs every 15 min via cron. Alerts: Zulip private DM to the owner plus a stream post to #agent-hub on topic 'zulip-health'.
# Legs: global Zulip server, Platform A pi/Abiba (the Zulip bridge), Platform B
# Tanko (DSH), Platform C Agent Zero (kagentz). The former Platform B Hermes
# agent leg is retired — see the note after the Tanko leg.
set -euo pipefail set -euo pipefail
ZULIP_SITE="https://chat.sysloggh.net" ZULIP_SITE="https://chat.sysloggh.net"
@@ -21,7 +24,8 @@ notify() {
# Zulip DM to owner # Zulip DM to owner
local content="${severity} Zulip Monitor: ${msg}" local content="${severity} Zulip Monitor: ${msg}"
local form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")" local form
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \ curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \ -u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
-d "${form}" > /dev/null 2>&1 || true -d "${form}" > /dev/null 2>&1 || true
@@ -141,23 +145,14 @@ else
esac esac
fi fi
# ── Platform B: Hermes (Mumuni) ── # ── Removed: the former "Platform B: Hermes" agent leg ──
MUMUNI_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.24 \ # Captain ruling 2026-09-10: that agent moved off this host onto her own
"cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}") # container (kagentz CT 105 on minipve, dedicated `hermes` user) and is now
MUMUNI_ZULIP=$(echo "$MUMUNI_STATE" | python3 -c " # monitored on her side — see the out-of-scope note in zulip-health.prose.md.
import sys,json # The old leg ssh'd to her former CT 100 deployment and read its Hermes gateway
d=json.load(sys.stdin) # state, which reported "unknown" on every run and posted a false 🔴 DM plus an
p=d.get('platforms',{}).get('zulip',{}) # #agent-hub stream alert. Do NOT re-add a probe for her: this monitor must
print(p.get('state','unknown')) # never contact her former host.
" 2>/dev/null)
if [ "$MUMUNI_ZULIP" != "connected" ]; then
notify "🔴" "Mumuni (Hermes) Zulip state: $MUMUNI_ZULIP"
ISSUES=$((ISSUES + 1))
echo " Mumuni: ❌ state=$MUMUNI_ZULIP" >> "$LOG"
else
echo " Mumuni: ✅ Zulip connected" >> "$LOG"
fi
# ── Platform C: Agent Zero (kagentz) ── # ── Platform C: Agent Zero (kagentz) ──
AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \ AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
+284
View File
@@ -0,0 +1,284 @@
"""Regression tests for the 2026-09-10 retirement of the Mumuni monitoring leg.
WHY THIS FILE EXISTS: captain ruling 2026-09-10 — Mumuni moved off this host
onto her own container (kagentz CT 105 on minipve, 192.168.68.14, dedicated
`hermes` user) and is monitored from her side. The monitor nevertheless kept
ssh'ing to root@192.168.68.24 for `~/.hermes/gateway_state.json` on the
decommissioned deployment, read "unknown" on every run, and posted a false 🔴
"Mumuni (Hermes) Zulip state: unknown" DM + #agent-hub stream alert to the
captain. The daily infra digest published a matching `mumuni:unknown` row.
CONTRACT UNDER TEST:
* `scripts/zulip-monitor.sh` carries NO Mumuni probe and NO 192.168.68.24
reference; it never ssh'es .24, and even on a failing run it emits no Mumuni
notify (stdout alert, Zulip payload, or log line).
* The Abiba (pi — the Zulip bridge), Tanko (DSH) and Agent Zero (kagentz) legs
still work: deleting the Mumuni leg must not have gutted the rest.
* `scripts/daily-infra-report.py` no longer probes .24 for a Hermes gateway
state and no longer emits a `mumuni` agent entry.
* `scripts/agent-health-check.py`'s AGENTS roster has no mumuni entry. This is
a pin, not a behavior change — verify the probe was already gone.
* `zulip-health.prose.md` retires the Mumuni-only steps and says explicitly
that Mumuni is not monitored from this host.
HOW: structural greps plus a behavioral sandbox. The sandbox copies the shipped
script verbatim and rewrites only its LOG constant, then runs it with stub
ssh/curl on PATH. The ssh stub records every host it is asked to reach, so
"never probes .24" is asserted from observed behavior, not from source text.
Usage: python3 -m pytest tests/test_mumuni_monitor_removal.py
"""
from __future__ import annotations
import importlib.util
import os
import pathlib
import stat
import subprocess
import pytest
ROOT = pathlib.Path(__file__).resolve().parents[1]
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
DAILY_REPORT = ROOT / "scripts" / "daily-infra-report.py"
AHC = ROOT / "scripts" / "agent-health-check.py"
HEALTH_CONTRACT = ROOT / "zulip-health.prose.md"
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
# ── scripts/zulip-monitor.sh: structural guards ──────────────────────
def test_zulip_monitor_has_no_mumuni_reference():
text = ZULIP_MONITOR.read_text()
assert "mumuni" not in text.lower()
assert "gateway_state.json" not in text
def test_zulip_monitor_has_no_24_reference():
assert MUMUNI_IP not in ZULIP_MONITOR.read_text()
def test_zulip_monitor_keeps_bridge_tanko_and_agent_zero_legs():
text = ZULIP_MONITOR.read_text()
assert "Platform A: pi (Abiba)" in text # the Zulip bridge
assert "Platform B: Tanko" in text
assert "Platform C: Agent Zero" in text
assert TANKO_VANTAGE in text
assert AGENT_ZERO_HOST in text
assert "kagentz" in text
# ── scripts/zulip-monitor.sh: behavioral sandbox ─────────────────────
SSH_STUB = r"""#!/usr/bin/env bash
# Stub ssh: record the target host, then answer by host + remote command.
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
host=""
for a in "$@"; do
case "$a" in
*@192.168.*) host="${a##*@}" ;;
esac
done
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
cmd="${*: -1}"
case "$host" in
192.168.68.15)
case "$cmd" in
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
*curl*) printf '%s' "$TANKO_HTTP" ;;
esac ;;
192.168.68.14)
case "$cmd" in
*agent.json*) printf '%s' "$AZ_A2A" ;;
*"ps aux"*) printf '%s\n' "$AZ_PS" ;;
esac ;;
*)
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
esac
exit 0
"""
CURL_STUB = r"""#!/usr/bin/env bash
# Stub curl: serve the Abiba health fixture and the Zulip server 200, and
# record every call (including notify) payloads.
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
case "$*" in
*:9200/health*)
case " $* " in
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
*) printf '%s' "$PI_BODY" ;; # body probe
esac ;;
*server_settings*)
printf '%s' "$SERVER_HTTP" ;;
esac
exit 0
"""
def _write_exec(path: pathlib.Path, body: str) -> None:
path.write_text(body)
path.chmod(path.stat().st_mode
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
az_a2a='{"name":"kagentz"}',
az_ps="root 111 0.1 0.2 /opt/venv-a0/bin/python3 -u adapter.py"):
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
Only the LOG constant is rewritten (to keep the run inside the worktree).
Everything else — legs, labels, notify logic — is the shipped script.
"""
sandbox = tmp_path / "sandbox"
bindir = sandbox / "bin"
record = sandbox / "record"
bindir.mkdir(parents=True)
record.mkdir()
_write_exec(bindir / "ssh", SSH_STUB)
_write_exec(bindir / "curl", CURL_STUB)
source = ZULIP_MONITOR.read_text()
log_line = 'LOG="/root/zulip-health-monitor.log"'
assert log_line in source, "LOG constant moved — update the sandbox harness"
log_path = sandbox / "zulip-health-monitor.log"
script = sandbox / "zulip-monitor.sh"
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
env = dict(os.environ)
env.update({
"PATH": f"{bindir}:{env['PATH']}",
"RECORD_DIR": str(record),
"TANKO_SVC": tanko_svc,
"TANKO_HTTP": tanko_http,
"AZ_A2A": az_a2a,
"AZ_PS": az_ps,
"PI_HTTP": "200",
"PI_BODY": CONNECTED_FIXTURE.read_text(),
"SERVER_HTTP": "200",
})
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
capture_output=True, text=True)
return proc, record, log_path
def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
proc, record, log_path = _run_monitor(tmp_path)
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
# Every retained leg actually ran and passed.
assert "Server: ✅ HTTP 200" in log
assert "Abiba: ✅ Connected" in log
assert "Tanko: ✅ service=active http=200" in log
assert "kagentz: ✅ A2A alive" in log
assert "kagentz: ✅ Adapter running" in log
assert "Result: ✅ All healthy" in log
# A healthy run emits no notify at all — and certainly no Mumuni one.
assert proc.stdout == ""
assert "Mumuni" not in log
assert "🔴" not in log
# Observed behavior: .24 is never resolved, only Tanko's vantage and the
# Agent Zero host are contacted.
hosts = record.joinpath("ssh.hosts").read_text().split()
assert MUMUNI_IP not in hosts
assert set(hosts) == {TANKO_VANTAGE, AGENT_ZERO_HOST}
assert not record.joinpath("unexpected-ssh").exists()
def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
# Failure path: exercises notify() end to end so "no Mumuni notify" is
# proven on the alert path, not only on the quiet healthy path.
proc, record, log_path = _run_monitor(tmp_path, tanko_svc="inactive",
tanko_http="000")
assert proc.returncode == 0, proc.stderr
alerts = proc.stdout
assert "Tanko (DSH dsh-web) service state: inactive" in alerts
assert "1 issue(s) found" in alerts
# No Mumuni text in stdout, the log, or any Zulip DM/stream payload.
assert "Mumuni" not in alerts
assert "Mumuni" not in log_path.read_text()
assert MUMUNI_IP not in alerts + log_path.read_text()
payloads = record.joinpath("curl.calls").read_text()
assert "Mumuni" not in payloads
assert MUMUNI_IP not in payloads
# The rest of the monitor still ran alongside the failing Tanko leg.
log = log_path.read_text()
assert "Abiba: ✅ Connected" in log
assert "kagentz: ✅ A2A alive" in log
assert "Result: 🔴 1 issue(s) found" in log
# ── scripts/daily-infra-report.py: no Mumuni agent probe ─────────────
def test_daily_report_has_no_mumuni_agent_probe():
text = DAILY_REPORT.read_text()
assert 'report["agents"]["mumuni"]' not in text
assert f'ssh("{MUMUNI_IP}"' not in text
assert "hermes --version" not in text
def test_daily_report_keeps_abiba_and_tanko_agent_legs():
text = DAILY_REPORT.read_text()
assert 'report["agents"]["abiba"]' in text
assert 'report["agents"]["tanko"]' in text
def test_daily_report_still_describes_abiba_at_its_own_ct100_ip():
# .24 is abiba's own CT 100 address — the IP itself is legitimate; only a
# Mumuni gateway probe against it was the fault.
assert '"platform": "pi", "ct": 100, "ip": "192.168.68.24"' in \
DAILY_REPORT.read_text()
# ── scripts/agent-health-check.py: roster pin ───────────────────────
@pytest.fixture(scope="module")
def ahc():
spec = importlib.util.spec_from_file_location("agent_health_check_roster", AHC)
assert spec and spec.loader
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
def test_agent_health_roster_has_no_mumuni_entry(ahc):
assert "mumuni" not in ahc.AGENTS
def test_agent_health_roster_comment_no_longer_places_mumuni_at_24():
text = AHC.read_text()
assert "mumuni (.24" not in text
assert "mumuni (.24, inside abiba CT100)" not in text
# ── zulip-health.prose.md: contract reconciliation ──────────────────
def test_health_contract_retires_mumuni_only_steps():
text = HEALTH_CONTRACT.read_text()
assert MUMUNI_IP not in text
for step in ("**B4:", "**B5:", "**B6:"):
assert step not in text
def test_health_contract_states_mumuni_is_not_monitored_from_this_host():
text = HEALTH_CONTRACT.read_text()
assert "Mumuni is NOT monitored from this host" in text
assert "monitored on her side" in text
assert "her own container" in text
def test_health_contract_keeps_tanko_agent_zero_and_bridge_steps():
text = HEALTH_CONTRACT.read_text()
for marker in ("**B1:", "**B2:", "**B3:", "Step 4: Platform C",
"Step 2: Platform A", "Step 1: Zulip Server Liveness"):
assert marker in text, marker
+25 -43
View File
@@ -1,9 +1,9 @@
--- ---
kind: responsibility kind: responsibility
name: zulip-health name: zulip-health
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Tanko on DSH / Mumuni on Hermes), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity. description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
title: Zulip Mesh Health Monitor — Multi-Platform title: Zulip Mesh Health Monitor — Multi-Platform
version: 3.0.0 version: 3.1.0
runtime_contract: 2 runtime_contract: 2
agent: abiba agent: abiba
report_only_agents: report_only_agents:
@@ -12,13 +12,23 @@ report_only_agents:
# Zulip Mesh Health Monitor # Zulip Mesh Health Monitor
Monitors ALL Zulip-connected agents across platforms (pi, Hermes, DSH, Agent Zero). Monitors the Zulip-connected agents under this host's operational control (pi,
Runs every 15 minutes in the background. Also triggers on session start. DSH, Agent Zero). Runs every 15 minutes in the background. Also triggers on
session start.
> **Mumuni is NOT monitored from this host (captain ruling 2026-09-10).** She
> moved off this host onto her own container — kagentz CT 105 on minipve
> (192.168.68.14), running a dedicated `hermes` user — and is monitored on her
> side. No step in this contract, and no leg of `scripts/zulip-monitor.sh`, may
> ssh to her old deployment, read her `~/.hermes/gateway_state.json`, or alert on
> her state. The former Platform-B-for-Mumuni steps (gateway process, heartbeat,
> response delivery) are retired: they always read "unknown" against the
> decommissioned deployment and produced a false 🔴 alert on every run.
## Requires ## Requires
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY` - **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); Mumuni (192.168.68.14, kagentz CT105 on minipve); and Agent Zero Docker host (192.168.68.14) - **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
- **PM2** on localhost for pi process management - **PM2** on localhost for pi process management
- **Network access** to `chat.sysloggh.net`, `localhost:9200` - **Network access** to `chat.sysloggh.net`, `localhost:9200`
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce` - **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
@@ -81,13 +91,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse
## Streaming Support (2026-07-05) ## Streaming Support (2026-07-05)
Zulip agents now support progressive message editing during agent generation. Zulip agents now support progressive message editing during agent generation.
When a Zulip agent (Tanko on DSH, Mumuni on Hermes) processes a message, the response is When a Zulip agent under this monitor's scope (Tanko on DSH) processes a message, the response is
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API: streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
- Adapter implements `edit_message()` using `_api_patch()` helper - Adapter implements `edit_message()` using `_api_patch()` helper
- Gateway stream consumer progressively edits the Zulip message - Gateway stream consumer progressively edits the Zulip message
- User sees real-time agent thinking instead of waiting for full response - User sees real-time agent thinking instead of waiting for full response
- Verified: Tanko (CT 112) and Mumuni (kagentz CT 105) both have streaming active - Verified: Tanko (CT 112) has streaming active; Mumuni's (kagentz CT 105) is verified on her own host, not from here
### Verification ### Verification
```bash ```bash
@@ -183,7 +193,10 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
| Crash loop >10/h | Alert user | | Crash loop >10/h | Alert user |
### Step 3: Platform B — Tanko (DSH on amdpve CT 112) & Mumuni (Hermes) ### Step 3: Platform B — Tanko (DSH on amdpve CT 112)
Mumuni is out of scope for this host (see the note above): she runs on her own
container and is monitored on her side.
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so
there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs
@@ -240,44 +253,13 @@ timeout only. Never expect a bare `200` — the public URL terminates in the
token-gated authentik chain. Statuses outside the healthy set are token-gated authentik chain. Statuses outside the healthy set are
logged/reported as a warning — reported, never healed on. logged/reported as a warning — reported, never healed on.
**B4: Gateway Process** (Hermes agent Mumuni only — Tanko runs no Hermes gateway)
```bash
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
```
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more
than one `gateway run` process is found, the gateway has a collision (typically
one `--force` and one `--replace` process). Kill the newer/duplicate process,
then restart the remaining gateway per-agent (parameterized 2026-08-09, captain
ruling). Check the gateway log for "Gateway running with 2 platform(s)" (not 1)
to confirm Zulip reloaded.
**B5: Heartbeat Verification** (Hermes agent Mumuni only — Tanko has no Hermes gateway)
```bash
ssh root@192.168.68.24 "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
```
Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
Silence > 300s → warning. Silence > 600s → critical.
**B6: Response Delivery** (Hermes agent Mumuni only)
```bash
ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
```
> 50% fail rate → critical.
**Platform B Actions** **Platform B Actions**
| Condition | Action | | Condition | Action |
|-----------|--------| |-----------|--------|
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` (Mumuni) / restart Tanko via DSH service | | `dsh-web` service not `active` | Restart Tanko via DSH service |
| No heartbeat in 10min | Same as above | | HTTP `:3080` connection refused/timeout (`000`) | Same as above |
| `Failed to finalize` > 50% | Check PATCH API, Zulip server | | HTTP status outside the expected set | Log/report as a warning — reported, never healed on |
| Response empty/short | Check A2A endpoint / LiteLLM model |
### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14) ### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)
@@ -333,7 +315,7 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`. I
Check each agent's log for excessive bot-to-bot chatter: Check each agent's log for excessive bot-to-bot chatter:
- Abiba: `Skipped.*bot msgs` count - Abiba: `Skipped.*bot msgs` count
- Tanko/Mumuni: Repeated DM exchanges between bots - Tanko: Repeated DM exchanges between bots
- kagentz: Adapter log for bot DMs being processed - kagentz: Adapter log for bot DMs being processed
If any bot processes >50 bot-originated messages in 15min → warning. If any bot processes >50 bot-originated messages in 15min → warning.