Captain ruling 2026-09-10: Mumuni moved off this host onto her own container (kagentz CT 105 on minipve, 192.168.68.14, dedicated `hermes` user) and is monitored from her side. This host must not monitor anything Mumuni. The stale probes fired false alerts repeatedly: * scripts/zulip-monitor.sh ssh'd to root@192.168.68.24 for the decommissioned deployment's ~/.hermes/gateway_state.json, read "unknown" on every run, and posted a 🔴 "Mumuni (Hermes) Zulip state: unknown" DM + #agent-hub stream alert each cycle. * scripts/daily-infra-report.py published a matching "mumuni:unknown" row in every digest. Changes: * zulip-monitor.sh: delete the "Platform B: Hermes (Mumuni)" leg and its notify; keep the Zulip-server, Platform A pi/Abiba (bridge), Platform B Tanko and Platform C Agent Zero legs. A comment records why the leg is retired so it is not re-added. Also fixes SC2155 so shellcheck is clean. * daily-infra-report.py: delete the .24 ~/.hermes/gateway_state.json agent probe, its hermes --version probe, and the now-dead mumuni render branch. Abiba (CT 100, its own .24 address) and Tanko legs unchanged; the PVE API token name is untouched. * zulip-health.prose.md (v3.1.0): drop the Mumuni-only B4 gateway-process, B5 heartbeat and B6 response-delivery steps and the stale 192.168.68.24 references; state explicitly that Mumuni is not monitored from this host. Tanko/Agent-Zero/bridge steps retained. * agent-health-check.py: correct the v2 changelog roster comment that still placed mumuni at .24/CT100. No behavior change — the mumuni probe was already absent from the AGENTS dict; v5 changelog notes the correction. Tests: tests/test_mumuni_monitor_removal.py pins the removal structurally and behaviorally — the shipped zulip-monitor.sh is run in a sandbox (only its LOG constant rewritten) with stub ssh/curl on PATH; the ssh stub records every target host, so "never reaches .24" and "no Mumuni notify even on the alert path" are asserted from observed behavior. A mutation check (re-inject the old leg) fails the suite, so the guarantee is not vacuous.
194 lines
9.2 KiB
Bash
Executable File
194 lines
9.2 KiB
Bash
Executable File
#!/bin/bash
|
|
# /root/scripts/zulip-monitor.sh — Zulip Mesh Health Monitor
|
|
# Implements zulip-health.prose.md v3
|
|
# Runs every 15 min via cron. Alerts: Zulip private DM to the owner plus a stream post to #agent-hub on topic 'zulip-health'.
|
|
# Legs: global Zulip server, Platform A pi/Abiba (the Zulip bridge), Platform B
|
|
# Tanko (DSH), Platform C Agent Zero (kagentz). The former Platform B Hermes
|
|
# agent leg is retired — see the note after the Tanko leg.
|
|
set -euo pipefail
|
|
|
|
ZULIP_SITE="https://chat.sysloggh.net"
|
|
ZULIP_EMAIL="abiba-bot@chat.sysloggh.net"
|
|
ZULIP_KEY="cKTDMZAPW08dk3zl05sStzO7HRztzyn8"
|
|
OWNER_ZULIP_ID="9"
|
|
|
|
|
|
LOG="/root/zulip-health-monitor.log"
|
|
TIMESTAMP=$(date -u '+%Y-%m-%d %H:%M UTC')
|
|
ISSUES=0
|
|
echo "=== Zulip Health Check — $TIMESTAMP ===" >> "$LOG"
|
|
|
|
notify() {
|
|
local severity="$1" msg="$2"
|
|
echo "[$severity] $msg"
|
|
|
|
# Zulip DM to owner
|
|
local content="${severity} Zulip Monitor: ${msg}"
|
|
local form
|
|
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
|
|
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
|
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
|
-d "${form}" > /dev/null 2>&1 || true
|
|
# Zulip stream post to #agent-hub on topic 'zulip-health'
|
|
local stream_content="${severity} Zulip Monitor: ${msg}"
|
|
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
|
|
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
|
|
-d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \
|
|
> /dev/null 2>&1 \
|
|
|| echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)" >> "$LOG"
|
|
}
|
|
|
|
# ── Global: Zulip Server ──
|
|
SERVER_CODE=$(curl -s -o /dev/null -w "%{http_code}" --connect-timeout 10 \
|
|
https://chat.sysloggh.net/api/v1/server_settings \
|
|
-u 'abiba-bot@chat.sysloggh.net:cKTDMZAPW08dk3zl05sStzO7HRztzyn8' 2>/dev/null || echo "000")
|
|
if [ "$SERVER_CODE" != "200" ]; then
|
|
notify "🔴" "Zulip server returned HTTP $SERVER_CODE"
|
|
ISSUES=$((ISSUES + 1))
|
|
else
|
|
echo " Server: ✅ HTTP 200" >> "$LOG"
|
|
fi
|
|
|
|
# ── Platform A: pi (Abiba) ──
|
|
# Probes the pi Zulip extension health endpoint (:9200/health, served by the
|
|
# extension's startHealthServer; shape documented in zulip-health.prose.md).
|
|
# FAIL-SAFE contract (pinned by tests/zulip-monitor-abiba.sh): connection state
|
|
# lives NESTED at zulip.connected / zulip.last_error — there is no top-level
|
|
# `connected` and no retry counter in the payload. A fetch error, non-2xx
|
|
# response, empty/unparseable body, or payload missing a boolean
|
|
# zulip.connected is a PROBE FAILURE: it alerts and NEVER calls pm2 restart.
|
|
# pm2 restart runs ONLY on affirmative zulip.connected=false.
|
|
# -- abiba-leg-start (verbatim-extracted by tests/zulip-monitor-abiba.sh)
|
|
PI_HTTP=$(curl -s -o /dev/null --connect-timeout 5 --max-time 10 -w '%{http_code}' http://localhost:9200/health 2>/dev/null || echo "000")
|
|
PI_BODY=$(curl -s --connect-timeout 5 --max-time 10 http://localhost:9200/health 2>/dev/null || true)
|
|
PI_STATE=$(printf '%s' "$PI_BODY" | python3 -c '
|
|
import sys, json
|
|
code = sys.argv[1]
|
|
body = sys.stdin.read()
|
|
try:
|
|
d = json.loads(body)
|
|
except Exception:
|
|
sys.stdout.write("probe-failed|unparseable body")
|
|
sys.exit(0)
|
|
if not code.startswith("2"):
|
|
sys.stdout.write("probe-failed|HTTP %s" % code)
|
|
sys.exit(0)
|
|
if not isinstance(d, dict) or not isinstance(d.get("zulip"), dict):
|
|
sys.stdout.write("probe-failed|missing zulip.connected")
|
|
sys.exit(0)
|
|
z = d["zulip"]
|
|
if "connected" not in z or not isinstance(z["connected"], bool):
|
|
sys.stdout.write("probe-failed|missing or non-boolean zulip.connected")
|
|
sys.exit(0)
|
|
err = z.get("last_error") or ""
|
|
if z["connected"]:
|
|
if err:
|
|
sys.stdout.write("degraded|%s" % err)
|
|
else:
|
|
sys.stdout.write("healthy|%s" % z.get("messages_processed", 0))
|
|
else:
|
|
sys.stdout.write("disconnected|")
|
|
' "$PI_HTTP" 2>/dev/null) || PI_STATE="probe-failed|python error"
|
|
PI_VERDICT=${PI_STATE%%|*}
|
|
PI_DETAIL=${PI_STATE#*|}
|
|
|
|
case "$PI_VERDICT" in
|
|
healthy)
|
|
echo " Abiba: ✅ Connected (processed=$PI_DETAIL)" >> "$LOG" ;;
|
|
degraded)
|
|
notify "🟡" "Abiba pi extension error: ${PI_DETAIL:0:100}"
|
|
echo " Abiba: 🟡 Error: ${PI_DETAIL:0:100}" >> "$LOG" ;;
|
|
disconnected)
|
|
notify "🔴" "Abiba pi extension DISCONNECTED — restarting"
|
|
pm2 restart abiba-zulip 2>/dev/null || true
|
|
ISSUES=$((ISSUES + 1))
|
|
echo " Abiba: ❌ Disconnected — restarted" >> "$LOG" ;;
|
|
probe-failed)
|
|
notify "🟠" "Abiba pi extension health probe FAILED (${PI_DETAIL}; HTTP $PI_HTTP) — NOT restarting, manual check needed"
|
|
ISSUES=$((ISSUES + 1))
|
|
echo " Abiba: ⚠️ Probe failed (${PI_DETAIL}; HTTP $PI_HTTP) — NOT restarted" >> "$LOG" ;;
|
|
*)
|
|
notify "🟠" "Abiba pi extension health probe returned unexpected verdict (${PI_STATE}) — NOT restarting, manual check needed"
|
|
ISSUES=$((ISSUES + 1))
|
|
echo " Abiba: ⚠️ Unexpected probe verdict (${PI_STATE}) — NOT restarted" >> "$LOG" ;;
|
|
esac
|
|
# -- abiba-leg-end
|
|
|
|
# ── Platform B: Tanko (DSH dsh-web on amdpve CT 112) ──
|
|
# Direct SSH to 192.168.68.122 is not a dependency of this monitor — per-worker
|
|
# key availability varies — so probes run from the amdpve vantage via `pct exec`.
|
|
# Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112 on amdpve
|
|
# (192.168.68.15). The gateway binds 127.0.0.1:3080 loopback-only by design — a
|
|
# remote :3080 probe is refused and is NOT a fault.
|
|
TANKO_SVC=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
|
|
"pct exec 112 -- systemctl is-active dsh-web" 2>/dev/null || true)
|
|
[ -n "$TANKO_SVC" ] || TANKO_SVC="unknown"
|
|
TANKO_HTTP=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \
|
|
"pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/" 2>/dev/null || true)
|
|
[ -n "$TANKO_HTTP" ] || TANKO_HTTP="000"
|
|
|
|
if [ "$TANKO_SVC" != "active" ]; then
|
|
notify "🔴" "Tanko (DSH dsh-web) service state: $TANKO_SVC — needs restart"
|
|
ISSUES=$((ISSUES + 1))
|
|
echo " Tanko: ❌ service=$TANKO_SVC" >> "$LOG"
|
|
elif [ "$TANKO_HTTP" = "000" ]; then
|
|
notify "🔴" "Tanko (DSH dsh-web) HTTP :3080 connection refused/timeout — needs restart"
|
|
ISSUES=$((ISSUES + 1))
|
|
echo " Tanko: ❌ http=000 (refused/timeout)" >> "$LOG"
|
|
else
|
|
case "$TANKO_HTTP" in
|
|
200|301|302|307|308|401|403)
|
|
echo " Tanko: ✅ service=active http=$TANKO_HTTP" >> "$LOG" ;;
|
|
*)
|
|
notify "🟡" "Tanko (DSH dsh-web) HTTP :3080 answered $TANKO_HTTP — running, unexpected status"
|
|
echo " Tanko: 🟡 service=active http=$TANKO_HTTP (running, warning)" >> "$LOG" ;;
|
|
esac
|
|
fi
|
|
|
|
# ── Removed: the former "Platform B: Hermes" agent leg ──
|
|
# Captain ruling 2026-09-10: that agent moved off this host onto her own
|
|
# container (kagentz CT 105 on minipve, dedicated `hermes` user) and is now
|
|
# monitored on her side — see the out-of-scope note in zulip-health.prose.md.
|
|
# The old leg ssh'd to her former CT 100 deployment and read its Hermes gateway
|
|
# state, which reported "unknown" on every run and posted a false 🔴 DM plus an
|
|
# #agent-hub stream alert. Do NOT re-add a probe for her: this monitor must
|
|
# never contact her former host.
|
|
|
|
# ── Platform C: Agent Zero (kagentz) ──
|
|
AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
|
"docker exec agent-zero curl -s --connect-timeout 5 http://127.0.0.1:8001/.well-known/agent.json 2>/dev/null" 2>/dev/null || echo "")
|
|
AZ_ALIVE=$(echo "$AZ_A2A" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('name',''))" 2>/dev/null)
|
|
|
|
if [ "$AZ_ALIVE" != "kagentz" ]; then
|
|
notify "🔴" "kagentz A2A server DOWN — restarting"
|
|
ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
|
"docker exec agent-zero bash -c 'pkill -9 -f a2a_agent; sleep 1; cd /a0 && /opt/venv-a0/bin/python3 -u /a0/usr/a2a_agent.py > /tmp/a2a.log 2>&1 &'" 2>/dev/null || true
|
|
ISSUES=$((ISSUES + 1))
|
|
echo " kagentz: ❌ A2A down — restarted" >> "$LOG"
|
|
else
|
|
echo " kagentz: ✅ A2A alive" >> "$LOG"
|
|
|
|
# Check adapter process
|
|
AZ_ADAPTER=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
|
"docker exec agent-zero ps aux 2>/dev/null | grep adapter | grep -v grep | wc -l" 2>/dev/null || echo "0")
|
|
if [ "$AZ_ADAPTER" -lt 1 ]; then
|
|
notify "🔴" "kagentz Zulip adapter DOWN — restarting"
|
|
ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
|
|
"docker exec agent-zero bash -c 'cd /a0/usr/kagentz-zulip && ZULIP_SITE=https://chat.sysloggh.net ZULIP_EMAIL=kagentz-bot@chat.sysloggh.net ZULIP_API_KEY=E9q9PXJTxftPYBkb5pBDWupDO7KK21ty ZULIP_AGENT_NAME=kagentz A2A_URL=http://localhost:8001/a2a A2A_TOKEN=8zNgdOEXzYxjQvTl /opt/venv-a0/bin/python3 -u adapter.py > /tmp/zulip-adapter.log 2>&1 &'" 2>/dev/null || true
|
|
ISSUES=$((ISSUES + 1))
|
|
echo " kagentz: ❌ Adapter down — restarted" >> "$LOG"
|
|
else
|
|
echo " kagentz: ✅ Adapter running" >> "$LOG"
|
|
fi
|
|
fi
|
|
|
|
# ── Summary ──
|
|
if [ "$ISSUES" -eq 0 ]; then
|
|
echo " Result: ✅ All healthy" >> "$LOG"
|
|
else
|
|
echo " Result: 🔴 $ISSUES issue(s) found" >> "$LOG"
|
|
notify "🔴" "$ISSUES issue(s) found — check /root/zulip-health-monitor.log"
|
|
fi
|
|
|
|
echo "" >> "$LOG"
|