Compare commits

...
Author SHA1 Message Date
root 287657a77a Merge remote-tracking branch 'origin/master' into update/docker-ecosystems-20260908
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-10 23:01:32 +00:00
root 21f9073e0b fix(daily-infra-report): use pve_probe_status in render + add regression tests (PR #64 round 2) 2026-09-10 23:01:18 +00:00
root 32fe7c0652 fix(daily-infra-report): complete Zulip nested-key fix + Proxmox port/auth fix (PR #64 review fix 1-2)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-10 22:03:48 +00:00
root 25cf2f5eef fix(daily-infra-report): read Zulip state from nested 'zulip' key (PR #64 fix 3a)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-10 20:55:17 +00:00
root 26f2301188 fix(pm2-self-heal): remove stray fragment so script parses (PR #64 fix 2)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-10 20:43:01 +00:00
root a6a459acc0 fix: correct Firecrawl GET / response code from 404 to 200 (PR #64)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
2026-09-10 20:42:15 +00:00
abiba-bot 14bed6e916 Merge pull request 'fix(monitoring): stop monitoring Mumuni from this host (moved to CT105/.14)' (#71) from fm/denya-mumuni-monitor-removal-20260910 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
2026-09-10 12:30:19 +00:00
agent-zero 2e70c834cb docs(infrastructure-update): add mandatory digest-pin sweep to Wave 3 verification
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Digest-pinned images (image: repo@sha256:...) are invisible to 'docker compose
pull' — the pin re-pulls the same digest forever, so new releases never appear.
Verified 2026-09-10: audiobookshelf ran 2.34.0 for 7 weeks despite weekly pulls;
un-pinned to :latest, now 2.36.0 (HTTP 200). Dockhand stack (also pinned) was
removed 2026-09-10 as unused (user decision). Weekly task qSOOVzsU now sweeps
for digest pins every run and flags them for user-approved un-pinning.
2026-09-10 06:30:07 -04:00
abiba-bot 4f59b82404 no-mistakes(document): Refresh stale Mumuni roster docs and registry metadata
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-09-10 10:26:00 +00:00
abiba-bot 8a4ee4cf5a no-mistakes(review): Pin Mumuni removal behaviorally and fix digest import bug 2026-09-10 10:18:38 +00:00
abiba 952fca9c92 fix(monitoring): retire the Mumuni leg — stop probing her decommissioned .24 deployment
Captain ruling 2026-09-10: Mumuni moved off this host onto her own container
(kagentz CT 105 on minipve, 192.168.68.14, dedicated `hermes` user) and is
monitored from her side. This host must not monitor anything Mumuni.

The stale probes fired false alerts repeatedly:
  * scripts/zulip-monitor.sh ssh'd to root@192.168.68.24 for the
    decommissioned deployment's ~/.hermes/gateway_state.json, read "unknown"
    on every run, and posted a 🔴 "Mumuni (Hermes) Zulip state: unknown" DM +
    #agent-hub stream alert each cycle.
  * scripts/daily-infra-report.py published a matching "mumuni:unknown" row in
    every digest.

Changes:
  * zulip-monitor.sh: delete the "Platform B: Hermes (Mumuni)" leg and its
    notify; keep the Zulip-server, Platform A pi/Abiba (bridge), Platform B
    Tanko and Platform C Agent Zero legs. A comment records why the leg is
    retired so it is not re-added. Also fixes SC2155 so shellcheck is clean.
  * daily-infra-report.py: delete the .24 ~/.hermes/gateway_state.json agent
    probe, its hermes --version probe, and the now-dead mumuni render branch.
    Abiba (CT 100, its own .24 address) and Tanko legs unchanged; the PVE API
    token name is untouched.
  * zulip-health.prose.md (v3.1.0): drop the Mumuni-only B4 gateway-process,
    B5 heartbeat and B6 response-delivery steps and the stale 192.168.68.24
    references; state explicitly that Mumuni is not monitored from this host.
    Tanko/Agent-Zero/bridge steps retained.
  * agent-health-check.py: correct the v2 changelog roster comment that still
    placed mumuni at .24/CT100. No behavior change — the mumuni probe was
    already absent from the AGENTS dict; v5 changelog notes the correction.

Tests: tests/test_mumuni_monitor_removal.py pins the removal structurally and
behaviorally — the shipped zulip-monitor.sh is run in a sandbox (only its LOG
constant rewritten) with stub ssh/curl on PATH; the ssh stub records every
target host, so "never reaches .24" and "no Mumuni notify even on the alert
path" are asserted from observed behavior. A mutation check (re-inject the old
leg) fails the suite, so the guarantee is not vacuous.
2026-09-10 10:11:19 +00:00
abiba-bot b1462f3e79 Merge pull request 'fix(monitoring): probe-drift round 2 — gpu port, PVE liveness, report-only visibility, report provenance' (#70) from fm/probe-drift-round2-20260909 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Failing after 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Skipped
2026-09-10 02:17:01 +00:00
agent-zero bfdff13ae7 docs(infrastructure-update): point weekly trigger at scheduler task qSOOVzsU
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
Weekly Sunday 03:00 America/New_York run is now implemented as Agent Zero
scheduled task 'weekly-fleet-docker-update' (qSOOVzsU) with built-in
post-update service verification and old-image pruning.
2026-09-08 05:27:38 -04:00
agent-zero 8bf32f6f0f update(infrastructure-update): v1.3.0 — full Docker ecosystem coverage from 2026-09-08 live run
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
- Add 5 Docker ecosystems: docker-vm .7, CT 116 .116, CT 117 (zulip+jitsi),
  hwpve .11 (Authentik), NetBird VPS 72.61.0.17
- Wave 3: add trove-test, docker-stats, monitoring stack, Zulip (pct exec + compose
  recreate preserves zulip_default network), Jitsi, Authentik, NetBird
- Firecrawl test is now POST /v1/search (GET / returns 404 by design)
- Add Authentik 302 + NetBird dashboard 200 checks
- Document harness-litellm 3-5 min cold start after recreate (verified 2026-09-08)
- Add hwpve + VPS compose files to config backup list; success criteria covers 5 hosts
2026-09-08 03:23:03 -04:00
10 changed files with 619 additions and 131 deletions
+3 -3
View File
@@ -628,18 +628,18 @@ contracts:
sensitivity: high
status: active
owner: abiba
version: 3.0.0
version: 3.1.0
trigger:
type: scheduled
cadence: '*/15 * * * *'
description: "Every 15 minutes \u2014 monitors all Zulip-connected agents"
description: "Every 15 minutes \u2014 monitors the Zulip-connected agents under this host's control (pi, DSH, Agent Zero)"
cron_job_id: null
execution:
agent: abiba
timeout: 120
requires:
- Zulip API key for abiba-bot@chat.sysloggh.net
- SSH access to all Hermes agents
- SSH access to amdpve (192.168.68.15) for Tanko (CT 112) and the Agent Zero Docker host (.14)
verification:
postconditions:
- check: bot registration active
+19 -6
View File
@@ -3,15 +3,16 @@ kind: responsibility
name: infrastructure-update
description: >
Autonomous system-wide update contract covering all 5 Proxmox nodes,
15+ containers/VMs, and 4 Docker ecosystems. Updates apt packages,
15+ containers/VMs, and 5 Docker ecosystems (docker-vm .7, CT 116 .116,
CT 117, hwpve .11, NetBird VPS 72.61.0.17). Updates apt packages,
Docker images, and container stacks in safe waves with health checks
and automatic rollback on failure.
agent: abiba
triggers:
- on "infra update" command
- weekly (Sunday 03:00 EDT) via cron
- weekly (Sunday 03:00 America/New_York) via Agent Zero scheduler task "weekly-fleet-docker-update" (qSOOVzsU) — implemented 2026-09-08
- on security advisory relay from Mumuni
version: 1.2.0
version: 1.3.0
---
## Maintains
@@ -80,7 +81,13 @@ Before ANY update wave:
| VM 109 (.7) | Home stack (Pulse, Stirling PDF) — JDownloader moved to CT 118 LXC 2026-08-01 | `cd /opt/home_stack && docker compose pull && docker compose up -d` |
| VM 109 (.7) | Audiobookshelf | `cd /opt/audiobookshelf && docker compose pull && docker compose up -d` |
| CT 116 (.116) | Inference Harness (LiteLLM, Prometheus, Grafana) | `cd /opt/inference-harness && docker compose pull && docker compose up -d` |
| CT 117 (zulip, storepve) | Zulip | `docker pull zulip/docker-zulip:latest && docker restart zulip-zulip-1` |
| CT 117 (storepve) | Zulip | `pct exec 117 -- bash -c 'cd /opt/zulip && docker compose pull && docker compose up -d'` (from storepve; compose recreates on zulip_default network) |
| CT 117 (storepve) | Jitsi | `pct exec 117 -- bash -c 'cd /opt/jitsi && docker compose pull && docker compose up -d'` (from storepve) |
| hwpve (.11) | Authentik (server, worker, postgres) | `ssh root@192.168.68.11 'cd /root && docker compose pull && docker compose up -d'` |
| NetBird VPS (72.61.0.17) | NetBird (server, dashboard, proxy, traefik, crowdsec) | `ssh root@72.61.0.17 'cd /root && docker compose pull && docker compose up -d'` |
| VM 109 (.7) | Trove test | `cd /opt/trove-test && docker compose pull && docker compose up -d` |
| VM 109 (.7) | docker-stats | `cd /opt/docker-stats && docker compose pull && docker compose up -d` |
| CT 116 (.116) | Monitoring (Grafana, Prometheus, Alertmanager, PVE exporter) | `cd /opt/monitoring && docker compose pull && docker compose up -d` |
**Verify after Wave 3:**
- All containers healthy: `docker ps` on each host
@@ -88,8 +95,12 @@ Before ANY update wave:
- MCP integration test: `curl localhost:4000/mcp-rest/tools/list -H "Authorization: Bearer $MASTER_KEY"` → 90 tools (23 RA-H OS + 67 GitHub)
- Zulip test: send test message to #agent-hub
- Dashboard loading: `curl localhost:3001/` (via CT 116)
- Firecrawl test: `curl :3002/`
- Firecrawl test: `curl -X POST http://192.168.68.7:3002/v1/search -H 'Content-Type: application/json' -d '{"query":"health","limit":1}'` → `"success":true` (GET `/` returns 200)
- Authentik test: `curl http://192.168.68.11:9000/` → 302 redirect to login
- NetBird test: `curl -s -o /dev/null -w '%{http_code}' https://netbird.sysloggh.net/` → 200
- harness-litellm cold start: allow 3-5 min after recreate — reports unhealthy and :4000 refuses connections while loading config/DB, then recovers to 200 on its own (verified 2026-09-08)
- SearXNG test: `curl :8888`
- Digest-pin sweep: `grep -rn '@sha256:' /opt/*/docker-compose.y*` on every host — digest-pinned images are INVISIBLE to `docker compose pull` (the pin re-pulls the same digest forever, so new releases never appear). Flag every pin in the run report and propose un-pinning to a floating tag with user approval before editing. Found 2026-09-10: audiobookshelf was digest-pinned at 2.34.0 (container created 2026-07-18) and silently missed by every sweep; dockhand stack was also pinned (stack removed 2026-09-10, unused). After un-pinning audiobookshelf to :latest it updated to 2.36.0 and verified HTTP 200.
## Wave 4: Proxmox Kernel Reboot
@@ -158,6 +169,8 @@ Before Wave 1, snapshot these files:
/opt/search-stack/searxng/docker-compose.yml (VM 109 .7)
/opt/home_stack/docker-compose.yml (VM 109 .7)
/opt/audiobookshelf/docker-compose.yml (VM 109 .7)
/root/compose.yml (hwpve .11 — Authentik server/worker/postgres)
/root/docker-compose.yml (NetBird VPS — netbird server/dashboard/proxy, traefik, crowdsec)
/root/.pi/agent/extensions/config.yaml (CT 100 .24)
/etc/systemd/system/strix-server.service (amdpve .15 — strix-moe)
/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110)
@@ -220,7 +233,7 @@ When LiteLLM is upgraded to a version supporting per-key MCP grants:
- [ ] All 5 PVE nodes updated, no reboot-loop
- [ ] All VMs/CTs running post-update
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117)
- [ ] All Docker containers healthy (VM 109 + CT 116 + CT 117 + hwpve .11 + NetBird VPS)
- [ ] LiteLLM inference passing (syslog-auto test)
- [ ] Zulip server + all 3 agents connected
- [ ] GPU fleet at full capacity (3/3)
+1 -1
View File
@@ -113,7 +113,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only.
- **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`).
- **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault).
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames.
- **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM.
## Maintains
+8 -2
View File
@@ -17,8 +17,9 @@ Changelog:
v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity,
vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key
name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}).
Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114),
abiba (.24).
Fleet roster: tanko (.122), koby (.129), koonimo (.114), abiba (.24).
(v2 also carried a mumuni probe; see v5 — mumuni is no longer probed: she
moved to her own container, kagentz CT 105 / .14, and is monitored there.)
v3 (2026-09-08): GPU unit repoint verified live (.8 llama-chat-api.service,
.110 llama-server.service, .15 strix-server.service) — .8 was probing a stale
llama-server unit that reads inactive, producing false UNREACHABLE legs.
@@ -44,6 +45,11 @@ Changelog:
separate from `failures`. Every run prints absolute execution provenance
(script + cwd) in the header, in the cron ALERT line, and in --json output so
a stale-consumer report is distinguishable from a fault at read time.
v5 (2026-09-10): roster correction only, no behavior change. mumuni was removed
from the AGENTS dict when she moved off this host onto her own container
(kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored
from her side. This script must not probe mumuni or .24 — the v2 changelog
roster line was the last reference still placing her at .24 / CT100.
"""
import subprocess, json, sys, os, time, re, io, contextlib
+58 -55
View File
@@ -15,7 +15,7 @@ import smtplib, json, subprocess, os, sys, datetime, re
from email.mime.text import MIMEText
from email.mime.multipart import MIMEMultipart
PVE = "https://minipve.sysloggh.net"
PVE = "https://192.168.68.12:8006"
AUTH = "Authorization: PVEAPIToken=monitoring@pve!mumuni=eafd56c5-93d4-4d40-a41d-e688be0987f3"
# ── Shared credentials —─
@@ -38,10 +38,16 @@ TIME_STR = NOW.strftime("%Y-%m-%d %H:%M UTC")
# ── Helpers ──
def pve_get(path):
cmd = f'curl -sfk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
"""Fetch PVE API data. Returns list on success, None on error (to distinguish from empty list)."""
cmd = f'curl -sk --connect-timeout 10 "{PVE}{path}" -H "{AUTH}"'
try:
return json.loads(subprocess.check_output(cmd, shell=True))["data"]
except: return []
r = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=12)
if r.returncode != 0:
return None
data = json.loads(r.stdout)
return data.get("data", [])
except:
return None
def ssh(host, cmd):
try:
@@ -97,21 +103,33 @@ def collect():
# ── Proxmox Nodes ──
nodes = pve_get("/api2/json/nodes")
report["nodes"] = {n["node"]: {
"cpu_pct": round(n.get('cpu',0)*100, 1),
"ram": f"{n.get('mem',0)//1024//1024}/{n.get('maxmem',0)//1024//1024}MB",
"ram_pct": round(n.get('mem',0)/n.get('maxmem',1)*100, 0),
"disk": f"{n.get('disk',0)//1024//1024//1024}/{n.get('maxdisk',0)//1024//1024//1024}GB",
"disk_pct": round(n.get('disk',0)/n.get('maxdisk',1)*100, 0),
"uptime_h": n.get('uptime',0)//3600,
"status": n["status"]
} for n in nodes}
report["node_count"] = len(nodes)
report["nodes_online"] = sum(1 for n in nodes if n["status"] == "online")
if nodes is None:
report["nodes"] = {}
report["node_count"] = 0
report["nodes_online"] = 0
report["pve_probe_status"] = "unreachable"
else:
report["nodes"] = {n["node"]: {
"cpu_pct": round(n.get('cpu',0)*100, 1),
"ram": f"{n.get('mem',0)//1024//1024}/{n.get('maxmem',0)//1024//1024}MB",
"ram_pct": round(n.get('mem',0)/n.get('maxmem',1)*100, 0),
"disk": f"{n.get('disk',0)//1024//1024//1024}/{n.get('maxdisk',0)//1024//1024//1024}GB",
"disk_pct": round(n.get('disk',0)/n.get('maxdisk',1)*100, 0),
"uptime_h": n.get('uptime',0)//3600,
"status": n["status"]
} for n in nodes}
report["node_count"] = len(nodes)
report["nodes_online"] = sum(1 for n in nodes if n["status"] == "online")
report["pve_probe_status"] = "ok"
# ── VMs/CTs ──
resources = pve_get("/api2/json/cluster/resources")
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
if resources is None:
vms = []
report["resources_probe_status"] = "unreachable"
else:
vms = [r for r in resources if r.get("type") in ("qemu","lxc")]
report["resources_probe_status"] = "ok"
report["total_vms"] = len(vms)
report["running_vms"] = sum(1 for v in vms if v.get("status") == "running")
stopped = [v for v in vms if v.get("status") != "running"]
@@ -183,7 +201,7 @@ def collect():
("Authentik", "https://auth.sysloggh.net"),
("Zulip", "https://chat.sysloggh.net"),
("Pulse", "https://pulse.sysloggh.net"),
("Proxmox", "https://minipve.sysloggh.net"),
("Proxmox", "https://192.168.68.12:8006"),
("SearXNG", "http://192.168.68.7:8888"),
("Firecrawl", "http://192.168.68.7:3002/health"),
]
@@ -247,11 +265,13 @@ def collect():
zulip_health = json.loads(health_body) if health_body else {}
except:
zulip_health = {}
report["zulip_ext"]["connected"] = zulip_health.get("connected", False)
report["zulip_ext"]["queue_id"] = zulip_health.get("queue_id")
report["zulip_ext"]["last_error"] = zulip_health.get("last_error")
report["zulip_ext"]["messages_processed"] = zulip_health.get("messages_processed", 0)
report["zulip_ext"]["retry_count"] = zulip_health.get("retry_count", 0)
# Live state is nested under 'zulip' key
zulip_state = zulip_health.get("zulip", {})
report["zulip_ext"]["connected"] = zulip_state.get("connected", False)
report["zulip_ext"]["queue_id"] = zulip_state.get("queue_id")
report["zulip_ext"]["last_error"] = zulip_state.get("last_error")
report["zulip_ext"]["messages_processed"] = zulip_state.get("messages_processed", 0)
report["zulip_ext"]["skipped"] = zulip_state.get("skipped", 0)
# Phase 2: PM2 process check
pm2_raw = subprocess.check_output(
@@ -289,8 +309,8 @@ def collect():
# Abiba (pi)
report["agents"]["abiba"] = {
"platform": "pi", "ct": 100, "ip": "192.168.68.24",
"zulip_connected": zulip_health.get("connected", False),
"zulip_processed": zulip_health.get("messages_processed", 0),
"zulip_connected": zulip_state.get("connected", False),
"zulip_processed": zulip_state.get("messages_processed", 0),
"pm2_status": pm2.get("status", "unknown"),
"pm2_restarts": pm2.get("restarts", "?"),
"pm2_uptime": pm2.get("uptime", "?"),
@@ -308,26 +328,13 @@ def collect():
"updated_at": "",
}
# Mumuni (CT 100, IP 192.168.68.24)
mumuni_state = ssh("192.168.68.24", "cat ~/.hermes/gateway_state.json 2>/dev/null")
mumuni_data = {}
try:
mumuni_data = json.loads(mumuni_state) if mumuni_state else {}
except:
mumuni_data = {}
mumuni_platforms = mumuni_data.get("platforms", {})
report["agents"]["mumuni"] = {
"platform": "hermes", "ct": 100, "ip": "192.168.68.24",
"gateway_state": mumuni_data.get("gateway_state", "unknown"),
"telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"),
"zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"),
"email_state": mumuni_platforms.get("email", {}).get("state", "unknown"),
"hermes_version": "",
}
# Get Hermes version
ver = ssh("192.168.68.24", "hermes --version 2>/dev/null | head -1")
if ver:
report["agents"]["mumuni"]["hermes_version"] = ver.split("·")[0].replace("Hermes Agent ","").strip()
# Mumuni is deliberately absent from this digest: captain ruling 2026-09-10.
# She moved off this host onto her own container (kagentz CT 105 on minipve,
# 192.168.68.14, dedicated `hermes` user) and is monitored from her side. The
# former probe ssh'd to 192.168.68.24 for the decommissioned deployment's
# ~/.hermes/gateway_state.json, always read "unknown", and published a false
# "mumuni:unknown" line in the agent table and the gateway-unknown issue
# count of every digest. Do NOT re-add an .24 / gateway_state probe.
return report
@@ -419,7 +426,7 @@ th {{ color: #8b949e; font-weight: normal; }}
<div class="alert {'good' if not issues else 'bad' if any('🔴' in i for i in issues) else 'warn'}">
<p style="margin:0;font-size:16px"><b>{status}</b></p>
<p style="margin:4px 0 0 0;font-size:13px">
{r['node_count']} PVE nodes · {r['total_vms']} VMs/CTs · {r['running_vms']} running ·
Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']}) · {r['total_vms']} VMs/CTs · {r['running_vms']} running ·
{r['docker_vm']['total'] + r['docker_syslog']['total'] + r['docker_netbird']['total']} containers ·
{len(r['endpoints'])} endpoints · {len(r.get('agents',{}))} agents
</p>
@@ -435,8 +442,10 @@ th {{ color: #8b949e; font-weight: normal; }}
# ── Quick Stats ──
html += '<div class="card"><h2>📊 Quick Stats</h2><div class="grid">'
pve_status_label = "unreachable" if r.get('pve_probe_status') == 'unreachable' else f"{r['nodes_online']}/{r['node_count']}"
pve_status_color = "red" if r.get('pve_probe_status') == 'unreachable' or r['nodes_online'] != r['node_count'] else "green"
stats = [
("PVE Nodes", f"{r['nodes_online']}/{r['node_count']}", "green" if r['nodes_online'] == r['node_count'] else "red"),
("PVE Nodes", pve_status_label, pve_status_color),
("VMs/CTs", f"{r['running_vms']}/{r['total_vms']}", "green" if r['running_vms'] == r['total_vms'] else "red"),
("Containers", f"{r['docker_vm']['running']}/{r['docker_vm']['total']}", "green" if r['docker_vm']['running'] == r['docker_vm']['total'] else "yellow"),
("LiteLLM Ctrs", f"{r['docker_syslog']['running']}/{r['docker_syslog']['total']}", "green" if r['docker_syslog']['running'] == r['docker_syslog']['total'] else "red"),
@@ -539,12 +548,6 @@ th {{ color: #8b949e; font-weight: normal; }}
zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜")
gateway = agent.get("gateway_state", "?")
processed = "DSH"
elif name == "mumuni":
zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜")
gateway = agent.get("gateway_state", "?")
tg = "✅" if agent.get("telegram_state") == "connected" else "❌"
ver = agent.get("hermes_version", "")
processed = f"TG:{tg} v{ver}"
else:
zulip_state = "⬜"
gateway = agent.get("gateway_state", "?")
@@ -698,6 +701,6 @@ if __name__ == "__main__":
print(f" Zulip Ext: {'✅' if report.get('zulip_ext',{}).get('connected') else '❌'}")
print(f" LiteLLM: {sum(1 for c in report.get('litellm',{}).get('checks',[]) if c['status']=='pass')}/{len(report.get('litellm',{}).get('checks',[]))} checks pass")
agent_parts = []
for k,v in report.get('agents',{}).items():
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
print(f" Agents: {', '.join(agent_parts)}")
for k,v in report.get('agents',{}).items():
agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}")
print(f" Agents: {', '.join(agent_parts)}")
+1 -2
View File
@@ -5,6 +5,7 @@
# Field positions (awk -F'│'): $7=pid $8=uptime $9=restarts $10=status
TELEGRAM_BOT_TOKEN="$(grep TELEGRAM_BOT_TOKEN /root/.pi/agent/extensions/telegram/.env 2>/dev/null | cut -d= -f2 || echo '')"
LOG="/root/pm2-self-heal.log"
TELEGRAM_CHAT_ID="5822977936"
notify_tg() {
@@ -16,8 +17,6 @@ notify_tg() {
-d "text=${msg}" \
-d "parse_mode=HTML" > /dev/null 2>&1 || true
}
ALERTS="${ALERTS}$msg"
}
# Log-only mode: replaced by prose contract pm2-self-heal.prose.md
# Only alerts Telegram on actual failure (status != online)
+14 -19
View File
@@ -1,7 +1,10 @@
#!/bin/bash
# /root/scripts/zulip-monitor.sh — Zulip Mesh Health Monitor
# Implements zulip-health.prose.md v2
# Implements zulip-health.prose.md v3
# Runs every 15 min via cron. Alerts: Zulip private DM to the owner plus a stream post to #agent-hub on topic 'zulip-health'.
# Legs: global Zulip server, Platform A pi/Abiba (the Zulip bridge), Platform B
# Tanko (DSH), Platform C Agent Zero (kagentz). The former Platform B Hermes
# agent leg is retired — see the note after the Tanko leg.
set -euo pipefail
ZULIP_SITE="https://chat.sysloggh.net"
@@ -21,7 +24,8 @@ notify() {
# Zulip DM to owner
local content="${severity} Zulip Monitor: ${msg}"
local form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
local form
form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")"
curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \
-u "${ZULIP_EMAIL}:${ZULIP_KEY}" \
-d "${form}" > /dev/null 2>&1 || true
@@ -141,23 +145,14 @@ else
esac
fi
# ── Platform B: Hermes (Mumuni) ──
MUMUNI_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.24 \
"cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}")
MUMUNI_ZULIP=$(echo "$MUMUNI_STATE" | python3 -c "
import sys,json
d=json.load(sys.stdin)
p=d.get('platforms',{}).get('zulip',{})
print(p.get('state','unknown'))
" 2>/dev/null)
if [ "$MUMUNI_ZULIP" != "connected" ]; then
notify "🔴" "Mumuni (Hermes) Zulip state: $MUMUNI_ZULIP"
ISSUES=$((ISSUES + 1))
echo " Mumuni: ❌ state=$MUMUNI_ZULIP" >> "$LOG"
else
echo " Mumuni: ✅ Zulip connected" >> "$LOG"
fi
# ── Removed: the former "Platform B: Hermes" agent leg ──
# Captain ruling 2026-09-10: that agent moved off this host onto her own
# container (kagentz CT 105 on minipve, dedicated `hermes` user) and is now
# monitored on her side — see the out-of-scope note in zulip-health.prose.md.
# The old leg ssh'd to her former CT 100 deployment and read its Hermes gateway
# state, which reported "unknown" on every run and posted a false 🔴 DM plus an
# #agent-hub stream alert. Do NOT re-add a probe for her: this monitor must
# never contact her former host.
# ── Platform C: Agent Zero (kagentz) ──
AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \
+138
View File
@@ -0,0 +1,138 @@
"""
Regression tests for daily-infra-report.py fixes (PR #64).
Tests:
(a) Asserts the nested zulip read feeds the agent-card fields
(b) Asserts an unreachable pve_get renders labelled-unreachable, not "0/0"
"""
import json
import subprocess
import sys
from pathlib import Path
from unittest.mock import patch, MagicMock
# Add scripts to path
sys.path.insert(0, str(Path(__file__).parent.parent / "scripts"))
import importlib.util
def load_script():
"""Load the daily-infra-report script as a module."""
script_path = Path(__file__).parent.parent / "scripts" / "daily-infra-report.py"
spec = importlib.util.spec_from_file_location("daily_infra_report", script_path)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
def test_nested_zulip_read_feeds_agent_card():
"""Test that Zulip state is read from the nested 'zulip' key and feeds agent-card fields."""
# Mock the http_get_body response with nested structure
mock_health_response = json.dumps({
"status": "ok",
"platform": "pi",
"agent": "abiba",
"zulip": {
"connected": True,
"queue_id": "test-queue-id",
"messages_processed": 42,
"skipped": 5,
"last_error": None
}
})
# Import and patch
report_mod = load_script()
with patch.object(report_mod, 'http_get_body', return_value=mock_health_response):
# Simulate the collect() function's Zulip section
zulip_health = json.loads(report_mod.http_get_body("http://localhost:9200/health"))
zulip_state = zulip_health.get("zulip", {})
# Assert the nested key is read correctly
assert zulip_state.get("connected") == True, "Zulip connected should be True from nested key"
assert zulip_state.get("messages_processed") == 42, "messages_processed should be 42 from nested key"
assert zulip_state.get("queue_id") == "test-queue-id", "queue_id should be read from nested key"
# Simulate the agent card field population
agent_card = {
"zulip_connected": zulip_state.get("connected", False),
"zulip_processed": zulip_state.get("messages_processed", 0),
}
assert agent_card["zulip_connected"] == True, "Agent card should show Zulip connected"
assert agent_card["zulip_processed"] == 42, "Agent card should show 42 processed messages"
def test_unreachable_pve_get_renders_labelled_unreachable():
"""Test that an unreachable PVE API renders 'unreachable' instead of '0/0'."""
# Import and patch
report_mod = load_script()
# Test pve_get returns None on error
with patch.object(report_mod.subprocess, 'run') as mock_run:
mock_run.return_value.returncode = 7 # Connection failure
result = report_mod.pve_get("/api2/json/nodes")
assert result is None, "pve_get should return None on connection failure"
# Test the render logic
report = {
"nodes": {},
"node_count": 0,
"nodes_online": 0,
"pve_probe_status": "unreachable",
"total_vms": 0,
"running_vms": 0,
}
# The render should show "unreachable" not "0/0"
pve_status_label = "unreachable" if report.get('pve_probe_status') == 'unreachable' else f"{report['nodes_online']}/{report['node_count']}"
assert pve_status_label == "unreachable", "PVE status should show 'unreachable' when probe fails, not '0/0'"
def test_unreachable_resources_renders_labelled_unreachable():
"""Test that unreachable resources probe renders 'unreachable' instead of '0/0'."""
report_mod = load_script()
# Test resources probe returns None
with patch.object(report_mod.subprocess, 'run') as mock_run:
mock_run.return_value.returncode = 7
result = report_mod.pve_get("/api2/json/cluster/resources")
assert result is None, "pve_get for resources should return None on connection failure"
# Test the render logic
report = {
"resources_probe_status": "unreachable",
"total_vms": 0,
"running_vms": 0,
}
resources_label = "unreachable" if report.get('resources_probe_status') == 'unreachable' else f"{report['running_vms']}/{report['total_vms']}"
assert resources_label == "unreachable", "Resources status should show 'unreachable' when probe fails, not '0/0'"
if __name__ == "__main__":
print("Running tests...")
try:
test_nested_zulip_read_feeds_agent_card()
print("✓ test_nested_zulip_read_feeds_agent_card passed")
except AssertionError as e:
print(f"✗ test_nested_zulip_read_feeds_agent_card failed: {e}")
sys.exit(1)
try:
test_unreachable_pve_get_renders_labelled_unreachable()
print("✓ test_unreachable_pve_get_renders_labelled_unreachable passed")
except AssertionError as e:
print(f"✗ test_unreachable_pve_get_renders_labelled_unreachable failed: {e}")
sys.exit(1)
try:
test_unreachable_resources_renders_labelled_unreachable()
print("✓ test_unreachable_resources_renders_labelled_unreachable passed")
except AssertionError as e:
print(f"✗ test_unreachable_resources_renders_labelled_unreachable failed: {e}")
sys.exit(1)
print("All tests passed!")
+352
View File
@@ -0,0 +1,352 @@
"""Regression tests for the 2026-09-10 retirement of the Mumuni monitoring leg.
WHY THIS FILE EXISTS: captain ruling 2026-09-10 — Mumuni moved off this host
onto her own container (kagentz CT 105 on minipve, 192.168.68.14, dedicated
`hermes` user) and is monitored from her side. The monitor nevertheless kept
ssh'ing to root@192.168.68.24 for `~/.hermes/gateway_state.json` on the
decommissioned deployment, read "unknown" on every run, and posted a false 🔴
"Mumuni (Hermes) Zulip state: unknown" DM + #agent-hub stream alert to the
captain. The daily infra digest published a matching `mumuni:unknown` row.
CONTRACT UNDER TEST:
* `scripts/zulip-monitor.sh` carries NO Mumuni probe and NO 192.168.68.24
reference; it never ssh'es .24, and even on a failing run it emits no Mumuni
notify (stdout alert, Zulip payload, or log line).
* The Abiba (pi — the Zulip bridge), Tanko (DSH) and Agent Zero (kagentz) legs
still work: deleting the Mumuni leg must not have gutted the rest.
* `scripts/daily-infra-report.py` no longer probes .24 for a Hermes gateway
state and no longer emits a `mumuni` agent entry.
* `scripts/agent-health-check.py`'s AGENTS roster has no mumuni entry. This is
a pin, not a behavior change — verify the probe was already gone.
* `zulip-health.prose.md` retires the Mumuni-only steps and says explicitly
that Mumuni is not monitored from this host.
HOW: behavioral execution plus one named deliverable-text contract. The sandbox
copies the shipped monitor verbatim and rewrites only its LOG constant, then
runs it with stub ssh/curl on PATH; the ssh stub records every host it is asked
to reach, so "never probes .24" and "no Mumuni notify" are asserted from
observed behavior. The daily digest is pinned by importing it and exercising
collect() and build_html() directly. The single source-text assertion is the
deliverable-text contract the captain acceptance names for the shipped monitor.
Usage: python3 -m pytest tests/test_mumuni_monitor_removal.py
"""
from __future__ import annotations
import importlib.util
import os
import pathlib
import stat
import subprocess
import pytest
ROOT = pathlib.Path(__file__).resolve().parents[1]
ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh"
DAILY_REPORT = ROOT / "scripts" / "daily-infra-report.py"
AHC = ROOT / "scripts" / "agent-health-check.py"
HEALTH_CONTRACT = ROOT / "zulip-health.prose.md"
CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json"
MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment
TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec
AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker
# ── scripts/zulip-monitor.sh: deliverable-text contract ─────────────
def test_zulip_monitor_deliverable_text_contract():
"""Owned deliverable-text contract for scripts/zulip-monitor.sh.
Captain acceptance requires the shipped monitor to contain no Mumuni probe
identifier and no 192.168.68.24 literal. Behavioral proof that the monitor
never contacts that host and never emits a Mumuni notify lives in the
sandbox tests below; this only pins the named text contract.
"""
text = ZULIP_MONITOR.read_text()
assert "mumuni" not in text.lower()
assert MUMUNI_IP not in text
# ── scripts/zulip-monitor.sh: behavioral sandbox ─────────────────────
SSH_STUB = r"""#!/usr/bin/env bash
# Stub ssh: record the target host, then answer by host + remote command.
printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls"
host=""
for a in "$@"; do
case "$a" in
*@192.168.*) host="${a##*@}" ;;
esac
done
printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts"
cmd="${*: -1}"
case "$host" in
192.168.68.15)
case "$cmd" in
*"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;;
*curl*) printf '%s' "$TANKO_HTTP" ;;
esac ;;
192.168.68.14)
case "$cmd" in
*agent.json*) printf '%s' "$AZ_A2A" ;;
*"ps aux"*) printf '%s\n' "$AZ_PS" ;;
esac ;;
*)
printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;;
esac
exit 0
"""
CURL_STUB = r"""#!/usr/bin/env bash
# Stub curl: serve the Abiba health fixture and the Zulip server 200, and
# record every call (including notify) payloads.
printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls"
case "$*" in
*:9200/health*)
case " $* " in
*" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe
*) printf '%s' "$PI_BODY" ;; # body probe
esac ;;
*server_settings*)
printf '%s' "$SERVER_HTTP" ;;
esac
exit 0
"""
def _write_exec(path: pathlib.Path, body: str) -> None:
path.write_text(body)
path.chmod(path.stat().st_mode
| stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200",
az_a2a='{"name":"kagentz"}',
az_ps="root 111 0.1 0.2 /opt/venv-a0/bin/python3 -u adapter.py"):
"""Run the shipped monitor in a sandbox; return (proc, record_dir, log_path).
Only the LOG constant is rewritten (to keep the run inside the worktree).
Everything else — legs, labels, notify logic — is the shipped script.
"""
sandbox = tmp_path / "sandbox"
bindir = sandbox / "bin"
record = sandbox / "record"
bindir.mkdir(parents=True)
record.mkdir()
_write_exec(bindir / "ssh", SSH_STUB)
_write_exec(bindir / "curl", CURL_STUB)
source = ZULIP_MONITOR.read_text()
log_line = 'LOG="/root/zulip-health-monitor.log"'
assert log_line in source, "LOG constant moved — update the sandbox harness"
log_path = sandbox / "zulip-health-monitor.log"
script = sandbox / "zulip-monitor.sh"
script.write_text(source.replace(log_line, f'LOG="{log_path}"'))
env = dict(os.environ)
env.update({
"PATH": f"{bindir}:{env['PATH']}",
"RECORD_DIR": str(record),
"TANKO_SVC": tanko_svc,
"TANKO_HTTP": tanko_http,
"AZ_A2A": az_a2a,
"AZ_PS": az_ps,
"PI_HTTP": "200",
"PI_BODY": CONNECTED_FIXTURE.read_text(),
"SERVER_HTTP": "200",
})
proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env,
capture_output=True, text=True)
return proc, record, log_path
def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path):
proc, record, log_path = _run_monitor(tmp_path)
assert proc.returncode == 0, proc.stderr
log = log_path.read_text()
# Every retained leg actually ran and passed.
assert "Server: ✅ HTTP 200" in log
assert "Abiba: ✅ Connected" in log
assert "Tanko: ✅ service=active http=200" in log
assert "kagentz: ✅ A2A alive" in log
assert "kagentz: ✅ Adapter running" in log
assert "Result: ✅ All healthy" in log
# A healthy run emits no notify at all — and certainly no Mumuni one.
assert proc.stdout == ""
assert "Mumuni" not in log
assert "🔴" not in log
# Observed behavior: .24 is never resolved, only Tanko's vantage and the
# Agent Zero host are contacted.
hosts = record.joinpath("ssh.hosts").read_text().split()
assert MUMUNI_IP not in hosts
assert set(hosts) == {TANKO_VANTAGE, AGENT_ZERO_HOST}
assert not record.joinpath("unexpected-ssh").exists()
def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path):
# Failure path: exercises notify() end to end so "no Mumuni notify" is
# proven on the alert path, not only on the quiet healthy path.
proc, record, log_path = _run_monitor(tmp_path, tanko_svc="inactive",
tanko_http="000")
assert proc.returncode == 0, proc.stderr
alerts = proc.stdout
assert "Tanko (DSH dsh-web) service state: inactive" in alerts
assert "1 issue(s) found" in alerts
# No Mumuni text in stdout, the log, or any Zulip DM/stream payload.
assert "Mumuni" not in alerts
assert "Mumuni" not in log_path.read_text()
assert MUMUNI_IP not in alerts + log_path.read_text()
payloads = record.joinpath("curl.calls").read_text()
assert "Mumuni" not in payloads
assert MUMUNI_IP not in payloads
# The rest of the monitor still ran alongside the failing Tanko leg.
log = log_path.read_text()
assert "Abiba: ✅ Connected" in log
assert "kagentz: ✅ A2A alive" in log
assert "Result: 🔴 1 issue(s) found" in log
# ── scripts/daily-infra-report.py: behavioral digest checks ──────────
@pytest.fixture(scope="module")
def daily():
spec = importlib.util.spec_from_file_location("daily_infra_report", DAILY_REPORT)
assert spec and spec.loader
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
DAILY_AGENTS = {
"abiba": {
"platform": "pi", "ct": 100, "ip": MUMUNI_IP,
"zulip_connected": True, "zulip_processed": 5,
"pm2_status": "online", "pm2_restarts": "0", "pm2_uptime": "1h",
},
"tanko": {
"platform": "dsh", "ct": 112, "ip": "192.168.68.122",
"gateway_state": "n/a (DSH)", "zulip_state": "connected",
"telegram_state": "unknown", "gateway_pid": None, "updated_at": "",
},
}
def _fabricated_report(agents):
return {
"nodes": {},
"node_count": 1,
"nodes_online": 1,
"total_vms": 0,
"running_vms": 0,
"stopped_vms": [],
"vms_by_node": {n: [] for n in
["amdpve", "minipve", "storepve", "acerpve", "ocupve"]},
"storage": [],
"docker_vm": {"total": 0, "running": 0, "unhealthy": [],
"containers": [], "reclaimable": "", "disk_used": "1%"},
"docker_syslog": {"total": 0, "running": 0, "containers": []},
"docker_netbird": {"total": 0, "running": 0, "containers": []},
"endpoints": [],
"litellm": {"checks": []},
"nfs": [],
"zulip_ext": {
"connected": True, "queue_id": "queue", "last_error": None,
"messages_processed": 0, "retry_count": 0, "pm2": {},
"pm2_healthy": True, "bot_skipped_15min": 0, "finalized_1h": 0,
"failed_finalize_1h": 0, "finalize_fail_pct": 0,
"server_status": "200",
},
"agents": agents,
}
def _agent_status_card(html):
start = html.index("🤖 Agent Status")
end = html.index("💬 Zulip Extension")
return html[start:end]
def test_daily_report_renders_only_abiba_and_tanko_agents(daily):
"""build_html() over a Mumuni-free agent set must render no Mumuni row and
no Mumuni gateway-unknown issue, while abiba and tanko rows still render."""
html = daily.build_html(_fabricated_report(dict(DAILY_AGENTS)))
card = _agent_status_card(html)
assert "mumuni" not in card.lower()
assert "abiba" in card
assert "tanko" in card
assert "mumuni" not in html.lower()
def test_daily_report_collect_never_probes_mumuni(monkeypatch, daily):
"""collect() with ssh stubbed must add no mumuni agent and must never ssh
its decommissioned .24 host."""
probed = []
class _NoSubprocess:
@staticmethod
def check_output(*args, **kwargs):
return b""
def fake_ssh(host, cmd):
probed.append(host)
return ""
monkeypatch.setattr(daily, "pve_get", lambda path: [])
monkeypatch.setattr(daily, "ssh_jerome", lambda host, cmd: "")
monkeypatch.setattr(daily, "ssh", fake_ssh)
monkeypatch.setattr(daily, "http_get",
lambda url, auth=None, timeout=10: "200")
monkeypatch.setattr(daily, "http_get_body",
lambda url, auth=None, timeout=10: "")
monkeypatch.setattr(daily, "count_in_log", lambda *a, **k: 0)
monkeypatch.setattr(daily, "subprocess", _NoSubprocess)
report = daily.collect()
assert "mumuni" not in report["agents"]
assert MUMUNI_IP not in probed
# ── scripts/agent-health-check.py: roster pin ───────────────────────
@pytest.fixture(scope="module")
def ahc():
spec = importlib.util.spec_from_file_location("agent_health_check_roster", AHC)
assert spec and spec.loader
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
def test_agent_health_roster_has_no_mumuni_entry(ahc):
assert "mumuni" not in ahc.AGENTS
# ── zulip-health.prose.md: contract reconciliation ──────────────────
def test_health_contract_retires_mumuni_only_steps():
text = HEALTH_CONTRACT.read_text()
assert MUMUNI_IP not in text
for step in ("**B4:", "**B5:", "**B6:"):
assert step not in text
def test_health_contract_states_mumuni_is_not_monitored_from_this_host():
text = HEALTH_CONTRACT.read_text()
assert "Mumuni is NOT monitored from this host" in text
assert "monitored on her side" in text
assert "her own container" in text
def test_health_contract_keeps_tanko_agent_zero_and_bridge_steps():
text = HEALTH_CONTRACT.read_text()
for marker in ("**B1:", "**B2:", "**B3:", "Step 4: Platform C",
"Step 2: Platform A", "Step 1: Zulip Server Liveness"):
assert marker in text, marker
+25 -43
View File
@@ -1,9 +1,9 @@
---
kind: responsibility
name: zulip-health
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Tanko on DSH / Mumuni on Hermes), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity.
description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side.
title: Zulip Mesh Health Monitor — Multi-Platform
version: 3.0.0
version: 3.1.0
runtime_contract: 2
agent: abiba
report_only_agents:
@@ -12,13 +12,23 @@ report_only_agents:
# Zulip Mesh Health Monitor
Monitors ALL Zulip-connected agents across platforms (pi, Hermes, DSH, Agent Zero).
Runs every 15 minutes in the background. Also triggers on session start.
Monitors the Zulip-connected agents under this host's operational control (pi,
DSH, Agent Zero). Runs every 15 minutes in the background. Also triggers on
session start.
> **Mumuni is NOT monitored from this host (captain ruling 2026-09-10).** She
> moved off this host onto her own container — kagentz CT 105 on minipve
> (192.168.68.14), running a dedicated `hermes` user — and is monitored on her
> side. No step in this contract, and no leg of `scripts/zulip-monitor.sh`, may
> ssh to her old deployment, read her `~/.hermes/gateway_state.json`, or alert on
> her state. The former Platform-B-for-Mumuni steps (gateway process, heartbeat,
> response delivery) are retired: they always read "unknown" against the
> decommissioned deployment and produced a false 🔴 alert on every run.
## Requires
- **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY`
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); Mumuni (192.168.68.14, kagentz CT105 on minipve); and Agent Zero Docker host (192.168.68.14)
- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14)
- **PM2** on localhost for pi process management
- **Network access** to `chat.sysloggh.net`, `localhost:9200`
- **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce`
@@ -81,13 +91,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse
## Streaming Support (2026-07-05)
Zulip agents now support progressive message editing during agent generation.
When a Zulip agent (Tanko on DSH, Mumuni on Hermes) processes a message, the response is
When a Zulip agent under this monitor's scope (Tanko on DSH) processes a message, the response is
streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API:
- Adapter implements `edit_message()` using `_api_patch()` helper
- Gateway stream consumer progressively edits the Zulip message
- User sees real-time agent thinking instead of waiting for full response
- Verified: Tanko (CT 112) and Mumuni (kagentz CT 105) both have streaming active
- Verified: Tanko (CT 112) has streaming active; Mumuni's (kagentz CT 105) is verified on her own host, not from here
### Verification
```bash
@@ -183,7 +193,10 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta
| Crash loop >10/h | Alert user |
### Step 3: Platform B — Tanko (DSH on amdpve CT 112) & Mumuni (Hermes)
### Step 3: Platform B — Tanko (DSH on amdpve CT 112)
Mumuni is out of scope for this host (see the note above): she runs on her own
container and is monitored on her side.
Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so
there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs
@@ -240,44 +253,13 @@ timeout only. Never expect a bare `200` — the public URL terminates in the
token-gated authentik chain. Statuses outside the healthy set are
logged/reported as a warning — reported, never healed on.
**B4: Gateway Process** (Hermes agent Mumuni only — Tanko runs no Hermes gateway)
```bash
ssh root@<CT> "ps aux | grep 'gateway run' | grep -v grep"
```
Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more
than one `gateway run` process is found, the gateway has a collision (typically
one `--force` and one `--replace` process). Kill the newer/duplicate process,
then restart the remaining gateway per-agent (parameterized 2026-08-09, captain
ruling). Check the gateway log for "Gateway running with 2 platform(s)" (not 1)
to confirm Zulip reloaded.
**B5: Heartbeat Verification** (Hermes agent Mumuni only — Tanko has no Hermes gateway)
```bash
ssh root@192.168.68.24 "grep Heartbeat ~/.hermes/logs/agent.log | tail -3"
```
Expected: recent heartbeat (within 5 min), `polls=N` incrementing.
Silence > 300s → warning. Silence > 600s → critical.
**B6: Response Delivery** (Hermes agent Mumuni only)
```bash
ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10"
```
> 50% fail rate → critical.
**Platform B Actions**
| Condition | Action |
|-----------|--------|
| `zulip.state != "connected"` | `ssh root@<CT> "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` (Mumuni) / restart Tanko via DSH service |
| No heartbeat in 10min | Same as above |
| `Failed to finalize` > 50% | Check PATCH API, Zulip server |
| Response empty/short | Check A2A endpoint / LiteLLM model |
| `dsh-web` service not `active` | Restart Tanko via DSH service |
| HTTP `:3080` connection refused/timeout (`000`) | Same as above |
| HTTP status outside the expected set | Log/report as a warning — reported, never healed on |
### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14)
@@ -333,7 +315,7 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`. I
Check each agent's log for excessive bot-to-bot chatter:
- Abiba: `Skipped.*bot msgs` count
- Tanko/Mumuni: Repeated DM exchanges between bots
- Tanko: Repeated DM exchanges between bots
- kagentz: Adapter log for bot DMs being processed
If any bot processes >50 bot-originated messages in 15min → warning.