diff --git a/contract-registry.yaml b/contract-registry.yaml index 2553260..fb50e4a 100644 --- a/contract-registry.yaml +++ b/contract-registry.yaml @@ -628,18 +628,18 @@ contracts: sensitivity: high status: active owner: abiba - version: 3.0.0 + version: 3.1.0 trigger: type: scheduled cadence: '*/15 * * * *' - description: "Every 15 minutes \u2014 monitors all Zulip-connected agents" + description: "Every 15 minutes \u2014 monitors the Zulip-connected agents under this host's control (pi, DSH, Agent Zero)" cron_job_id: null execution: agent: abiba timeout: 120 requires: - Zulip API key for abiba-bot@chat.sysloggh.net - - SSH access to all Hermes agents + - SSH access to amdpve (192.168.68.15) for Tanko (CT 112) and the Agent Zero Docker host (.14) verification: postconditions: - check: bot registration active diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 624d3c7..734c26a 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -113,7 +113,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only. - **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`). - **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault). -- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames. +- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames. - **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM. ## Maintains diff --git a/scripts/agent-health-check.py b/scripts/agent-health-check.py index 37966ed..c5d2618 100755 --- a/scripts/agent-health-check.py +++ b/scripts/agent-health-check.py @@ -17,8 +17,9 @@ Changelog: v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity, vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}). - Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), - abiba (.24). + Fleet roster: tanko (.122), koby (.129), koonimo (.114), abiba (.24). + (v2 also carried a mumuni probe; see v5 — mumuni is no longer probed: she + moved to her own container, kagentz CT 105 / .14, and is monitored there.) v3 (2026-09-08): GPU unit repoint verified live (.8 llama-chat-api.service, .110 llama-server.service, .15 strix-server.service) — .8 was probing a stale llama-server unit that reads inactive, producing false UNREACHABLE legs. @@ -44,6 +45,11 @@ Changelog: separate from `failures`. Every run prints absolute execution provenance (script + cwd) in the header, in the cron ALERT line, and in --json output so a stale-consumer report is distinguishable from a fault at read time. + v5 (2026-09-10): roster correction only, no behavior change. mumuni was removed + from the AGENTS dict when she moved off this host onto her own container + (kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored + from her side. This script must not probe mumuni or .24 — the v2 changelog + roster line was the last reference still placing her at .24 / CT100. """ import subprocess, json, sys, os, time, re, io, contextlib diff --git a/scripts/daily-infra-report.py b/scripts/daily-infra-report.py index d387c6b..08f149e 100755 --- a/scripts/daily-infra-report.py +++ b/scripts/daily-infra-report.py @@ -308,26 +308,13 @@ def collect(): "updated_at": "", } - # Mumuni (CT 100, IP 192.168.68.24) - mumuni_state = ssh("192.168.68.24", "cat ~/.hermes/gateway_state.json 2>/dev/null") - mumuni_data = {} - try: - mumuni_data = json.loads(mumuni_state) if mumuni_state else {} - except: - mumuni_data = {} - mumuni_platforms = mumuni_data.get("platforms", {}) - report["agents"]["mumuni"] = { - "platform": "hermes", "ct": 100, "ip": "192.168.68.24", - "gateway_state": mumuni_data.get("gateway_state", "unknown"), - "telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"), - "zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"), - "email_state": mumuni_platforms.get("email", {}).get("state", "unknown"), - "hermes_version": "", - } - # Get Hermes version - ver = ssh("192.168.68.24", "hermes --version 2>/dev/null | head -1") - if ver: - report["agents"]["mumuni"]["hermes_version"] = ver.split("·")[0].replace("Hermes Agent ","").strip() + # Mumuni is deliberately absent from this digest: captain ruling 2026-09-10. + # She moved off this host onto her own container (kagentz CT 105 on minipve, + # 192.168.68.14, dedicated `hermes` user) and is monitored from her side. The + # former probe ssh'd to 192.168.68.24 for the decommissioned deployment's + # ~/.hermes/gateway_state.json, always read "unknown", and published a false + # "mumuni:unknown" line in the agent table and the gateway-unknown issue + # count of every digest. Do NOT re-add an .24 / gateway_state probe. return report @@ -539,12 +526,6 @@ th {{ color: #8b949e; font-weight: normal; }} zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜") gateway = agent.get("gateway_state", "?") processed = "DSH" - elif name == "mumuni": - zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜") - gateway = agent.get("gateway_state", "?") - tg = "✅" if agent.get("telegram_state") == "connected" else "❌" - ver = agent.get("hermes_version", "") - processed = f"TG:{tg} v{ver}" else: zulip_state = "⬜" gateway = agent.get("gateway_state", "?") @@ -698,6 +679,6 @@ if __name__ == "__main__": print(f" Zulip Ext: {'✅' if report.get('zulip_ext',{}).get('connected') else '❌'}") print(f" LiteLLM: {sum(1 for c in report.get('litellm',{}).get('checks',[]) if c['status']=='pass')}/{len(report.get('litellm',{}).get('checks',[]))} checks pass") agent_parts = [] -for k,v in report.get('agents',{}).items(): - agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}") -print(f" Agents: {', '.join(agent_parts)}") + for k,v in report.get('agents',{}).items(): + agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}") + print(f" Agents: {', '.join(agent_parts)}") diff --git a/scripts/zulip-monitor.sh b/scripts/zulip-monitor.sh index 812f039..32b884a 100755 --- a/scripts/zulip-monitor.sh +++ b/scripts/zulip-monitor.sh @@ -1,7 +1,10 @@ #!/bin/bash # /root/scripts/zulip-monitor.sh — Zulip Mesh Health Monitor -# Implements zulip-health.prose.md v2 +# Implements zulip-health.prose.md v3 # Runs every 15 min via cron. Alerts: Zulip private DM to the owner plus a stream post to #agent-hub on topic 'zulip-health'. +# Legs: global Zulip server, Platform A pi/Abiba (the Zulip bridge), Platform B +# Tanko (DSH), Platform C Agent Zero (kagentz). The former Platform B Hermes +# agent leg is retired — see the note after the Tanko leg. set -euo pipefail ZULIP_SITE="https://chat.sysloggh.net" @@ -21,7 +24,8 @@ notify() { # Zulip DM to owner local content="${severity} Zulip Monitor: ${msg}" - local form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")" + local form + form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")" curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \ -u "${ZULIP_EMAIL}:${ZULIP_KEY}" \ -d "${form}" > /dev/null 2>&1 || true @@ -141,23 +145,14 @@ else esac fi -# ── Platform B: Hermes (Mumuni) ── -MUMUNI_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.24 \ - "cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}") -MUMUNI_ZULIP=$(echo "$MUMUNI_STATE" | python3 -c " -import sys,json -d=json.load(sys.stdin) -p=d.get('platforms',{}).get('zulip',{}) -print(p.get('state','unknown')) -" 2>/dev/null) - -if [ "$MUMUNI_ZULIP" != "connected" ]; then - notify "🔴" "Mumuni (Hermes) Zulip state: $MUMUNI_ZULIP" - ISSUES=$((ISSUES + 1)) - echo " Mumuni: ❌ state=$MUMUNI_ZULIP" >> "$LOG" -else - echo " Mumuni: ✅ Zulip connected" >> "$LOG" -fi +# ── Removed: the former "Platform B: Hermes" agent leg ── +# Captain ruling 2026-09-10: that agent moved off this host onto her own +# container (kagentz CT 105 on minipve, dedicated `hermes` user) and is now +# monitored on her side — see the out-of-scope note in zulip-health.prose.md. +# The old leg ssh'd to her former CT 100 deployment and read its Hermes gateway +# state, which reported "unknown" on every run and posted a false 🔴 DM plus an +# #agent-hub stream alert. Do NOT re-add a probe for her: this monitor must +# never contact her former host. # ── Platform C: Agent Zero (kagentz) ── AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \ diff --git a/tests/test_mumuni_monitor_removal.py b/tests/test_mumuni_monitor_removal.py new file mode 100644 index 0000000..5f9b9eb --- /dev/null +++ b/tests/test_mumuni_monitor_removal.py @@ -0,0 +1,352 @@ +"""Regression tests for the 2026-09-10 retirement of the Mumuni monitoring leg. + +WHY THIS FILE EXISTS: captain ruling 2026-09-10 — Mumuni moved off this host +onto her own container (kagentz CT 105 on minipve, 192.168.68.14, dedicated +`hermes` user) and is monitored from her side. The monitor nevertheless kept +ssh'ing to root@192.168.68.24 for `~/.hermes/gateway_state.json` on the +decommissioned deployment, read "unknown" on every run, and posted a false 🔴 +"Mumuni (Hermes) Zulip state: unknown" DM + #agent-hub stream alert to the +captain. The daily infra digest published a matching `mumuni:unknown` row. + +CONTRACT UNDER TEST: + * `scripts/zulip-monitor.sh` carries NO Mumuni probe and NO 192.168.68.24 + reference; it never ssh'es .24, and even on a failing run it emits no Mumuni + notify (stdout alert, Zulip payload, or log line). + * The Abiba (pi — the Zulip bridge), Tanko (DSH) and Agent Zero (kagentz) legs + still work: deleting the Mumuni leg must not have gutted the rest. + * `scripts/daily-infra-report.py` no longer probes .24 for a Hermes gateway + state and no longer emits a `mumuni` agent entry. + * `scripts/agent-health-check.py`'s AGENTS roster has no mumuni entry. This is + a pin, not a behavior change — verify the probe was already gone. + * `zulip-health.prose.md` retires the Mumuni-only steps and says explicitly + that Mumuni is not monitored from this host. + +HOW: behavioral execution plus one named deliverable-text contract. The sandbox +copies the shipped monitor verbatim and rewrites only its LOG constant, then +runs it with stub ssh/curl on PATH; the ssh stub records every host it is asked +to reach, so "never probes .24" and "no Mumuni notify" are asserted from +observed behavior. The daily digest is pinned by importing it and exercising +collect() and build_html() directly. The single source-text assertion is the +deliverable-text contract the captain acceptance names for the shipped monitor. + +Usage: python3 -m pytest tests/test_mumuni_monitor_removal.py +""" +from __future__ import annotations + +import importlib.util +import os +import pathlib +import stat +import subprocess + +import pytest + +ROOT = pathlib.Path(__file__).resolve().parents[1] +ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh" +DAILY_REPORT = ROOT / "scripts" / "daily-infra-report.py" +AHC = ROOT / "scripts" / "agent-health-check.py" +HEALTH_CONTRACT = ROOT / "zulip-health.prose.md" +CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json" + +MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment +TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec +AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker + + +# ── scripts/zulip-monitor.sh: deliverable-text contract ───────────── + +def test_zulip_monitor_deliverable_text_contract(): + """Owned deliverable-text contract for scripts/zulip-monitor.sh. + + Captain acceptance requires the shipped monitor to contain no Mumuni probe + identifier and no 192.168.68.24 literal. Behavioral proof that the monitor + never contacts that host and never emits a Mumuni notify lives in the + sandbox tests below; this only pins the named text contract. + """ + text = ZULIP_MONITOR.read_text() + assert "mumuni" not in text.lower() + assert MUMUNI_IP not in text + + +# ── scripts/zulip-monitor.sh: behavioral sandbox ───────────────────── + +SSH_STUB = r"""#!/usr/bin/env bash +# Stub ssh: record the target host, then answer by host + remote command. +printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls" +host="" +for a in "$@"; do + case "$a" in + *@192.168.*) host="${a##*@}" ;; + esac +done +printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts" +cmd="${*: -1}" +case "$host" in + 192.168.68.15) + case "$cmd" in + *"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;; + *curl*) printf '%s' "$TANKO_HTTP" ;; + esac ;; + 192.168.68.14) + case "$cmd" in + *agent.json*) printf '%s' "$AZ_A2A" ;; + *"ps aux"*) printf '%s\n' "$AZ_PS" ;; + esac ;; + *) + printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;; +esac +exit 0 +""" + +CURL_STUB = r"""#!/usr/bin/env bash +# Stub curl: serve the Abiba health fixture and the Zulip server 200, and +# record every call (including notify) payloads. +printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls" +case "$*" in + *:9200/health*) + case " $* " in + *" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe + *) printf '%s' "$PI_BODY" ;; # body probe + esac ;; + *server_settings*) + printf '%s' "$SERVER_HTTP" ;; +esac +exit 0 +""" + + +def _write_exec(path: pathlib.Path, body: str) -> None: + path.write_text(body) + path.chmod(path.stat().st_mode + | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH) + + +def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200", + az_a2a='{"name":"kagentz"}', + az_ps="root 111 0.1 0.2 /opt/venv-a0/bin/python3 -u adapter.py"): + """Run the shipped monitor in a sandbox; return (proc, record_dir, log_path). + + Only the LOG constant is rewritten (to keep the run inside the worktree). + Everything else — legs, labels, notify logic — is the shipped script. + """ + sandbox = tmp_path / "sandbox" + bindir = sandbox / "bin" + record = sandbox / "record" + bindir.mkdir(parents=True) + record.mkdir() + + _write_exec(bindir / "ssh", SSH_STUB) + _write_exec(bindir / "curl", CURL_STUB) + + source = ZULIP_MONITOR.read_text() + log_line = 'LOG="/root/zulip-health-monitor.log"' + assert log_line in source, "LOG constant moved — update the sandbox harness" + log_path = sandbox / "zulip-health-monitor.log" + script = sandbox / "zulip-monitor.sh" + script.write_text(source.replace(log_line, f'LOG="{log_path}"')) + + env = dict(os.environ) + env.update({ + "PATH": f"{bindir}:{env['PATH']}", + "RECORD_DIR": str(record), + "TANKO_SVC": tanko_svc, + "TANKO_HTTP": tanko_http, + "AZ_A2A": az_a2a, + "AZ_PS": az_ps, + "PI_HTTP": "200", + "PI_BODY": CONNECTED_FIXTURE.read_text(), + "SERVER_HTTP": "200", + }) + proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env, + capture_output=True, text=True) + return proc, record, log_path + + +def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path): + proc, record, log_path = _run_monitor(tmp_path) + assert proc.returncode == 0, proc.stderr + log = log_path.read_text() + + # Every retained leg actually ran and passed. + assert "Server: ✅ HTTP 200" in log + assert "Abiba: ✅ Connected" in log + assert "Tanko: ✅ service=active http=200" in log + assert "kagentz: ✅ A2A alive" in log + assert "kagentz: ✅ Adapter running" in log + assert "Result: ✅ All healthy" in log + + # A healthy run emits no notify at all — and certainly no Mumuni one. + assert proc.stdout == "" + assert "Mumuni" not in log + assert "🔴" not in log + + # Observed behavior: .24 is never resolved, only Tanko's vantage and the + # Agent Zero host are contacted. + hosts = record.joinpath("ssh.hosts").read_text().split() + assert MUMUNI_IP not in hosts + assert set(hosts) == {TANKO_VANTAGE, AGENT_ZERO_HOST} + assert not record.joinpath("unexpected-ssh").exists() + + +def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path): + # Failure path: exercises notify() end to end so "no Mumuni notify" is + # proven on the alert path, not only on the quiet healthy path. + proc, record, log_path = _run_monitor(tmp_path, tanko_svc="inactive", + tanko_http="000") + assert proc.returncode == 0, proc.stderr + + alerts = proc.stdout + assert "Tanko (DSH dsh-web) service state: inactive" in alerts + assert "1 issue(s) found" in alerts + + # No Mumuni text in stdout, the log, or any Zulip DM/stream payload. + assert "Mumuni" not in alerts + assert "Mumuni" not in log_path.read_text() + assert MUMUNI_IP not in alerts + log_path.read_text() + payloads = record.joinpath("curl.calls").read_text() + assert "Mumuni" not in payloads + assert MUMUNI_IP not in payloads + + # The rest of the monitor still ran alongside the failing Tanko leg. + log = log_path.read_text() + assert "Abiba: ✅ Connected" in log + assert "kagentz: ✅ A2A alive" in log + assert "Result: 🔴 1 issue(s) found" in log + + +# ── scripts/daily-infra-report.py: behavioral digest checks ────────── + +@pytest.fixture(scope="module") +def daily(): + spec = importlib.util.spec_from_file_location("daily_infra_report", DAILY_REPORT) + assert spec and spec.loader + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +DAILY_AGENTS = { + "abiba": { + "platform": "pi", "ct": 100, "ip": MUMUNI_IP, + "zulip_connected": True, "zulip_processed": 5, + "pm2_status": "online", "pm2_restarts": "0", "pm2_uptime": "1h", + }, + "tanko": { + "platform": "dsh", "ct": 112, "ip": "192.168.68.122", + "gateway_state": "n/a (DSH)", "zulip_state": "connected", + "telegram_state": "unknown", "gateway_pid": None, "updated_at": "", + }, +} + + +def _fabricated_report(agents): + return { + "nodes": {}, + "node_count": 1, + "nodes_online": 1, + "total_vms": 0, + "running_vms": 0, + "stopped_vms": [], + "vms_by_node": {n: [] for n in + ["amdpve", "minipve", "storepve", "acerpve", "ocupve"]}, + "storage": [], + "docker_vm": {"total": 0, "running": 0, "unhealthy": [], + "containers": [], "reclaimable": "", "disk_used": "1%"}, + "docker_syslog": {"total": 0, "running": 0, "containers": []}, + "docker_netbird": {"total": 0, "running": 0, "containers": []}, + "endpoints": [], + "litellm": {"checks": []}, + "nfs": [], + "zulip_ext": { + "connected": True, "queue_id": "queue", "last_error": None, + "messages_processed": 0, "retry_count": 0, "pm2": {}, + "pm2_healthy": True, "bot_skipped_15min": 0, "finalized_1h": 0, + "failed_finalize_1h": 0, "finalize_fail_pct": 0, + "server_status": "200", + }, + "agents": agents, + } + + +def _agent_status_card(html): + start = html.index("🤖 Agent Status") + end = html.index("💬 Zulip Extension") + return html[start:end] + + +def test_daily_report_renders_only_abiba_and_tanko_agents(daily): + """build_html() over a Mumuni-free agent set must render no Mumuni row and + no Mumuni gateway-unknown issue, while abiba and tanko rows still render.""" + html = daily.build_html(_fabricated_report(dict(DAILY_AGENTS))) + card = _agent_status_card(html) + assert "mumuni" not in card.lower() + assert "abiba" in card + assert "tanko" in card + assert "mumuni" not in html.lower() + + +def test_daily_report_collect_never_probes_mumuni(monkeypatch, daily): + """collect() with ssh stubbed must add no mumuni agent and must never ssh + its decommissioned .24 host.""" + probed = [] + + class _NoSubprocess: + @staticmethod + def check_output(*args, **kwargs): + return b"" + + def fake_ssh(host, cmd): + probed.append(host) + return "" + + monkeypatch.setattr(daily, "pve_get", lambda path: []) + monkeypatch.setattr(daily, "ssh_jerome", lambda host, cmd: "") + monkeypatch.setattr(daily, "ssh", fake_ssh) + monkeypatch.setattr(daily, "http_get", + lambda url, auth=None, timeout=10: "200") + monkeypatch.setattr(daily, "http_get_body", + lambda url, auth=None, timeout=10: "") + monkeypatch.setattr(daily, "count_in_log", lambda *a, **k: 0) + monkeypatch.setattr(daily, "subprocess", _NoSubprocess) + + report = daily.collect() + assert "mumuni" not in report["agents"] + assert MUMUNI_IP not in probed + + +# ── scripts/agent-health-check.py: roster pin ─────────────────────── + +@pytest.fixture(scope="module") +def ahc(): + spec = importlib.util.spec_from_file_location("agent_health_check_roster", AHC) + assert spec and spec.loader + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def test_agent_health_roster_has_no_mumuni_entry(ahc): + assert "mumuni" not in ahc.AGENTS + + +# ── zulip-health.prose.md: contract reconciliation ────────────────── + +def test_health_contract_retires_mumuni_only_steps(): + text = HEALTH_CONTRACT.read_text() + assert MUMUNI_IP not in text + for step in ("**B4:", "**B5:", "**B6:"): + assert step not in text + + +def test_health_contract_states_mumuni_is_not_monitored_from_this_host(): + text = HEALTH_CONTRACT.read_text() + assert "Mumuni is NOT monitored from this host" in text + assert "monitored on her side" in text + assert "her own container" in text + + +def test_health_contract_keeps_tanko_agent_zero_and_bridge_steps(): + text = HEALTH_CONTRACT.read_text() + for marker in ("**B1:", "**B2:", "**B3:", "Step 4: Platform C", + "Step 2: Platform A", "Step 1: Zulip Server Liveness"): + assert marker in text, marker diff --git a/zulip-health.prose.md b/zulip-health.prose.md index 3e200ce..6ebafd4 100644 --- a/zulip-health.prose.md +++ b/zulip-health.prose.md @@ -1,9 +1,9 @@ --- kind: responsibility name: zulip-health -description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Tanko on DSH / Mumuni on Hermes), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity. +description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side. title: Zulip Mesh Health Monitor — Multi-Platform -version: 3.0.0 +version: 3.1.0 runtime_contract: 2 agent: abiba report_only_agents: @@ -12,13 +12,23 @@ report_only_agents: # Zulip Mesh Health Monitor -Monitors ALL Zulip-connected agents across platforms (pi, Hermes, DSH, Agent Zero). -Runs every 15 minutes in the background. Also triggers on session start. +Monitors the Zulip-connected agents under this host's operational control (pi, +DSH, Agent Zero). Runs every 15 minutes in the background. Also triggers on +session start. + +> **Mumuni is NOT monitored from this host (captain ruling 2026-09-10).** She +> moved off this host onto her own container — kagentz CT 105 on minipve +> (192.168.68.14), running a dedicated `hermes` user — and is monitored on her +> side. No step in this contract, and no leg of `scripts/zulip-monitor.sh`, may +> ssh to her old deployment, read her `~/.hermes/gateway_state.json`, or alert on +> her state. The former Platform-B-for-Mumuni steps (gateway process, heartbeat, +> response delivery) are retired: they always read "unknown" against the +> decommissioned deployment and produced a false 🔴 alert on every run. ## Requires - **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY` -- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); Mumuni (192.168.68.14, kagentz CT105 on minipve); and Agent Zero Docker host (192.168.68.14) +- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14) - **PM2** on localhost for pi process management - **Network access** to `chat.sysloggh.net`, `localhost:9200` - **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce` @@ -81,13 +91,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse ## Streaming Support (2026-07-05) Zulip agents now support progressive message editing during agent generation. -When a Zulip agent (Tanko on DSH, Mumuni on Hermes) processes a message, the response is +When a Zulip agent under this monitor's scope (Tanko on DSH) processes a message, the response is streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API: - Adapter implements `edit_message()` using `_api_patch()` helper - Gateway stream consumer progressively edits the Zulip message - User sees real-time agent thinking instead of waiting for full response -- Verified: Tanko (CT 112) and Mumuni (kagentz CT 105) both have streaming active +- Verified: Tanko (CT 112) has streaming active; Mumuni's (kagentz CT 105) is verified on her own host, not from here ### Verification ```bash @@ -183,7 +193,10 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta | Crash loop >10/h | Alert user | -### Step 3: Platform B — Tanko (DSH on amdpve CT 112) & Mumuni (Hermes) +### Step 3: Platform B — Tanko (DSH on amdpve CT 112) + +Mumuni is out of scope for this host (see the note above): she runs on her own +container and is monitored on her side. Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs @@ -240,44 +253,13 @@ timeout only. Never expect a bare `200` — the public URL terminates in the token-gated authentik chain. Statuses outside the healthy set are logged/reported as a warning — reported, never healed on. -**B4: Gateway Process** (Hermes agent Mumuni only — Tanko runs no Hermes gateway) - -```bash -ssh root@ "ps aux | grep 'gateway run' | grep -v grep" -``` - -Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more -than one `gateway run` process is found, the gateway has a collision (typically -one `--force` and one `--replace` process). Kill the newer/duplicate process, -then restart the remaining gateway per-agent (parameterized 2026-08-09, captain -ruling). Check the gateway log for "Gateway running with 2 platform(s)" (not 1) -to confirm Zulip reloaded. - -**B5: Heartbeat Verification** (Hermes agent Mumuni only — Tanko has no Hermes gateway) - -```bash -ssh root@192.168.68.24 "grep Heartbeat ~/.hermes/logs/agent.log | tail -3" -``` - -Expected: recent heartbeat (within 5 min), `polls=N` incrementing. -Silence > 300s → warning. Silence > 600s → critical. - -**B6: Response Delivery** (Hermes agent Mumuni only) - -```bash -ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10" -``` - -> 50% fail rate → critical. - **Platform B Actions** | Condition | Action | |-----------|--------| -| `zulip.state != "connected"` | `ssh root@ "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` (Mumuni) / restart Tanko via DSH service | -| No heartbeat in 10min | Same as above | -| `Failed to finalize` > 50% | Check PATCH API, Zulip server | -| Response empty/short | Check A2A endpoint / LiteLLM model | +| `dsh-web` service not `active` | Restart Tanko via DSH service | +| HTTP `:3080` connection refused/timeout (`000`) | Same as above | +| HTTP status outside the expected set | Log/report as a warning — reported, never healed on | ### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14) @@ -333,7 +315,7 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`. I Check each agent's log for excessive bot-to-bot chatter: - Abiba: `Skipped.*bot msgs` count -- Tanko/Mumuni: Repeated DM exchanges between bots +- Tanko: Repeated DM exchanges between bots - kagentz: Adapter log for bot DMs being processed If any bot processes >50 bot-originated messages in 15min → warning.