diff --git a/CLAUDE.md b/CLAUDE.md deleted file mode 120000 index 47dc3e3..0000000 --- a/CLAUDE.md +++ /dev/null @@ -1 +0,0 @@ -AGENTS.md \ No newline at end of file diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 0000000..a9d4d26 --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1,2 @@ + +@AGENTS.md diff --git a/README.md b/README.md index b97b0d1..46b8968 100644 --- a/README.md +++ b/README.md @@ -162,7 +162,7 @@ Companion shell scripts that contracts delegate to. |---|---| | `daily-infra-report.py` | Generates the daily infrastructure dashboard (HTML email to jerome@sysloggh.com). | | `pm2-self-heal.sh` | Shell companion to the pm2-self-heal contract — restarts crashed PM2 processes (runs every 5 min). | -| `agent-health-check.py` | Consolidated agent health: LiteLLM key validation + GPU port conflict + streaming checks (every 10 min). Replaced zulip-monitor.sh. | +| `agent-health-check.py` | Consolidated agent health: LiteLLM key validation + GPU port conflict + streaming checks (every 10 min). | ## Contract Structure diff --git a/contract-registry.yaml b/contract-registry.yaml index dc27ba8..fb50e4a 100644 --- a/contract-registry.yaml +++ b/contract-registry.yaml @@ -628,18 +628,18 @@ contracts: sensitivity: high status: active owner: abiba - version: 3.0.0 + version: 3.1.0 trigger: type: scheduled cadence: '*/15 * * * *' - description: "Every 15 minutes \u2014 monitors all Zulip-connected agents" + description: "Every 15 minutes \u2014 monitors the Zulip-connected agents under this host's control (pi, DSH, Agent Zero)" cron_job_id: null execution: agent: abiba timeout: 120 requires: - Zulip API key for abiba-bot@chat.sysloggh.net - - SSH access to all Hermes agents + - SSH access to amdpve (192.168.68.15) for Tanko (CT 112) and the Agent Zero Docker host (.14) verification: postconditions: - check: bot registration active @@ -1155,7 +1155,7 @@ contracts: type: scheduled cadence: 0 3 * * * description: Daily at 3am ET - cron_job_id: null + cron_job_id: b59f3cc21f4c # provisioned on kagentz 2026-09-08 (okyeame-memory-audit, glm-5.3-flash) execution: agent: mumuni timeout: 600 diff --git a/docs/AUTHORING-GUIDE.md b/docs/AUTHORING-GUIDE.md index bfe4c8b..f73129e 100644 --- a/docs/AUTHORING-GUIDE.md +++ b/docs/AUTHORING-GUIDE.md @@ -211,6 +211,40 @@ Errors tell the operator what went wrong and what to do about it. Be specific: Check router /health/unified at http://192.168.68.116/health/unified instead." ``` +### Report provenance + +Every report a contract produces must lead with the **absolute path the probe +executed from** — `pwd -P`, or the running script's absolute path. A live-state +report without provenance is unactionable: a report from a stale copy (a worktree +clone, a retired cron entry, a diverged consumer) looks identical to a live +fault, and the team burns rounds repairing healthy infrastructure. This is not +optional. The 2026-09-09 probe-drift rounds cost three false `DEGRADED` reports +because a stale consumer probed the wrong port and nothing in the report said +where it ran. + +Pair it with the **scoped any-HTTP-response liveness rule**: for unauthenticated +or auth-gated endpoints — where any HTTP answer proves a listener is up (the +PVE API's `401`, LiteLLM health's `301` redirect) — a probe is ALIVE on ANY HTTP +status, including `301` redirects and `401`/`403` auth challenges. **DOWN = +connection refused (`000`) or timeout only.** + +Probes whose success condition is specifically a bare `200` are NOT covered by +the any-HTTP rule. On those — authenticated probes such as the Zulip message +POST and the router `/health` — an unexpected status (`401`/`403` from a bad or +missing credential, `5xx`, or anything other than the expected `200`) is an +**ALERT**, not "alive". + +```markdown +**Report format**: Begin every report with the absolute execution path +(`pwd -P` / script path). On auth-gated endpoints, alive = ANY HTTP status and +DOWN = `000`/timeout only; on probes whose expected result is `200`, any other +status is an alert. +``` + +The lint pipeline enforces the provenance clause: any contract with a +`**Report format**` line must state an absolute path (`pwd -P`, `absolute path`, +or `executed from`). + ### Comments Comments in contracts explain WHY, not WHAT. The execution steps say what to diff --git a/docs/probe-drift-round2-evidence.md b/docs/probe-drift-round2-evidence.md new file mode 100644 index 0000000..c6a5b2d --- /dev/null +++ b/docs/probe-drift-round2-evidence.md @@ -0,0 +1,278 @@ +# Probe-drift round 2 — per-leg before/after evidence + +**Date:** 2026-09-10 +**Worktree (absolute execution path):** `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts` +**Branch:** `fm/probe-drift-round2-20260909` + +Every command below was run from the absolute path above; output is pasted +verbatim. This is the evidence trail for the four scoped corrections; it is not +a contract (never `prose run` it). + +--- + +## Leg 1 — agent-health-check (item 1) + +**Before** — from `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts`, +`python3 scripts/agent-health-check.py --no-deploy` (v2, base of this branch): + +``` +🏥 Agent Health Check v2 — 2026-09-10 01:16 UTC + +🔑 LiteLLM Keys: + ✅ tanko: key valid → syslog-auto + ✅ abiba: key valid → syslog-auto + ✅ koby: key valid → syslog-auto + ✅ koonimo: key valid → syslog-auto + +🎮 GPU Port Health: + ✅ gpu-rtx3090 (.8): healthy (pid=472206) + ✅ gpu-rtx5070 (.110): healthy (pid=207601) + ✅ gpu-strixhalo (.15): healthy (pid=4098872) + +🤖 Agent Gateways: + ✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK) + ⚠️ abiba: gw=no-state-file zulip=? streaming=no errors_10m=0 pid=? + ✅ koby: gw=running zulip=connected streaming=no errors_10m=0 pid=360900 + ✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125 + +🖥️ CT Liveness: + ✅ tanko (CT 112 on amdpve): running + ✅ abiba (CT 100 on minipve): running + ❌ koby (CT 111 on amdpve): PVE UNREACHABLE + ✅ koonimo (CT 113 on amdpve): running + +📝 Config Integrity: + ⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27 + ✅ abiba: config.yaml valid YAML + ✅ koby: config.yaml valid YAML + ✅ koonimo: config.yaml valid YAML + +🔌 Wrapper/CLI Integrity: + ⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27 + ⚠️ abiba: wrapper infisical path may be wrong (infisical at /usr/bin/infisical) + ❌ abiba: hermes-real NOT FOUND (wrapper broken) + ⚠️ abiba: .env may be missing LITELLM_API_KEY entry + ⚠️ koby: wrapper infisical path may be wrong (infisical at /usr/bin/infisical) + ❌ koby: hermes-real NOT FOUND (wrapper broken) + ✅ koby: wrapper + .env key present + ⚠️ koonimo: wrapper infisical path may be wrong (infisical at /usr/bin/infisical) + ✅ koonimo: wrapper + .env key present + +🔐 Vault Secrets: + ✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw + ✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg + ✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ + +❌ 6 FAILURE(S): ct-unreachable:koby:192.168.68.15 | wrapper-infisical-path:abiba | wrapper-no-hermes-real:abiba | wrapper-infisical-path:koby | wrapper-no-hermes-real:koby | wrapper-infisical-path:koonimo +``` + +Root causes (all stale expectations; no live fault): + +| Failure | Why it was stale | +|---|---| +| `ct-unreachable:koby:192.168.68.15` | CT 111 (tdunna/koby) runs on **storepve (.6)**, not amdpve (.15). | +| `wrapper-*:abiba` | Abiba is pi-only since the harness purge. `/root/.local/bin/hermes` is a dangling symlink; no `hermes-real`, no `~/.hermes/.env`. | +| `wrapper-*:koby` | Koby is **report-only** (captain ruling 2026-08-17, Rule 17): detect and report, never repair — its legs must not count as fleet failures. Koby's wrapper is also the genuine no-infisical case: it sources `~/.hermes/.env` rather than `/usr/bin/infisical`, which the check now accepts. | +| `wrapper-infisical-path:koonimo` | Koonimo's wrapper **does** reference `/usr/bin/infisical` — but past the old check's `head -20` window, so the check looked for the path in the wrong slice and false-failed. The fix that mattered was reading the full wrapper body (and then verifying any absolute infisical path it finds actually exists). | + +**After** — same absolute path, `python3 scripts/agent-health-check.py --no-deploy` (v4): + +``` +$ pwd -P +/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts +$ python3 scripts/agent-health-check.py --no-deploy +🏥 Agent Health Check v4 — 2026-09-10 01:22 UTC +📍 executed from: script=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts/scripts/agent-health-check.py cwd=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts + +🔑 LiteLLM Keys: + ✅ tanko: key valid → syslog-auto + ✅ abiba: key valid → syslog-auto + ✅ koby: key valid → syslog-auto + ✅ koonimo: key valid → syslog-auto + +🎮 GPU Port Health: + ✅ gpu-rtx3090 (.8): healthy (pid=472206) + ✅ gpu-rtx5070 (.110): healthy (pid=207601) + ✅ gpu-strixhalo (.15): healthy (pid=4098872) + +🤖 Agent Gateways: + ✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK) + ✅ abiba: pi-only runtime — no Hermes gateway since the harness purge (CT 100, SSH OK) + 🔍 koby: REPORT-ONLY mode (diagnostic only, no repairs on .129) + ✅ koby: gateway running (pid=360900, report-only mode) + ✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125 + +🖥️ CT Liveness: + ✅ tanko (CT 112 on amdpve): running + ✅ abiba (CT 100 on minipve): running + ✅ koby (CT 111 on storepve): running + ✅ koonimo (CT 113 on amdpve): running + +📝 Config Integrity: + ⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27 + ⏭️ abiba: pi-only runtime — no Hermes config.yaml since the harness purge + ✅ koby: config.yaml valid YAML + ✅ koonimo: config.yaml valid YAML + +🔌 Wrapper/CLI Integrity: + ⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27 + ⏭️ abiba: pi-only runtime — no hermes CLI wrapper since the harness purge + ℹ️ koby: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK + ❌ koby: hermes-real NOT FOUND (wrapper broken) + 🔍 report-only (koby): wrapper-no-hermes-real:koby — reported, not counted/repaired + ✅ koby: wrapper + .env key present + ✅ koonimo: wrapper infisical path OK + ✅ koonimo: wrapper + .env key present + +🔐 Vault Secrets: + ✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw + ✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg + ✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ + +✅ All checks passed +exit=0 +``` + +**Live vantage proof** (same worktree): + +``` +$ ssh root@192.168.68.15 "pct status 111" +Configuration file 'nodes/amdpve/lxc/111.conf' does not exist +$ ssh root@192.168.68.6 "pct status 111; pct list | grep '^ *111'" +status: running +111 running tdunna +$ ssh root@192.168.68.129 "hostname" +tdunna +``` + +--- + +## Leg 2 — infrastructure-monitoring PVE API (item 2) + +**Before** — the contract's probe, aimed at the monitoring host CT 116: + +``` +$ curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json +000 +``` + +CT 116 runs no `pveproxy`, so it never answers on `:8006`. The probe target was +wrong, which is what read as PVE-API `000`. + +**After** — probing the five real cluster nodes (`:8006/api2/json/version`), +alive under the any-HTTP-response rule (`401` = up, unauthenticated): + +``` +$ for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do + printf '%s:8006 -> %s\n' "$node" "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")" + done +192.168.68.9:8006 -> 401 +192.168.68.5:8006 -> 401 +192.168.68.15:8006 -> 401 +192.168.68.6:8006 -> 401 +192.168.68.12:8006 -> 401 +``` + +`401` on every node = alive by design. `DOWN` is `000`/timeout only. (The +contract's LiteLLM probe was the same class: `/litellm/health` answers `301` → +`/litellm/health/liveliness`, so it is now specified as any-HTTP too.) + +--- + +## Leg 3 — gpu-monitor GPU probes (item 3) + +**Before** — the false alarm came from probing bare port 80 on GPU hosts: + +``` +$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8/health +000 +$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110/health +000 +``` + +Nothing listens on GPU port 80, so the monitor reported +`DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` three times on 2026-09-09. + +**After** — the real endpoints answer: + +``` +$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health +200 +$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health +200 +$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.15:8080/health +200 +$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified +301 # Location: http://192.168.68.116/gpu/gpu-data — the same payload +$ curl -s -o /dev/null -w '%{http_code}' -L http://192.168.68.116/health/unified +200 +``` + +`301` is healthy under the any-HTTP-response rule. The contract now requires GPU +health on `:8080` (or router `/health/unified`) and forbids bare port 80 on a +GPU host. + +--- + +## Leg 4 — report provenance (item 4) + +Every contract report must now lead with the absolute path it executed from. +`docs/AUTHORING-GUIDE.md` documents the rule and `scripts/prose-lint.sh` +enforces it: + +``` +$ pwd -P +/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts +$ bash scripts/prose-lint.sh + ✅ Report provenance present in all report-format contracts +... +✅ LINT PASSED (12 warning(s)) +``` + +The health script prints `📍 executed from: script=… cwd=…` and includes +`execution_path`/`cwd` in `--json` output. + +--- + +## Full suite + +``` +$ pwd -P +/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts +$ python3 -m pytest -q +24 passed +$ shellcheck scripts/prose-lint.sh +(clean) +``` + +--- + +## Follow-up findings (observed, intentionally NOT changed here) + +These are adjacent stale expectations discovered while verifying the four +scoped legs. Each touches a CRITICAL/HIGH-sensitivity artifact or an unrelated +script, so it is recorded for the captain/verify mate rather than silently +repaired. + +1. **`infrastructure-control.prose.md` (CRITICAL) CT 111 node assignment.** + Lines ~109 and ~615 place `tdunna` (CT 111, koby) on **amdpve**. Live + verification on 2026-09-10 shows `pct status 111` = `running` on + **storepve (.6)** and `Configuration file 'nodes/amdpve/lxc/111.conf' does + not exist` on .15. `agent-health-check.py` now carries the live-verified + `storepve` mapping (the script is not the topology source of truth); the + CRITICAL contract itself needs an authorized correction. +2. **Strix Halo `:8080` firewall claim is stale.** `prose-ai-review.sh` + ground-truth rule #4 and `gpu-monitor.prose.md` say `:8080` is firewalled to + `.116` only and `.24` cannot probe it. Live on .15: + `-A INPUT -s 192.168.68.24/32 -p tcp --dport 8080 -j ACCEPT`, and a probe + from .24 returns `200`. The contract keeps routing Strix via the router + (safe), but the claim no longer matches iptables. +3. **`contract-registry.yaml` references `agent-health-check.prose.md`**, which + does not exist in the repo. The registry entry (with `koby_action: skip_heal`) + is aspirational/stale. +4. **Pre-existing script defects, untouched:** `scripts/pm2-self-heal.sh` has a + bash syntax error at lines 19–20 (`bash -n` fails), and `shellcheck` fails on + five untouched scripts (`netbird-add-domain.sh`, `pct-run.sh`, + `pm2-self-heal.sh`, `prose-ai-review.sh`, `swap-gpu-dense-model.sh`). + `scripts/prose-lint.sh` — the one shell file touched here — is now + shellcheck-clean. diff --git a/gpu-monitor.prose.md b/gpu-monitor.prose.md index ee98873..881bac4 100644 --- a/gpu-monitor.prose.md +++ b/gpu-monitor.prose.md @@ -31,20 +31,31 @@ agent: abiba └──────┘ │qwen27B│ │LiteLLM │ └──────┘ │dashboard│ └────────┘ +``` Note: JSON sidecar exporters at :8090 were never deployed on any GPU host. Router falls back to GPU /health direct probe. Monitor should use router /health/unified as source of truth for GPU status. Strix Halo :8080 is firewalled to .116 only — monitor on .24 cannot poll .15:8080 directly; must go through router on .116. -``` + +**PORT RULE (verified 2026-09-10):** GPU per-host health lives on **:8080** +(`http://:8080/health`); Prometheus GPU exporters live on **:9400**. +There is NO listener on bare port 80 for any GPU host — `http://192.168.68.8/health` +and `http://192.168.68.110/health` answer `000`. Never use a bare-port-80 probe +as a GPU liveness signal: on 2026-09-09 that produced three false +`DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` rounds while +`http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health` +answered `200`. Port 80 is valid only on the router (.116), never on a GPU host. ### Subsystems Polled | Subsystem | Endpoint | Frequency | Metrics | |-----------|----------|-----------|---------| -| GPU Status (all, via router) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (router probes each GPU /health directly) | -| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status | +| GPU Status (all, via router) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (router probes each GPU /health directly) — `301` → `/gpu/gpu-data` is **alive** | +| GPU .8 (RTX 3090) health | `http://192.168.68.8:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** | +| GPU .110 (RTX 5070) health | `http://192.168.68.110:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** | +| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (`301` → `/gpu/gpu-data` = alive) | | Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness | | LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count | | Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) | @@ -61,6 +72,22 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and ## Alert Thresholds +### Liveness rule (scoped) + +The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness +endpoints, where any HTTP answer proves a listener is up. Applied here: the +router's `/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data` +(the same payload) and LiteLLM's `/litellm/health` answers `301` → +`/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on +**ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included — +and **DOWN = connection refused (`000`) or timeout only**. Same scoped rule as +zulip-health (Tanko) and infrastructure-monitoring. + +Probes whose success condition is specifically a bare `200` are NOT covered by +the any-HTTP rule. On those — the GPU `:8080/health` endpoints, the router +`/health`, and the dashboard — an unexpected status (`401`/`403`, `5xx`, or +anything other than the expected `200`) is an **ALERT**, not "alive". + | Metric | Warning | Critical | |--------|---------|----------| | GPU Temp | >80°C | >90°C | @@ -97,7 +124,47 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and `curl http://localhost:9100/gpu-data | jq` — Full fleet status ### check-health -`curl http://localhost:9100/health` — Monitor self-check + +**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** + +```bash +# Provenance — run first; paste the absolute path into the report +pwd -P + +# GPU Monitor health + curl http://localhost:9100/health | jq +# Expected: 200 with {"status": "healthy", "cache_age_seconds": } + +# GPU host health — DIRECT on :8080. NEVER probe bare port 80 on a GPU host: +# http://192.168.68.8/health has no listener and returns 000 → false DEGRADED. +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health +# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health +# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) + +# Router unified health (source of truth; 301 → /gpu/gpu-data is HEALTHY) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified +# Expected: 301 (or 200 after following the redirect) — any HTTP status = alive + +# Router basic health (via nginx on port 80 — router .116 only, never a GPU host) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health +# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) + +# LiteLLM health (via nginx on port 80) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health +# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive + +# Dashboard + curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/ +# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) +``` + +**Report format**: Begin every report with the **absolute path the probe executed +from** (`pwd -P`, or the monitor script's absolute path) so a stale-consumer +report is distinguishable from a real fault at read time. Summarize actual +results from each probe. Apply the scoped liveness rule above: on auth-gated +endpoints only connection-refused (`000`) or timeout is DOWN; on bare-200 probes +any other status is an alert. Never probe a GPU host on bare port 80. ### view-dashboard Open `http://localhost:9100/` in browser — Live HTML dashboard @@ -110,9 +177,11 @@ python3 /root/scripts/gpu-monitor-server.py & Or via PM2: `pm2 restart gpu-monitor` ### check-router -The router health is accessed through nginx on port 80 (NOT port 9000 directly). -`curl http://192.168.68.116/health/unified` — Router unified health via nginx proxy +The router health is accessed through nginx on port 80 on the **router** +(.116) — NOT port 9000 directly, and NOT bare port 80 on a GPU host. +`curl http://192.168.68.116/health/unified` — Router unified health via nginx proxy; answers `301` → `/gpu/gpu-data` (same payload) = alive `curl http://192.168.68.116:9000/health/unified` — ❌ WILL FAIL (port bound to 127.0.0.1 only) +`curl http://192.168.68.8/health` — ❌ NEVER USE (GPU host, no port-80 listener → false `000`/DEGRADED) ## Configuration Files @@ -124,13 +193,18 @@ The router health is accessed through nginx on port 80 (NOT port 9000 directly). ## Execution -1. **Poll router** (every 15s): GET .116/health/unified — single source of truth for all GPU status (router probes each GPU /health directly via sidecar fallback) -2. **Poll router** (every 15s): GET .116/health via nginx:80 -3. **Poll LiteLLM** (every 15s): GET .116/litellm/health via nginx:80 -4. **Poll Strix** (every 15s): via router /health/unified (cannot poll .15:8080 directly — firewalled to .116 only) -5. **Poll dashboard** (every 15s): GET .116/dashboard/ -6. **Check alerts**: Compare metrics against thresholds -7. **Compute summary**: Fleet-wide health aggregation -8. **Render dashboard**: Generate HTML at /root/dashboard/gpu-fleet.html -9. **Serve API**: HTTP server on port 9100 -10. **Repeat** every 15 seconds +**Port discipline:** probe GPU hosts on `:8080` (or the router's +`/health/unified`); probe port 80 only on the router (.116). Never bare port 80 +on a GPU host. + +1. **Poll router** (every 15s): GET .116/health/unified — single source of truth for all GPU status (router probes each GPU /health directly via sidecar fallback). `301` → `/gpu/gpu-data` counts as alive. +2. **Fallback direct GPU probe** (only if router /health/unified is DOWN): GET `http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health` — **:8080 only, never bare port 80**. +3. **Poll router** (every 15s): GET .116/health via nginx:80 +4. **Poll LiteLLM** (every 15s): GET .116/litellm/health via nginx:80 +5. **Poll Strix** (every 15s): via router /health/unified (cannot poll .15:8080 directly — firewalled to .116 only) +6. **Poll dashboard** (every 15s): GET .116/dashboard/ +7. **Check alerts**: Compare metrics against thresholds +8. **Compute summary**: Fleet-wide health aggregation +9. **Render dashboard**: Generate HTML at /root/dashboard/gpu-fleet.html +10. **Serve API**: HTTP server on port 9100 +11. **Repeat** every 15 seconds diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index 62d8d4b..2969ac6 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -102,13 +102,36 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo) - Stack persists across reboots (systemd for exporters, Docker restart policy) ## Execution + +### Liveness rule (scoped) + +The any-HTTP-response rule applies ONLY to unauthenticated/auth-gated endpoints, +where any HTTP answer proves a listener is up: the PVE API +(`https://:8006/api2/json/version`) and LiteLLM health +(`/litellm/health`, `301` → `/litellm/health/liveliness`). For those endpoints a +probe is **ALIVE** on **ANY** HTTP status — `401`/`403` auth challenges and `3xx` +redirects included — and **DOWN = connection refused (`000`) or timeout only**. +The PVE API legitimately answers `401` to an unauthenticated probe — that is the +healthy signal, not a failure. Same scoped rule as zulip-health (Tanko) and +gpu-monitor. + +Probes whose success condition is specifically a bare `200` are NOT covered by +the any-HTTP rule. On those — the authenticated Zulip POST and the router +`/health` — an unexpected status (`401`/`403` from a bad or missing credential, +`5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive". + ### check-health **RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** ```bash +# Provenance — run first; paste the absolute path into the report +pwd -P + # Zulip API health (POST ping) -curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u 'abiba-bot@chat.sysloggh.net:KEY' +source /etc/litellm-monitor.env +ZULIP_USER="abiba-bot@chat.sysloggh.net" +curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}" # Expected: 200 (HTTP 000 = unreachable/cache) # PM2 process health @@ -120,6 +143,27 @@ curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL" curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL" curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL" +# Router health (via nginx on port 80) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health +# Expected: 200 (Router is up and responding) + +# LiteLLM health (via nginx on port 80) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health +# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive + +# PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring +# host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 — +# that was the stale-vantage bug this replaces (CT 116 is the monitoring host, +# not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE +# by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response); +# DOWN = connection refused (000) or timeout only. +for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do + printf '%s:8006 -> %s\n' "$node" \ + "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")" +done +# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12) +# A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault. + # Prometheus targets curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets' # Expected: All targets UP (may show some down if exporters not deployed) @@ -128,12 +172,21 @@ curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets' curl -s http://192.168.68.116:3001/api/health | jq '{status, version}' # Expected: {"status":"ok","version":"..."} -# LiteLLM metrics +# LiteLLM metrics (Prometheus endpoint) curl -s http://192.168.68.116:4001/metrics | head -20 # Expected: Prometheus-formatted metrics output ``` -**Report format**: Summarize actual results from each probe. If any probe returns non-200 or empty output, flag as alert. +**Report format**: Begin every report with the **absolute path the probe executed +from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is +distinguishable from a real fault at read time. Summarize actual results from +each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE +API and LiteLLM endpoints above: only connection-refused (`000`) or timeout is +DOWN; empty output is a warning. For probes whose expected result is a bare `200` +(the authenticated Zulip POST, router `/health`), flag an alert on any unexpected +status (`401`/`403`/`5xx`) — do not summarize it as alive. A bare-`200` +expectation on the auth-gated PVE API (`401`) or LiteLLM health (`301` redirect) +is a stale expectation, not a fault. ### Phase 1: GPU Exporters diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index c7680a8..734c26a 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -113,7 +113,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only. - **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`). - **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault). -- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v2 (2026-07-26) — reads each agent's **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key). Covers: LiteLLM keys, GPU ports, agent gateways (all 5 agents now SSHa ble), CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames. +- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). The current fleet roster is owned by the script changelog (`scripts/agent-health-check.py`); mumuni is no longer probed from this host. Legacy `tdunna`/`baggy` replaced with canonical agent hostnames. - **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM. ## Maintains diff --git a/proxmox-monitor.prose.md b/proxmox-monitor.prose.md index 90fadaa..39e81d6 100644 --- a/proxmox-monitor.prose.md +++ b/proxmox-monitor.prose.md @@ -100,6 +100,35 @@ Edit `build-dashboards.py`, run it, `docker restart harness-grafana` ### check-targets `curl http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job,health}'` +### check-health + +**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** + +```bash +# Prometheus health (bound to 0.0.0.0:9090 on .116) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:9090/-/healthy +# Expected: 200 (Prometheus is up and healthy) + +# Grafana health (bound to 0.0.0.0:3001 on .116) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:3001/api/health +# Expected: 200 (Grafana is up and healthy) + +# Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — must probe from .116 localhost) +ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9324/metrics" +# Expected: 200 (docker-stats-exporter is up and responding) + +# PVE exporter (bound to 127.0.0.1:9221 on .116 — must probe from .116 localhost) +ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9221/metrics" +# Expected: 200 (pve-exporter is up and responding) +``` + +**Report format**: Begin every report with the **absolute path the probe executed +from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is +distinguishable from a real fault at read time. Summarize actual results from +each probe. If any probe returns non-200, flag as alert. + +**Note**: Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only) so they must be probed from .116 via SSH. Prometheus and Grafana are bound to 0.0.0.0 so they can be probed from the LAN. + ### restart-exporter `cd /opt/monitoring && docker compose restart pve-exporter docker-stats` diff --git a/scripts/agent-health-check.py b/scripts/agent-health-check.py index 22a3b98..c5d2618 100755 --- a/scripts/agent-health-check.py +++ b/scripts/agent-health-check.py @@ -1,6 +1,6 @@ #!/usr/bin/env python3 """ -/root/scripts/agent-health-check.py — Consolidated Agent Health Verification v2 +/root/scripts/agent-health-check.py — Consolidated Agent Health Verification v4 Verifies: LiteLLM keys (agent-specific), GPU port conflicts, agent Zulip streaming, gateway liveness, gateway log health, CT liveness, config YAML integrity, @@ -17,11 +17,42 @@ Changelog: v2 (2026-07-26): Added CT liveness, config validation, wrapper integrity, vault secret emptiness check. Fixed Koby/Koonimo SSH hosts and agent key name format ({NAME}_LITELLM_API_KEY not LITELLM_API_KEY_{NAME}). - Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), - abiba (.24). + Fleet roster: tanko (.122), koby (.129), koonimo (.114), abiba (.24). + (v2 also carried a mumuni probe; see v5 — mumuni is no longer probed: she + moved to her own container, kagentz CT 105 / .14, and is monitored there.) + v3 (2026-09-08): GPU unit repoint verified live (.8 llama-chat-api.service, + .110 llama-server.service, .15 strix-server.service) — .8 was probing a stale + llama-server unit that reads inactive, producing false UNREACHABLE legs. + systemctl is-active no longer swallows non-zero exit as SSH failure. + Fixed UnboundLocalError on the abiba/koonimo gateway leg (pid unbound in the + summary f-string). Abiba's LiteLLM key now comes from /root/.pi/agent/env.sh + (#735 agent separation; creds moved out of shared /root/.bashrc). + v4 (2026-09-10): probe-drift round 2 (prose-contracts follow-up to #65/#66/#68). + abiba declared pi-only runtime — Hermes-era config/wrapper/gateway checks are + skipped (harness purge). koby declared report_only per the captain's + 2026-08-17 ruling: every koby leg is detected and reported, never counted as a + fleet failure and never repaired. koby's PVE mapping corrected to storepve + (CT 111 tdunna lives on .6 — the old amdpve mapping produced a false + ct-unreachable). The wrapper infisical-path check had two stale-expectation + bugs: it read only the first 20 lines of the wrapper, so koonimo (whose + wrapper does reference /usr/bin/infisical, just past line 20) was falsely + FAILed as "path may be wrong"; and it treated the absence of any infisical + reference as a fault, though koby's wrapper sources the key from + ~/.hermes/.env and never invokes infisical. The check now reads the full + wrapper body, accepts a no-infisical wrapper, and verifies that any absolute + infisical path the wrapper references actually exists. Report-only findings + are surfaced in a machine-readable `report_only` array in --json output, + separate from `failures`. Every run prints absolute execution provenance + (script + cwd) in the header, in the cron ALERT line, and in --json output so + a stale-consumer report is distinguishable from a fault at read time. + v5 (2026-09-10): roster correction only, no behavior change. mumuni was removed + from the AGENTS dict when she moved off this host onto her own container + (kagentz CT 105 on minipve, .14, dedicated `hermes` user) and is monitored + from her side. This script must not probe mumuni or .24 — the v2 changelog + roster line was the last reference still placing her at .24 / CT100. """ -import subprocess, json, sys, os, time +import subprocess, json, sys, os, time, re, io, contextlib from datetime import datetime LITELLM = "http://192.168.68.116:80" @@ -40,18 +71,55 @@ PVE_NODES = { # Agent definitions: ct, host, user, pve_node, vault_key_name AGENTS = { "tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"}, - "abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve", "vault_key": None}, # Pi agent + Mumuni Zulip, no vault key - "koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"}, + # abiba = pi agent (.24) — no vault key; its LiteLLM key is read from its + # local env file (key_env below), not from the shared vault or .bashrc. + # runtime=pi: abiba has run pi-only since the harness purge. There is no + # Hermes gateway, no ~/.hermes/config.yaml and no hermes CLI wrapper on .24 + # (the /root/.local/bin/hermes symlink is dangling), so the Hermes-era + # config/wrapper/gateway legs are skipped rather than reported as faults. + "abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve", + "vault_key": None, "runtime": "pi", + "key_env": {"file": "/root/.pi/agent/env.sh", "var": "LITELLM_API_KEY"}}, + # koby = report-only (captain's 2026-08-17 ruling, Rule 17): detect and + # report, NEVER repair, and never count against fleet failures. CT 111 + # (tdunna) lives on storepve (.6) — verified live 2026-09-10; the previous + # amdpve mapping made `pct status 111` fail and read as ct-unreachable. + "koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "storepve", "vault_key": "KOBY_LITELLM_API_KEY", "report_only": True}, "koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"}, } +# Systemd units verified live 2026-09-08 (systemctl list-units on each host): +# .8 rtx3090 (gpu-dense) -> llama-chat-api.service (active; the old +# llama-server.service unit file is stale/inactive — probing it read as +# UNREACHABLE for a healthy process) +# .110 rtx5070 (ocu-llm VM) -> llama-server.service (active) +# .15 strixhalo (amdpve) -> strix-server.service (active) GPU_HOSTS = { - "gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-server"}, - "gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server"}, - "gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server"}, + "gpu-rtx3090 (.8)": {"host": "192.168.68.8", "port": 8080, "service": "llama-chat-api.service"}, + "gpu-rtx5070 (.110)": {"host": "192.168.68.110", "port": 8080, "service": "llama-server.service"}, + "gpu-strixhalo (.15)": {"host": "192.168.68.15", "port": 8080, "service": "strix-server.service"}, } FAIL = [] +REPORT_ONLY = [] + + +def _fail(key, agent_name=None): + """Record a failure, except for report-only agents. + + Koby is report-only per the captain's 2026-08-17 ruling (Rule 17): its legs + are detected and reported, never repaired and never counted as fleet + failures. A red fleet alert on a known report-only leg is a false alarm. + Report-only findings are tracked separately so --json consumers can still + see them without them counting as fleet failures. Any non-report-only agent + (or a leg with no agent, e.g. GPU hosts) records normally. + """ + if agent_name and AGENTS.get(agent_name, {}).get("report_only"): + REPORT_ONLY.append(key) + print(f" 🔍 report-only ({agent_name}): {key} — reported, not counted/repaired") + return + FAIL.append(key) + INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN") INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net") @@ -161,11 +229,46 @@ def _get_agent_key(agent_name, vault_key_name): return None -# Inject keys from vault for each agent -for agent_name in AGENTS: - info = AGENTS[agent_name] - key = _get_agent_key(agent_name, info.get("vault_key")) - AGENTS[agent_name]["key"] = key + +def _read_env_export(path, var): + """Parse `export VAR=value` (or `VAR=value`) out of a local env file. + + #735 agent separation (2026-09-06): agent creds moved out of the shared + /root/.bashrc into per-agent env files under /root/.pi/agent/ (bashrc's + source line keeps abiba shells resolving them, but the file of record is + env.sh). Do NOT fall back to /root/.bashrc here: desktop (.200) SSH + sessions override LITELLM_API_KEY with mumuni's key, so sourcing bashrc + would validate the wrong identity. + """ + try: + with open(os.path.expanduser(path)) as _f: + for line in _f: + line = line.strip() + if not (line.startswith("export " + var + "=") or line.startswith(var + "=")): + continue + value = line.split("=", 1)[1].strip().strip('"').strip("'") + if value: + return value + except (OSError, UnicodeDecodeError): + pass + return None + + +def load_agent_keys(): + """Populate AGENTS[*]["key"] from the vault or the agent's local env file. + + Called from main(), not at import: keeping this out of module scope lets the + module be imported (and unit tested) without live vault/SSH access. Vault + format is {NAME}_LITELLM_API_KEY (project 322fceab-39da-4854-a55a-568e76c0f13f, + env prod); abiba has no vault key and reads LITELLM_API_KEY from its local + /root/.pi/agent/env.sh (moved there from /root/.bashrc in #735). + """ + for agent_name in AGENTS: + info = AGENTS[agent_name] + key = _get_agent_key(agent_name, info.get("vault_key")) + if not key and info.get("key_env"): + key = _read_env_export(info["key_env"]["file"], info["key_env"]["var"]) + AGENTS[agent_name]["key"] = key # ═══════════════════════════════════════════════════════════════════ @@ -176,8 +279,8 @@ def check_keys(): for name, agent in AGENTS.items(): key = agent.get("key") if not key: - print(f" ❌ {name}: NO KEY FOUND (vault empty or unreachable)") - FAIL.append(f"key:{name}:no-key") + print(f" ❌ {name}: NO KEY FOUND (vault/env empty or unreachable)") + _fail(f"key:{name}:no-key", name) continue data = http_json(f"{LITELLM}/v1/models", headers={"Authorization": f"Bearer {key}"}) @@ -186,11 +289,11 @@ def check_keys(): print(f" ✅ {name}: key valid → {model}") else: print(f" ❌ {name}: KEY FAILURE — auth rejected or unreachable") - FAIL.append(f"key:{name}") + _fail(f"key:{name}", name) # ═══════════════════════════════════════════════════════════════════ -# CHECK 2: GPU Port Conflict Detection (unchanged) +# CHECK 2: GPU Port Conflict Detection (unit names verified live 2026-09-08) # ═══════════════════════════════════════════════════════════════════ def check_gpu_ports(): @@ -199,7 +302,11 @@ def check_gpu_ports(): port = gpu["port"] svc = gpu["service"] - svc_status = ssh(host, f"systemctl is-active {svc}") + # `systemctl is-active` exits non-zero when the unit is inactive or + # missing, which the ssh() helper would swallow as an SSH failure and + # report as UNREACHABLE. `|| true` keeps the real state word so we can + # tell "unit inactive" from "host unreachable". + svc_status = ssh(host, f"systemctl is-active {svc} || true") port_owner = ssh(host, f"ss -tlnp 2>/dev/null | grep -Po ':{port}\\s+.*pid=\\K[0-9]+' | head -1") if not svc_status: @@ -242,28 +349,41 @@ def check_agents(): # Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — it no longer runs a # Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks. - if agent.get("runtime") == "dsh": + # Non-Hermes runtimes have no gateway to probe. dsh = Tanko since + # 2026-08-27; pi = abiba since the harness purge (.24 is pi-only). + if agent.get("runtime") in ("dsh", "pi"): + is_dsh = agent.get("runtime") == "dsh" + label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime" + since = "since 2026-08-27" if is_dsh else "since the harness purge" live = ssh(host, "true", user=user) - print(f" {'✅' if live is not None else '❌'} {name}: DSH (DeepSeek Harness) — " - f"no Hermes gateway since 2026-08-27 (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})") + print(f" {'✅' if live is not None else '❌'} {name}: {label} — " + f"no Hermes gateway {since} (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})") if live is None: - FAIL.append(f"unreachable:{name}") + _fail(f"unreachable:{name}", name) continue if not host or not user: print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check") continue + # Resolve the Hermes gateway PID once, before the report-only branch: + # the summary line below renders `pid`, and it used to be bound only in + # the report-only path — leaving it unbound on the abiba/koonimo path + # raised UnboundLocalError and crashed the whole check. Agents without + # a gateway get pid=?. + pid = ssh(host, "pgrep -f '[h]ermes_cli.main gateway run' | grep -v infisical | head -1", user=user) + if not pid: + pid = ssh(host, "pgrep -f '[h]ermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user) + if not pid: + pid = "?" + # ⛔ KOBY IS NEVER REPAIRED — diagnostic only if report_only: print(f" 🔍 {name}: REPORT-ONLY mode (diagnostic only, no repairs on .129)") # Still check gateway status for reporting purposes - pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user) - if not pid: - pid = ssh(host, "pgrep -f 'hermes.*gateway' | grep -v infisical | grep -v bash | head -1", user=user) - if not pid: + if pid == "?": print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)") - FAIL.append(f"gateway-down:{name}") + _fail(f"gateway-down:{name}", name) continue else: print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)") @@ -326,12 +446,12 @@ def check_ct_liveness(): status = ssh(pve_ip, f"pct status {ct} 2>/dev/null", user="root") if not status: print(f" ❌ {name} (CT {ct} on {pve_node}): PVE UNREACHABLE") - FAIL.append(f"ct-unreachable:{name}:{pve_ip}") + _fail(f"ct-unreachable:{name}:{pve_ip}", name) elif "running" in status: print(f" ✅ {name} (CT {ct} on {pve_node}): running") elif "stopped" in status: print(f" ❌ {name} (CT {ct} on {pve_node}): STOPPED") - FAIL.append(f"ct-stopped:{name}") + _fail(f"ct-stopped:{name}", name) else: print(f" ⚠️ {name} (CT {ct} on {pve_node}): {status.strip()}") @@ -347,6 +467,9 @@ def check_config_integrity(): if agent.get("runtime") == "dsh": print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27") continue + if agent.get("runtime") == "pi": + print(f" ⏭️ {name}: pi-only runtime — no Hermes config.yaml since the harness purge") + continue host = agent.get("host") user = agent.get("user") if not host or not user: @@ -361,18 +484,38 @@ def check_config_integrity(): user=user) if not yaml_ok: print(f" ❌ {name}: SSH UNREACHABLE (config check skipped)") - FAIL.append(f"config-unreachable:{name}") + _fail(f"config-unreachable:{name}", name) elif "OK" in yaml_ok: print(f" ✅ {name}: config.yaml valid YAML") else: print(f" ❌ {name}: config.yaml YAML ERROR — {yaml_ok[:120]}") - FAIL.append(f"config-yaml-error:{name}") + _fail(f"config-yaml-error:{name}", name) # ═══════════════════════════════════════════════════════════════════ # CHECK 6: Wrapper/CLI Integrity (NEW) # ═══════════════════════════════════════════════════════════════════ +def _infisical_invocation_paths(wrapper_body): + """Absolute infisical paths the wrapper actually invokes. + + Only executed (non-comment) lines count, and only a path followed by a real + infisical subcommand (e.g. `/usr/bin/infisical run`) is treated as an + invocation. A note such as `# migrated from /usr/local/bin/infisical` is + prose, not a call, so it must not manufacture a dangling-path false alarm. + """ + paths = [] + for line in wrapper_body.splitlines(): + code = line.split("#", 1)[0] + for _m in re.finditer( + r"(/[A-Za-z0-9._/-]*infisical)\s+(?:run|export|secrets|login|logout)\b", + code, + ): + if _m.group(1) not in paths: + paths.append(_m.group(1)) + return paths + + def check_wrapper_integrity(): """Verify the hermes CLI wrapper exists and can reach hermes-real.""" for name, agent in AGENTS.items(): @@ -380,6 +523,9 @@ def check_wrapper_integrity(): if agent.get("runtime") == "dsh": print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27") continue + if agent.get("runtime") == "pi": + print(f" ⏭️ {name}: pi-only runtime — no hermes CLI wrapper since the harness purge") + continue host = agent.get("host") user = agent.get("user") if not host or not user: @@ -393,24 +539,56 @@ def check_wrapper_integrity(): wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user) if not wrapper: print(f" ❌ {name}: NO HERMES CLI WRAPPER FOUND") - FAIL.append(f"wrapper-missing:{name}") + _fail(f"wrapper-missing:{name}", name) continue else: print(f" ⚠️ {name}: hermes at {wrapper.strip()} (not ~/.local/bin/hermes)") - # Check wrapper has correct infisical path - infisical_path_valid = ssh(host, - "head -20 /root/.local/bin/hermes 2>/dev/null | grep -q '/usr/bin/infisical' && echo OK || echo MISS", - user=user) - if infisical_path_valid == "MISS": - # Check if infisical exists on path - inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user) - if not inf_actual: - print(f" ❌ {name}: INFISICAL NOT INSTALLED (wrapper broken)") - FAIL.append(f"wrapper-no-infisical:{name}") + # Credential-injection mechanism. The Hermes-era wrapper injected creds + # with `/usr/bin/infisical run`, but the mechanism is not required to be + # infisical at all: koby's wrapper sources the key from ~/.hermes/.env + # and never mentions infisical, which is valid. The old check read only + # the first 20 lines, so koonimo's wrapper — which DOES reference + # /usr/bin/infisical, just past line 20 — false-failed as "path may be + # wrong". Read the full body, accept a no-infisical wrapper, and verify + # the absolute infisical path(s) the wrapper actually invokes. Only + # executed (non-comment) lines count: a comment or dead prose mentioning + # a removed path (litellm-api-keys.prose.md documents + # `rm -f /usr/local/bin/infisical`) must neither produce a dangling path + # nor trigger the PATH check — it is not an invocation. + wrapper_body = ssh(host, "cat /root/.local/bin/hermes 2>/dev/null", user=user) or "" + wrapper_code = "\n".join(line.split("#", 1)[0] for line in wrapper_body.splitlines()) + invoked_paths = _infisical_invocation_paths(wrapper_body) + if "infisical" in wrapper_code: + if invoked_paths: + missing = [] + for _p in invoked_paths: + _exists = ssh(host, f"test -x {_p} && echo OK || echo MISS", user=user) + if not _exists or _exists.strip().splitlines()[-1] != "OK": + missing.append(_p) + if len(missing) == len(invoked_paths): + inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user) + suffix = f" (infisical at {inf_actual})" if inf_actual else "" + print(f" ❌ {name}: wrapper invokes infisical via missing path(s) " + f"{', '.join(missing)}{suffix}") + _fail(f"wrapper-infisical-path:{name}", name) + elif missing: + print(f" ⚠️ {name}: wrapper has an unused/missing infisical path " + f"({', '.join(missing)}) but a working invocation — informational") + elif "/usr/bin/infisical" not in invoked_paths: + print(f" ⚠️ {name}: wrapper infisical path differs " + f"({', '.join(invoked_paths)}) — informational") + else: + print(f" ✅ {name}: wrapper infisical path OK") else: - print(f" ⚠️ {name}: wrapper infisical path may be wrong (infisical at {inf_actual})") - FAIL.append(f"wrapper-infisical-path:{name}") + inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user) + if not inf_actual: + print(f" ❌ {name}: wrapper invokes infisical but the binary is MISSING") + _fail(f"wrapper-no-infisical:{name}", name) + else: + print(f" ✅ {name}: wrapper infisical resolves via PATH ({inf_actual})") + else: + print(f" ℹ️ {name}: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK") # Check hermes-real exists hermes_real = ssh(host, @@ -423,7 +601,7 @@ def check_wrapper_integrity(): user=user) if not hermes_real or hermes_real.strip() == "MISS": print(f" ❌ {name}: hermes-real NOT FOUND (wrapper broken)") - FAIL.append(f"wrapper-no-hermes-real:{name}") + _fail(f"wrapper-no-hermes-real:{name}", name) else: print(f" ✅ {name}: hermes-real at alt path") @@ -451,10 +629,10 @@ def check_vault_secrets(): key = agent.get("key") if not key: print(f" ❌ {name}: vault secret {vault_key_name} MISSING or EMPTY") - FAIL.append(f"vault-empty:{name}:{vault_key_name}") + _fail(f"vault-empty:{name}:{vault_key_name}", name) elif not key.startswith("sk-"): print(f" ❌ {name}: vault secret {vault_key_name} WRONG FORMAT (starts '{key[:8]}...')") - FAIL.append(f"vault-bad-format:{name}:{vault_key_name}") + _fail(f"vault-bad-format:{name}:{vault_key_name}", name) else: print(f" ✅ {name}: vault {vault_key_name}=sk-...{key[-4:]}") @@ -487,18 +665,7 @@ def deploy_self(): # MAIN # ═══════════════════════════════════════════════════════════════════ -def main(): - quiet = "--quiet" in sys.argv - as_json = "--json" in sys.argv - - # Self-deploy to canonical location - if not quiet and "--no-deploy" not in sys.argv: - deploy_self() - - if not quiet: - print(f"🏥 Agent Health Check v2 — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}") - print() - +def _run_checks(): print("🔑 LiteLLM Keys:") check_keys() print() @@ -526,16 +693,49 @@ def main(): print("🔐 Vault Secrets:") check_vault_secrets() + +def main(): + quiet = "--quiet" in sys.argv + as_json = "--json" in sys.argv + + # Self-deploy to canonical location + if not quiet and "--no-deploy" not in sys.argv: + deploy_self() + + # Provenance: a report is only actionable if the reader can tell WHICH copy + # of this script produced it. A normal run carries it in the header, --json + # carries it for machine consumers, and the cron ALERT line carries it on + # failure. --quiet is documented as "only output on failure", so the header + # is emitted only when not quiet and a healthy quiet run stays silent. + script_path = os.path.abspath(__file__) + cwd = os.getcwd() + + if quiet: + captured = io.StringIO() + with contextlib.redirect_stdout(captured): + load_agent_keys() + _run_checks() + if FAIL: + sys.stdout.write(captured.getvalue()) + else: + print(f"🏥 Agent Health Check v4 — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}") + print(f"📍 executed from: script={script_path} cwd={cwd}") + print() + load_agent_keys() + _run_checks() + if FAIL: print(f"\n❌ {len(FAIL)} FAILURE(S): {' | '.join(FAIL)}") if quiet: - print(f"ALERT agent-health:{','.join(FAIL)}") + print(f"ALERT agent-health:{','.join(FAIL)} script={script_path} cwd={cwd}") elif not quiet: print("\n✅ All checks passed") if as_json: print(json.dumps({"timestamp": datetime.now().isoformat(), - "failures": FAIL, "healthy": len(FAIL) == 0})) + "execution_path": script_path, "cwd": cwd, + "failures": FAIL, "report_only": REPORT_ONLY, + "healthy": len(FAIL) == 0})) sys.exit(1 if FAIL else 0) diff --git a/scripts/daily-infra-report.py b/scripts/daily-infra-report.py index 106b2f8..3386d4d 100755 --- a/scripts/daily-infra-report.py +++ b/scripts/daily-infra-report.py @@ -328,26 +328,13 @@ def collect(): "updated_at": "", } - # Mumuni (CT 100, IP 192.168.68.24) - mumuni_state = ssh("192.168.68.24", "cat ~/.hermes/gateway_state.json 2>/dev/null") - mumuni_data = {} - try: - mumuni_data = json.loads(mumuni_state) if mumuni_state else {} - except: - mumuni_data = {} - mumuni_platforms = mumuni_data.get("platforms", {}) - report["agents"]["mumuni"] = { - "platform": "hermes", "ct": 100, "ip": "192.168.68.24", - "gateway_state": mumuni_data.get("gateway_state", "unknown"), - "telegram_state": mumuni_platforms.get("telegram", {}).get("state", "unknown"), - "zulip_state": mumuni_platforms.get("zulip", {}).get("state", "not_installed"), - "email_state": mumuni_platforms.get("email", {}).get("state", "unknown"), - "hermes_version": "", - } - # Get Hermes version - ver = ssh("192.168.68.24", "hermes --version 2>/dev/null | head -1") - if ver: - report["agents"]["mumuni"]["hermes_version"] = ver.split("·")[0].replace("Hermes Agent ","").strip() + # Mumuni is deliberately absent from this digest: captain ruling 2026-09-10. + # She moved off this host onto her own container (kagentz CT 105 on minipve, + # 192.168.68.14, dedicated `hermes` user) and is monitored from her side. The + # former probe ssh'd to 192.168.68.24 for the decommissioned deployment's + # ~/.hermes/gateway_state.json, always read "unknown", and published a false + # "mumuni:unknown" line in the agent table and the gateway-unknown issue + # count of every digest. Do NOT re-add an .24 / gateway_state probe. return report @@ -561,12 +548,6 @@ Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count'] zulip_state = "✅" if agent.get("zulip_state") == "connected" else ("❌" if agent.get("zulip_state") == "disconnected" else "⬜") gateway = agent.get("gateway_state", "?") processed = "DSH" - elif name == "mumuni": - zulip_state = "⬜" if agent.get("zulip_state") == "not_installed" else ("✅" if agent.get("zulip_state") == "connected" else "⬜") - gateway = agent.get("gateway_state", "?") - tg = "✅" if agent.get("telegram_state") == "connected" else "❌" - ver = agent.get("hermes_version", "") - processed = f"TG:{tg} v{ver}" else: zulip_state = "⬜" gateway = agent.get("gateway_state", "?") @@ -720,6 +701,6 @@ if __name__ == "__main__": print(f" Zulip Ext: {'✅' if report.get('zulip_ext',{}).get('connected') else '❌'}") print(f" LiteLLM: {sum(1 for c in report.get('litellm',{}).get('checks',[]) if c['status']=='pass')}/{len(report.get('litellm',{}).get('checks',[]))} checks pass") agent_parts = [] -for k,v in report.get('agents',{}).items(): - agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}") -print(f" Agents: {', '.join(agent_parts)}") + for k,v in report.get('agents',{}).items(): + agent_parts.append(f"{k}:{v.get('gateway_state',v.get('pm2_status','?'))}") + print(f" Agents: {', '.join(agent_parts)}") diff --git a/scripts/prose-lint.sh b/scripts/prose-lint.sh index 6874267..1756856 100755 --- a/scripts/prose-lint.sh +++ b/scripts/prose-lint.sh @@ -57,7 +57,7 @@ echo "── 2. Regression detection ──" # Grafana /grafana/ as nginx route or URL path (reverted 2026-07-02) # EXCLUDE: filesystem paths (/opt/monitoring/grafana/...), directory creation, revert docs -GRAFANA_HITS=$(grep -rn '/grafana/' *.prose.md 2>/dev/null \ +GRAFANA_HITS=$(grep -rn '/grafana/' ./*.prose.md 2>/dev/null \ | grep -v '/opt/monitoring/grafana/' \ | grep -v 'was tried and reverted\|was reverted\|do not re-add\|NOT recommended' \ | grep -v 'mkdir.*grafana\|Create.*grafana' \ @@ -71,7 +71,7 @@ else fi # Stale CT IDs (CT 122, CT 123 as CT IDs — not IPs .122, .123) -CT_STALE=$(grep -rn '\bCT 122\b' *.prose.md 2>/dev/null || true) +CT_STALE=$(grep -rn '\bCT 122\b' ./*.prose.md 2>/dev/null || true) if [ -n "$CT_STALE" ]; then echo " ❌ REGRESSION: CT 122 used as CT ID — Tanko is CT 112" echo "$CT_STALE" @@ -84,6 +84,35 @@ fi # .122/.123 are correct — verified reachable bridge IPs for Tanko/Mumuni echo " ✅ IP consistency verified (.19=.122=.123 all reachable)" +# Report provenance — every contract report must state the absolute path it +# executed from, so a stale-consumer report is distinguishable from a real fault +# at read time (2026-09-09 probe-drift incident: three false DEGRADED rounds). +# Enforced only inside the **Report format** paragraph, and a check-health +# contract with no Report format paragraph FAILs rather than being skipped. +PROV_FILES=$(grep -rlE '^### check-health|\*\*Report format\*\*' ./*.prose.md 2>/dev/null || true) +if [ -z "$PROV_FILES" ]; then + echo " ❌ No check-health/report-format contracts found — provenance not enforced" + FAILED=1 +else + PROV_BAD=0 + while IFS= read -r f; do + [ -n "$f" ] || continue + REPORT_PARA=$(awk '/\*\*Report format\*\*/{found=1} found{print} found && /^[[:space:]]*$/{exit}' "$f") + if [ -z "$REPORT_PARA" ]; then + echo " ❌ $f: check-health contract has no **Report format** paragraph" + PROV_BAD=1 + elif ! printf '%s\n' "$REPORT_PARA" | grep -qE 'absolute path|pwd -P|executed from'; then + echo " ❌ $f: **Report format** lacks execution provenance (absolute path / pwd -P)" + PROV_BAD=1 + fi + done <<< "$PROV_FILES" + if [ "$PROV_BAD" -eq 1 ]; then + FAILED=1 + else + echo " ✅ Report provenance present in all report-format contracts" + fi +fi + # ── 3. Cross-contract consistency ── echo "" echo "── 3. Cross-contract consistency ──" diff --git a/scripts/zulip-monitor.sh b/scripts/zulip-monitor.sh index ab4d4cd..32b884a 100755 --- a/scripts/zulip-monitor.sh +++ b/scripts/zulip-monitor.sh @@ -1,7 +1,10 @@ #!/bin/bash # /root/scripts/zulip-monitor.sh — Zulip Mesh Health Monitor -# Implements zulip-health.prose.md v2 -# Runs every 15 min via cron. Alerts via Telegram. +# Implements zulip-health.prose.md v3 +# Runs every 15 min via cron. Alerts: Zulip private DM to the owner plus a stream post to #agent-hub on topic 'zulip-health'. +# Legs: global Zulip server, Platform A pi/Abiba (the Zulip bridge), Platform B +# Tanko (DSH), Platform C Agent Zero (kagentz). The former Platform B Hermes +# agent leg is retired — see the note after the Tanko leg. set -euo pipefail ZULIP_SITE="https://chat.sysloggh.net" @@ -21,7 +24,8 @@ notify() { # Zulip DM to owner local content="${severity} Zulip Monitor: ${msg}" - local form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")" + local form + form="type=private&to=%5B${OWNER_ZULIP_ID}%5D&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote('''${content}'''))")" curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \ -u "${ZULIP_EMAIL}:${ZULIP_KEY}" \ -d "${form}" > /dev/null 2>&1 || true @@ -29,8 +33,9 @@ notify() { local stream_content="${severity} Zulip Monitor: ${msg}" curl -sf -X POST "${ZULIP_SITE}/api/v1/messages" \ -u "${ZULIP_EMAIL}:${ZULIP_KEY}" \ - -d "type=stream\&to=%5B7%5D\&topic=zulip-health\&content=$(python3 -c "import urllib.parse; print(urllib.parse.quote(str()))")" \ - > /dev/null 2>&1 || true + -d "type=stream&to=%5B7%5D&topic=zulip-health&content=$(printf '%s' "${stream_content}" | python3 -c "import sys,urllib.parse; print(urllib.parse.quote_from_bytes(sys.stdin.buffer.read()))")" \ + > /dev/null 2>&1 \ + || echo " WARN: stream alert to #agent-hub (zulip-health) delivery failed (curl exit $?)" >> "$LOG" } # ── Global: Zulip Server ── @@ -45,62 +50,109 @@ else fi # ── Platform A: pi (Abiba) ── -PI_HEALTH=$(curl -sf --connect-timeout 5 http://localhost:9200/health 2>/dev/null || echo "{}") -PI_CONNECTED=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('connected',False))" 2>/dev/null) -PI_ERROR=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('last_error') or '')" 2>/dev/null) -PI_RETRIES=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('retry_count',0))" 2>/dev/null) +# Probes the pi Zulip extension health endpoint (:9200/health, served by the +# extension's startHealthServer; shape documented in zulip-health.prose.md). +# FAIL-SAFE contract (pinned by tests/zulip-monitor-abiba.sh): connection state +# lives NESTED at zulip.connected / zulip.last_error — there is no top-level +# `connected` and no retry counter in the payload. A fetch error, non-2xx +# response, empty/unparseable body, or payload missing a boolean +# zulip.connected is a PROBE FAILURE: it alerts and NEVER calls pm2 restart. +# pm2 restart runs ONLY on affirmative zulip.connected=false. +# -- abiba-leg-start (verbatim-extracted by tests/zulip-monitor-abiba.sh) +PI_HTTP=$(curl -s -o /dev/null --connect-timeout 5 --max-time 10 -w '%{http_code}' http://localhost:9200/health 2>/dev/null || echo "000") +PI_BODY=$(curl -s --connect-timeout 5 --max-time 10 http://localhost:9200/health 2>/dev/null || true) +PI_STATE=$(printf '%s' "$PI_BODY" | python3 -c ' +import sys, json +code = sys.argv[1] +body = sys.stdin.read() +try: + d = json.loads(body) +except Exception: + sys.stdout.write("probe-failed|unparseable body") + sys.exit(0) +if not code.startswith("2"): + sys.stdout.write("probe-failed|HTTP %s" % code) + sys.exit(0) +if not isinstance(d, dict) or not isinstance(d.get("zulip"), dict): + sys.stdout.write("probe-failed|missing zulip.connected") + sys.exit(0) +z = d["zulip"] +if "connected" not in z or not isinstance(z["connected"], bool): + sys.stdout.write("probe-failed|missing or non-boolean zulip.connected") + sys.exit(0) +err = z.get("last_error") or "" +if z["connected"]: + if err: + sys.stdout.write("degraded|%s" % err) + else: + sys.stdout.write("healthy|%s" % z.get("messages_processed", 0)) +else: + sys.stdout.write("disconnected|") +' "$PI_HTTP" 2>/dev/null) || PI_STATE="probe-failed|python error" +PI_VERDICT=${PI_STATE%%|*} +PI_DETAIL=${PI_STATE#*|} -if [ "$PI_CONNECTED" != "True" ]; then - notify "🔴" "Abiba pi extension DISCONNECTED — restarting" - pm2 restart abiba-zulip 2>/dev/null || true +case "$PI_VERDICT" in + healthy) + echo " Abiba: ✅ Connected (processed=$PI_DETAIL)" >> "$LOG" ;; + degraded) + notify "🟡" "Abiba pi extension error: ${PI_DETAIL:0:100}" + echo " Abiba: 🟡 Error: ${PI_DETAIL:0:100}" >> "$LOG" ;; + disconnected) + notify "🔴" "Abiba pi extension DISCONNECTED — restarting" + pm2 restart abiba-zulip 2>/dev/null || true + ISSUES=$((ISSUES + 1)) + echo " Abiba: ❌ Disconnected — restarted" >> "$LOG" ;; + probe-failed) + notify "🟠" "Abiba pi extension health probe FAILED (${PI_DETAIL}; HTTP $PI_HTTP) — NOT restarting, manual check needed" + ISSUES=$((ISSUES + 1)) + echo " Abiba: ⚠️ Probe failed (${PI_DETAIL}; HTTP $PI_HTTP) — NOT restarted" >> "$LOG" ;; + *) + notify "🟠" "Abiba pi extension health probe returned unexpected verdict (${PI_STATE}) — NOT restarting, manual check needed" + ISSUES=$((ISSUES + 1)) + echo " Abiba: ⚠️ Unexpected probe verdict (${PI_STATE}) — NOT restarted" >> "$LOG" ;; +esac +# -- abiba-leg-end + +# ── Platform B: Tanko (DSH dsh-web on amdpve CT 112) ── +# Direct SSH to 192.168.68.122 is not a dependency of this monitor — per-worker +# key availability varies — so probes run from the amdpve vantage via `pct exec`. +# Tanko's Zulip gateway runs as the dsh-web systemd unit inside CT 112 on amdpve +# (192.168.68.15). The gateway binds 127.0.0.1:3080 loopback-only by design — a +# remote :3080 probe is refused and is NOT a fault. +TANKO_SVC=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \ + "pct exec 112 -- systemctl is-active dsh-web" 2>/dev/null || true) +[ -n "$TANKO_SVC" ] || TANKO_SVC="unknown" +TANKO_HTTP=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.15 \ + "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/" 2>/dev/null || true) +[ -n "$TANKO_HTTP" ] || TANKO_HTTP="000" + +if [ "$TANKO_SVC" != "active" ]; then + notify "🔴" "Tanko (DSH dsh-web) service state: $TANKO_SVC — needs restart" ISSUES=$((ISSUES + 1)) - echo " Abiba: ❌ Disconnected — restarted" >> "$LOG" -elif [ -n "$PI_ERROR" ]; then - notify "🟡" "Abiba pi extension error: ${PI_ERROR:0:100}" - echo " Abiba: 🟡 Error: ${PI_ERROR:0:100}" >> "$LOG" -elif [ "$PI_RETRIES" -ge 3 ]; then - notify "🟡" "Abiba pi extension: $PI_RETRIES retries — restarting" - pm2 restart abiba-zulip 2>/dev/null || true - echo " Abiba: 🟡 $PI_RETRIES retries — restarted" >> "$LOG" + echo " Tanko: ❌ service=$TANKO_SVC" >> "$LOG" +elif [ "$TANKO_HTTP" = "000" ]; then + notify "🔴" "Tanko (DSH dsh-web) HTTP :3080 connection refused/timeout — needs restart" + ISSUES=$((ISSUES + 1)) + echo " Tanko: ❌ http=000 (refused/timeout)" >> "$LOG" else - echo " Abiba: ✅ Connected (processed=$(echo "$PI_HEALTH" | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('messages_processed',0))" 2>/dev/null))" >> "$LOG" + case "$TANKO_HTTP" in + 200|301|302|307|308|401|403) + echo " Tanko: ✅ service=active http=$TANKO_HTTP" >> "$LOG" ;; + *) + notify "🟡" "Tanko (DSH dsh-web) HTTP :3080 answered $TANKO_HTTP — running, unexpected status" + echo " Tanko: 🟡 service=active http=$TANKO_HTTP (running, warning)" >> "$LOG" ;; + esac fi -# ── Platform B: Hermes (Tanko) ── -TANKO_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 jerome@192.168.68.122 \ - "cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}") -TANKO_ZULIP=$(echo "$TANKO_STATE" | python3 -c " -import sys,json -d=json.load(sys.stdin) -p=d.get('platforms',{}).get('zulip',{}) -print(p.get('state','unknown')) -" 2>/dev/null) - -if [ "$TANKO_ZULIP" != "connected" ]; then - notify "🔴" "Tanko (Hermes) Zulip state: $TANKO_ZULIP — needs restart" - ISSUES=$((ISSUES + 1)) - echo " Tanko: ❌ state=$TANKO_ZULIP" >> "$LOG" -else - echo " Tanko: ✅ Zulip connected" >> "$LOG" -fi - -# ── Platform B: Hermes (Mumuni) ── -MUMUNI_STATE=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.24 \ - "cat ~/.hermes/gateway_state.json 2>/dev/null" 2>/dev/null || echo "{}") -MUMUNI_ZULIP=$(echo "$MUMUNI_STATE" | python3 -c " -import sys,json -d=json.load(sys.stdin) -p=d.get('platforms',{}).get('zulip',{}) -print(p.get('state','unknown')) -" 2>/dev/null) - -if [ "$MUMUNI_ZULIP" != "connected" ]; then - notify "🔴" "Mumuni (Hermes) Zulip state: $MUMUNI_ZULIP" - ISSUES=$((ISSUES + 1)) - echo " Mumuni: ❌ state=$MUMUNI_ZULIP" >> "$LOG" -else - echo " Mumuni: ✅ Zulip connected" >> "$LOG" -fi +# ── Removed: the former "Platform B: Hermes" agent leg ── +# Captain ruling 2026-09-10: that agent moved off this host onto her own +# container (kagentz CT 105 on minipve, dedicated `hermes` user) and is now +# monitored on her side — see the out-of-scope note in zulip-health.prose.md. +# The old leg ssh'd to her former CT 100 deployment and read its Hermes gateway +# state, which reported "unknown" on every run and posted a false 🔴 DM plus an +# #agent-hub stream alert. Do NOT re-add a probe for her: this monitor must +# never contact her former host. # ── Platform C: Agent Zero (kagentz) ── AZ_A2A=$(ssh -o StrictHostKeyChecking=no -o ConnectTimeout=5 root@192.168.68.14 \ diff --git a/tests/fixtures/zulip-health-connected.json b/tests/fixtures/zulip-health-connected.json new file mode 100644 index 0000000..61ecc6b --- /dev/null +++ b/tests/fixtures/zulip-health-connected.json @@ -0,0 +1,25 @@ +{ + "status": "ok", + "platform": "pi", + "agent": "abiba", + "zulip": { + "connected": true, + "site": "https://chat.sysloggh.net", + "email": "abiba-bot@chat.sysloggh.net", + "queue_id": "ee7f8b6d-9d53-48a7-ad58-f6e999771001", + "bot_user_id": 21, + "messages_processed": 0, + "skipped": 0, + "last_error": null + }, + "circuit_breaker": { + "state": "CLOSED", + "failures": 0, + "successes": 5, + "totalRequests": 5, + "failureRate": "0.000", + "openedAt": null + }, + "workers": [], + "worker_count": 0 +} diff --git a/tests/fixtures/zulip-health-disconnected.json b/tests/fixtures/zulip-health-disconnected.json new file mode 100644 index 0000000..0296d77 --- /dev/null +++ b/tests/fixtures/zulip-health-disconnected.json @@ -0,0 +1,25 @@ +{ + "status": "down", + "platform": "pi", + "agent": "abiba", + "zulip": { + "connected": false, + "site": "https://chat.sysloggh.net", + "email": "abiba-bot@chat.sysloggh.net", + "queue_id": null, + "bot_user_id": null, + "messages_processed": 0, + "skipped": 0, + "last_error": "Zulip API error 401: queue registration failed" + }, + "circuit_breaker": { + "state": "CLOSED", + "failures": 0, + "successes": 0, + "totalRequests": 0, + "failureRate": "0.000", + "openedAt": null + }, + "workers": [], + "worker_count": 0 +} diff --git a/tests/test_mumuni_monitor_removal.py b/tests/test_mumuni_monitor_removal.py new file mode 100644 index 0000000..5f9b9eb --- /dev/null +++ b/tests/test_mumuni_monitor_removal.py @@ -0,0 +1,352 @@ +"""Regression tests for the 2026-09-10 retirement of the Mumuni monitoring leg. + +WHY THIS FILE EXISTS: captain ruling 2026-09-10 — Mumuni moved off this host +onto her own container (kagentz CT 105 on minipve, 192.168.68.14, dedicated +`hermes` user) and is monitored from her side. The monitor nevertheless kept +ssh'ing to root@192.168.68.24 for `~/.hermes/gateway_state.json` on the +decommissioned deployment, read "unknown" on every run, and posted a false 🔴 +"Mumuni (Hermes) Zulip state: unknown" DM + #agent-hub stream alert to the +captain. The daily infra digest published a matching `mumuni:unknown` row. + +CONTRACT UNDER TEST: + * `scripts/zulip-monitor.sh` carries NO Mumuni probe and NO 192.168.68.24 + reference; it never ssh'es .24, and even on a failing run it emits no Mumuni + notify (stdout alert, Zulip payload, or log line). + * The Abiba (pi — the Zulip bridge), Tanko (DSH) and Agent Zero (kagentz) legs + still work: deleting the Mumuni leg must not have gutted the rest. + * `scripts/daily-infra-report.py` no longer probes .24 for a Hermes gateway + state and no longer emits a `mumuni` agent entry. + * `scripts/agent-health-check.py`'s AGENTS roster has no mumuni entry. This is + a pin, not a behavior change — verify the probe was already gone. + * `zulip-health.prose.md` retires the Mumuni-only steps and says explicitly + that Mumuni is not monitored from this host. + +HOW: behavioral execution plus one named deliverable-text contract. The sandbox +copies the shipped monitor verbatim and rewrites only its LOG constant, then +runs it with stub ssh/curl on PATH; the ssh stub records every host it is asked +to reach, so "never probes .24" and "no Mumuni notify" are asserted from +observed behavior. The daily digest is pinned by importing it and exercising +collect() and build_html() directly. The single source-text assertion is the +deliverable-text contract the captain acceptance names for the shipped monitor. + +Usage: python3 -m pytest tests/test_mumuni_monitor_removal.py +""" +from __future__ import annotations + +import importlib.util +import os +import pathlib +import stat +import subprocess + +import pytest + +ROOT = pathlib.Path(__file__).resolve().parents[1] +ZULIP_MONITOR = ROOT / "scripts" / "zulip-monitor.sh" +DAILY_REPORT = ROOT / "scripts" / "daily-infra-report.py" +AHC = ROOT / "scripts" / "agent-health-check.py" +HEALTH_CONTRACT = ROOT / "zulip-health.prose.md" +CONNECTED_FIXTURE = ROOT / "tests" / "fixtures" / "zulip-health-connected.json" + +MUMUNI_IP = "192.168.68.24" # Mumuni's old (decommissioned) deployment +TANKO_VANTAGE = "192.168.68.15" # amdpve — Tanko CT 112 via pct exec +AGENT_ZERO_HOST = "192.168.68.14" # kagentz host, Agent Zero docker + + +# ── scripts/zulip-monitor.sh: deliverable-text contract ───────────── + +def test_zulip_monitor_deliverable_text_contract(): + """Owned deliverable-text contract for scripts/zulip-monitor.sh. + + Captain acceptance requires the shipped monitor to contain no Mumuni probe + identifier and no 192.168.68.24 literal. Behavioral proof that the monitor + never contacts that host and never emits a Mumuni notify lives in the + sandbox tests below; this only pins the named text contract. + """ + text = ZULIP_MONITOR.read_text() + assert "mumuni" not in text.lower() + assert MUMUNI_IP not in text + + +# ── scripts/zulip-monitor.sh: behavioral sandbox ───────────────────── + +SSH_STUB = r"""#!/usr/bin/env bash +# Stub ssh: record the target host, then answer by host + remote command. +printf '%s\n' "$*" >> "$RECORD_DIR/ssh.calls" +host="" +for a in "$@"; do + case "$a" in + *@192.168.*) host="${a##*@}" ;; + esac +done +printf '%s\n' "$host" >> "$RECORD_DIR/ssh.hosts" +cmd="${*: -1}" +case "$host" in + 192.168.68.15) + case "$cmd" in + *"systemctl is-active"*) printf '%s' "$TANKO_SVC" ;; + *curl*) printf '%s' "$TANKO_HTTP" ;; + esac ;; + 192.168.68.14) + case "$cmd" in + *agent.json*) printf '%s' "$AZ_A2A" ;; + *"ps aux"*) printf '%s\n' "$AZ_PS" ;; + esac ;; + *) + printf 'UNEXPECTED-SSH-HOST %s\n' "$host" >> "$RECORD_DIR/unexpected-ssh" ;; +esac +exit 0 +""" + +CURL_STUB = r"""#!/usr/bin/env bash +# Stub curl: serve the Abiba health fixture and the Zulip server 200, and +# record every call (including notify) payloads. +printf '%s\n' "$*" >> "$RECORD_DIR/curl.calls" +case "$*" in + *:9200/health*) + case " $* " in + *" -w "*) printf '%s' "$PI_HTTP" ;; # -w '%{http_code}' probe + *) printf '%s' "$PI_BODY" ;; # body probe + esac ;; + *server_settings*) + printf '%s' "$SERVER_HTTP" ;; +esac +exit 0 +""" + + +def _write_exec(path: pathlib.Path, body: str) -> None: + path.write_text(body) + path.chmod(path.stat().st_mode + | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH) + + +def _run_monitor(tmp_path, *, tanko_svc="active", tanko_http="200", + az_a2a='{"name":"kagentz"}', + az_ps="root 111 0.1 0.2 /opt/venv-a0/bin/python3 -u adapter.py"): + """Run the shipped monitor in a sandbox; return (proc, record_dir, log_path). + + Only the LOG constant is rewritten (to keep the run inside the worktree). + Everything else — legs, labels, notify logic — is the shipped script. + """ + sandbox = tmp_path / "sandbox" + bindir = sandbox / "bin" + record = sandbox / "record" + bindir.mkdir(parents=True) + record.mkdir() + + _write_exec(bindir / "ssh", SSH_STUB) + _write_exec(bindir / "curl", CURL_STUB) + + source = ZULIP_MONITOR.read_text() + log_line = 'LOG="/root/zulip-health-monitor.log"' + assert log_line in source, "LOG constant moved — update the sandbox harness" + log_path = sandbox / "zulip-health-monitor.log" + script = sandbox / "zulip-monitor.sh" + script.write_text(source.replace(log_line, f'LOG="{log_path}"')) + + env = dict(os.environ) + env.update({ + "PATH": f"{bindir}:{env['PATH']}", + "RECORD_DIR": str(record), + "TANKO_SVC": tanko_svc, + "TANKO_HTTP": tanko_http, + "AZ_A2A": az_a2a, + "AZ_PS": az_ps, + "PI_HTTP": "200", + "PI_BODY": CONNECTED_FIXTURE.read_text(), + "SERVER_HTTP": "200", + }) + proc = subprocess.run(["bash", str(script)], cwd=sandbox, env=env, + capture_output=True, text=True) + return proc, record, log_path + + +def test_healthy_run_is_quiet_and_never_reaches_mumuni(tmp_path): + proc, record, log_path = _run_monitor(tmp_path) + assert proc.returncode == 0, proc.stderr + log = log_path.read_text() + + # Every retained leg actually ran and passed. + assert "Server: ✅ HTTP 200" in log + assert "Abiba: ✅ Connected" in log + assert "Tanko: ✅ service=active http=200" in log + assert "kagentz: ✅ A2A alive" in log + assert "kagentz: ✅ Adapter running" in log + assert "Result: ✅ All healthy" in log + + # A healthy run emits no notify at all — and certainly no Mumuni one. + assert proc.stdout == "" + assert "Mumuni" not in log + assert "🔴" not in log + + # Observed behavior: .24 is never resolved, only Tanko's vantage and the + # Agent Zero host are contacted. + hosts = record.joinpath("ssh.hosts").read_text().split() + assert MUMUNI_IP not in hosts + assert set(hosts) == {TANKO_VANTAGE, AGENT_ZERO_HOST} + assert not record.joinpath("unexpected-ssh").exists() + + +def test_failing_run_alerts_on_tanko_but_never_on_mumuni(tmp_path): + # Failure path: exercises notify() end to end so "no Mumuni notify" is + # proven on the alert path, not only on the quiet healthy path. + proc, record, log_path = _run_monitor(tmp_path, tanko_svc="inactive", + tanko_http="000") + assert proc.returncode == 0, proc.stderr + + alerts = proc.stdout + assert "Tanko (DSH dsh-web) service state: inactive" in alerts + assert "1 issue(s) found" in alerts + + # No Mumuni text in stdout, the log, or any Zulip DM/stream payload. + assert "Mumuni" not in alerts + assert "Mumuni" not in log_path.read_text() + assert MUMUNI_IP not in alerts + log_path.read_text() + payloads = record.joinpath("curl.calls").read_text() + assert "Mumuni" not in payloads + assert MUMUNI_IP not in payloads + + # The rest of the monitor still ran alongside the failing Tanko leg. + log = log_path.read_text() + assert "Abiba: ✅ Connected" in log + assert "kagentz: ✅ A2A alive" in log + assert "Result: 🔴 1 issue(s) found" in log + + +# ── scripts/daily-infra-report.py: behavioral digest checks ────────── + +@pytest.fixture(scope="module") +def daily(): + spec = importlib.util.spec_from_file_location("daily_infra_report", DAILY_REPORT) + assert spec and spec.loader + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +DAILY_AGENTS = { + "abiba": { + "platform": "pi", "ct": 100, "ip": MUMUNI_IP, + "zulip_connected": True, "zulip_processed": 5, + "pm2_status": "online", "pm2_restarts": "0", "pm2_uptime": "1h", + }, + "tanko": { + "platform": "dsh", "ct": 112, "ip": "192.168.68.122", + "gateway_state": "n/a (DSH)", "zulip_state": "connected", + "telegram_state": "unknown", "gateway_pid": None, "updated_at": "", + }, +} + + +def _fabricated_report(agents): + return { + "nodes": {}, + "node_count": 1, + "nodes_online": 1, + "total_vms": 0, + "running_vms": 0, + "stopped_vms": [], + "vms_by_node": {n: [] for n in + ["amdpve", "minipve", "storepve", "acerpve", "ocupve"]}, + "storage": [], + "docker_vm": {"total": 0, "running": 0, "unhealthy": [], + "containers": [], "reclaimable": "", "disk_used": "1%"}, + "docker_syslog": {"total": 0, "running": 0, "containers": []}, + "docker_netbird": {"total": 0, "running": 0, "containers": []}, + "endpoints": [], + "litellm": {"checks": []}, + "nfs": [], + "zulip_ext": { + "connected": True, "queue_id": "queue", "last_error": None, + "messages_processed": 0, "retry_count": 0, "pm2": {}, + "pm2_healthy": True, "bot_skipped_15min": 0, "finalized_1h": 0, + "failed_finalize_1h": 0, "finalize_fail_pct": 0, + "server_status": "200", + }, + "agents": agents, + } + + +def _agent_status_card(html): + start = html.index("🤖 Agent Status") + end = html.index("💬 Zulip Extension") + return html[start:end] + + +def test_daily_report_renders_only_abiba_and_tanko_agents(daily): + """build_html() over a Mumuni-free agent set must render no Mumuni row and + no Mumuni gateway-unknown issue, while abiba and tanko rows still render.""" + html = daily.build_html(_fabricated_report(dict(DAILY_AGENTS))) + card = _agent_status_card(html) + assert "mumuni" not in card.lower() + assert "abiba" in card + assert "tanko" in card + assert "mumuni" not in html.lower() + + +def test_daily_report_collect_never_probes_mumuni(monkeypatch, daily): + """collect() with ssh stubbed must add no mumuni agent and must never ssh + its decommissioned .24 host.""" + probed = [] + + class _NoSubprocess: + @staticmethod + def check_output(*args, **kwargs): + return b"" + + def fake_ssh(host, cmd): + probed.append(host) + return "" + + monkeypatch.setattr(daily, "pve_get", lambda path: []) + monkeypatch.setattr(daily, "ssh_jerome", lambda host, cmd: "") + monkeypatch.setattr(daily, "ssh", fake_ssh) + monkeypatch.setattr(daily, "http_get", + lambda url, auth=None, timeout=10: "200") + monkeypatch.setattr(daily, "http_get_body", + lambda url, auth=None, timeout=10: "") + monkeypatch.setattr(daily, "count_in_log", lambda *a, **k: 0) + monkeypatch.setattr(daily, "subprocess", _NoSubprocess) + + report = daily.collect() + assert "mumuni" not in report["agents"] + assert MUMUNI_IP not in probed + + +# ── scripts/agent-health-check.py: roster pin ─────────────────────── + +@pytest.fixture(scope="module") +def ahc(): + spec = importlib.util.spec_from_file_location("agent_health_check_roster", AHC) + assert spec and spec.loader + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def test_agent_health_roster_has_no_mumuni_entry(ahc): + assert "mumuni" not in ahc.AGENTS + + +# ── zulip-health.prose.md: contract reconciliation ────────────────── + +def test_health_contract_retires_mumuni_only_steps(): + text = HEALTH_CONTRACT.read_text() + assert MUMUNI_IP not in text + for step in ("**B4:", "**B5:", "**B6:"): + assert step not in text + + +def test_health_contract_states_mumuni_is_not_monitored_from_this_host(): + text = HEALTH_CONTRACT.read_text() + assert "Mumuni is NOT monitored from this host" in text + assert "monitored on her side" in text + assert "her own container" in text + + +def test_health_contract_keeps_tanko_agent_zero_and_bridge_steps(): + text = HEALTH_CONTRACT.read_text() + for marker in ("**B1:", "**B2:", "**B3:", "Step 4: Platform C", + "Step 2: Platform A", "Step 1: Zulip Server Liveness"): + assert marker in text, marker diff --git a/tests/test_probe_drift.py b/tests/test_probe_drift.py new file mode 100644 index 0000000..28a4e47 --- /dev/null +++ b/tests/test_probe_drift.py @@ -0,0 +1,466 @@ +"""Regression tests for the 2026-09-09/10 probe-drift corrections. + +WHY THIS FILE EXISTS: the monitoring contracts kept emitting false alarms from +stale expectations rather than live faults. + + * agent-health-check v3 reported 6 failures that were all stale expectations: + abiba (pi-only since the harness purge) was tested as a Hermes host, koby + (report-only per the captain's 2026-08-17 ruling) was counted as repairable, + koby's CT 111 was probed on amdpve where it does not exist (it runs on + storepve .6), and the wrapper infisical check had two bugs — it read only + the first 20 lines, so koonimo's wrapper (which references /usr/bin/infisical + past line 20) false-failed, and it treated koby's genuine no-infisical + (~/.hermes/.env) wrapper as broken. + * gpu-monitor emitted "DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000" three + times from probing bare port 80 on GPU hosts while :8080 answered 200. + * infrastructure-monitoring probed CT 116 for the PVE API (no pveproxy -> + 000) instead of the five real cluster nodes, which answer 401 = alive. + +These tests execute the health script (with SSH/vault stubbed) and the real +provenance consumer (scripts/prose-lint.sh), and parse the contracts' executable +check-health probe blocks into normalized probe sets. No live network, vault, or +SSH access is required. +""" +from __future__ import annotations + +import importlib.util +import json +import os +import pathlib +import re +import subprocess +import sys +import textwrap + +import pytest + +ROOT = pathlib.Path(__file__).resolve().parents[1] +AHC = ROOT / "scripts" / "agent-health-check.py" +LINT = ROOT / "scripts" / "prose-lint.sh" +GPU = ROOT / "gpu-monitor.prose.md" +INFRA = ROOT / "infrastructure-monitoring.prose.md" + +PVE_NODE_IPS = { + "192.168.68.9", + "192.168.68.5", + "192.168.68.15", + "192.168.68.6", + "192.168.68.12", +} + + +@pytest.fixture(scope="module") +def ahc(): + """Import agent-health-check.py without live network/SSH side effects.""" + spec = importlib.util.spec_from_file_location("agent_health_check", AHC) + assert spec and spec.loader + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +# ── helpers: execute the health script with SSH/vault stubbed ───────── + +def _run_main(ahc, monkeypatch, capsys, argv, ssh_result=None): + ahc.FAIL.clear() + ahc.REPORT_ONLY.clear() + monkeypatch.setattr(ahc, "load_agent_keys", lambda: None) + monkeypatch.setattr(ahc, "ssh", lambda *a, **k: ssh_result) + monkeypatch.setattr(sys, "argv", ["agent-health-check.py", "--no-deploy", *argv]) + with pytest.raises(SystemExit) as exc: + ahc.main() + return exc.value.code, capsys.readouterr().out + + +def _json_payload(out): + for line in reversed(out.splitlines()): + if line.startswith('{"timestamp"'): + return json.loads(line) + raise AssertionError(f"no JSON payload in output:\n{out}") + + +# ── agent-health-check: stale-expectation legs ─────────────────────── + +def test_import_does_not_contact_vault(ahc): + # Keys are loaded in main() via load_agent_keys(); importing must stay inert. + assert callable(ahc.load_agent_keys) + assert all(agent.get("key") is None for agent in ahc.AGENTS.values()) + + +def test_abiba_is_pi_only_runtime(ahc): + # .24 has run pi-only since the harness purge: no Hermes gateway, config, or + # wrapper. Probing those legs produced false failures. + assert ahc.AGENTS["abiba"]["runtime"] == "pi" + + +def test_koby_is_report_only(ahc): + # Captain's 2026-08-17 ruling (Rule 17): detect and report, never repair. + assert ahc.AGENTS["koby"]["report_only"] is True + + +def test_koby_ct111_is_on_storepve(ahc): + # Live-verified 2026-09-10: `pct status 111` = running on storepve (.6); + # amdpve has no lxc/111.conf, which is what false-failed before. + assert ahc.AGENTS["koby"]["pve"] == "storepve" + + +def test_report_only_legs_never_count_as_failures(ahc): + for agent, report_only in (("koby", True), ("koonimo", False), ("tanko", False)): + ahc.FAIL.clear() + ahc.REPORT_ONLY.clear() + ahc._fail(f"probe:{agent}", agent) + if report_only: + assert ahc.FAIL == [] + assert ahc.REPORT_ONLY == [f"probe:{agent}"] + else: + assert ahc.FAIL == [f"probe:{agent}"] + assert ahc.REPORT_ONLY == [] + ahc.FAIL.clear() + ahc.REPORT_ONLY.clear() + + +def test_failure_recording_accepts_agentless_keys(ahc): + ahc.FAIL.clear() + try: + ahc._fail("gpu-no-port:gpu-rtx3090 (.8)") + assert ahc.FAIL == ["gpu-no-port:gpu-rtx3090 (.8)"] + finally: + ahc.FAIL.clear() + + +def test_json_reports_absolute_execution_provenance(ahc, monkeypatch, capsys): + code, out = _run_main(ahc, monkeypatch, capsys, ["--json"]) + payload = _json_payload(out) + assert payload["execution_path"] == os.path.abspath(str(AHC)) + assert payload["cwd"] == os.getcwd() + assert code == 1 # stubbed SSH fails every leg, but provenance is still emitted + + +def test_quiet_run_still_carries_provenance_on_the_alert_path(ahc, monkeypatch, capsys): + # The cron runs --quiet; a failure report must still carry provenance. The + # header line is suppressed in quiet mode, so the ALERT line is the carrier. + code, out = _run_main(ahc, monkeypatch, capsys, ["--quiet"]) + assert code == 1 + assert "📍 executed from:" not in out + alerts = [ln for ln in out.splitlines() if ln.startswith("ALERT agent-health:")] + assert alerts, out + assert f"script={os.path.abspath(str(AHC))}" in alerts[0] + assert f"cwd={os.getcwd()}" in alerts[0] + + +def test_quiet_healthy_run_emits_no_stdout(ahc, monkeypatch, capsys): + # --quiet is documented as "only output on failure": a run with no fleet + # failures must produce no stdout at all (the production cron runs --quiet). + for name in ("check_keys", "check_gpu_ports", "check_agents", "check_ct_liveness", + "check_config_integrity", "check_wrapper_integrity", "check_vault_secrets"): + monkeypatch.setattr(ahc, name, lambda: None) + code, out = _run_main(ahc, monkeypatch, capsys, ["--quiet"]) + assert code == 0 + assert out == "" + + +def test_json_surfaces_report_only_findings_separately(ahc, monkeypatch, capsys): + # Koby's down legs are reported but must not count as fleet failures; the + # --json payload exposes them in their own array (item 1 + f8). + _, out = _run_main(ahc, monkeypatch, capsys, ["--json"]) + payload = _json_payload(out) + assert isinstance(payload["report_only"], list) + assert any(key.startswith(("gateway-down:koby", "ct-unreachable:koby")) + for key in payload["report_only"]) + assert not any("koby" in key for key in payload["failures"]) + + +# ── agent-health-check: wrapper infisical behavior (f3) ─────────────── + +def _stub_wrapper_ssh(ahc, monkeypatch, wrapper_body, test_x_result="OK", command_v="/usr/local/bin/infisical"): + def fake_ssh(host, cmd, user="root"): + if cmd.startswith("cat /root/.local/bin/hermes"): + return wrapper_body + if cmd.startswith("ls -la /root/.local/bin/hermes "): + return "-rwxr-xr-x 1 root root 0 Jan 1 00:00 /root/.local/bin/hermes" + if cmd.startswith("ls -la /root/.local/bin/hermes-real") or "venv/bin/hermes" in cmd: + return "-rwxr-xr-x 1 root root 0 Jan 1 00:00 /root/.local/bin/hermes-real" + if cmd.startswith("grep -c 'LITELLM_API_KEY'"): + return "1" + if cmd.startswith("test -x "): + path = cmd[len("test -x "):].split()[0] + if isinstance(test_x_result, dict): + return test_x_result.get(path, "MISS") + return test_x_result + if cmd.startswith("command -v infisical"): + return command_v + return None + + monkeypatch.setattr(ahc, "ssh", fake_ssh) + monkeypatch.setattr(ahc, "AGENTS", {"koonimo": dict(ahc.AGENTS["koonimo"])}) + ahc.FAIL.clear() + ahc.REPORT_ONLY.clear() + + +def test_env_based_wrapper_without_infisical_is_not_failed(ahc, monkeypatch, capsys): + _stub_wrapper_ssh(ahc, monkeypatch, + "#!/bin/bash\nsource ~/.hermes/.env\nexec hermes-real \"$@\"\n") + ahc.check_wrapper_integrity() + out = capsys.readouterr().out + assert ahc.FAIL == [] + assert "wrapper resolves creds without infisical" in out + + +def test_dangling_absolute_infisical_path_is_failed(ahc, monkeypatch, capsys): + # Wrapper hardcodes /usr/bin/infisical, which is absent, while PATH resolves + # infisical to /usr/local/bin/infisical. The literal path must be verified, + # not inferred from PATH resolution. + _stub_wrapper_ssh(ahc, monkeypatch, + "#!/bin/bash\n/usr/bin/infisical run -- hermes-real \"$@\"\n", + test_x_result="MISS", command_v="/usr/local/bin/infisical") + ahc.check_wrapper_integrity() + assert "wrapper-infisical-path:koonimo" in ahc.FAIL + + +def test_existing_absolute_infisical_path_passes(ahc, monkeypatch, capsys): + _stub_wrapper_ssh(ahc, monkeypatch, + "#!/bin/bash\n/usr/bin/infisical run -- hermes-real \"$@\"\n", + test_x_result="OK") + ahc.check_wrapper_integrity() + out = capsys.readouterr().out + assert ahc.FAIL == [] + assert "wrapper infisical path OK" in out + + +def test_comment_mentioning_removed_infisical_path_is_not_failed(ahc, monkeypatch, capsys): + # litellm-api-keys.prose.md documents `rm -f /usr/local/bin/infisical`; a + # wrapper comment about that migration must not manufacture a dangling path + # when the real invocation (/usr/bin/infisical) is present and executable. + _stub_wrapper_ssh(ahc, monkeypatch, + "#!/bin/bash\n# migrated from /usr/local/bin/infisical\n" + "exec /usr/bin/infisical run -- hermes-real \"$@\"\n", + test_x_result={"/usr/bin/infisical": "OK", + "/usr/local/bin/infisical": "MISS"}) + ahc.check_wrapper_integrity() + out = capsys.readouterr().out + assert ahc.FAIL == [] + assert "wrapper infisical path OK" in out + + +def test_comment_only_infisical_mention_does_not_reach_path_check(ahc, monkeypatch, capsys): + # A comment-only mention of a removed infisical path on a healthy .env-based + # wrapper is not an invocation: it must not fall through to the `command -v` + # PATH check and false-FAIL `wrapper-no-infisical`. + _stub_wrapper_ssh(ahc, monkeypatch, + "#!/bin/bash\n# migrated from /usr/local/bin/infisical\n" + "source ~/.hermes/.env\nexec hermes-real \"$@\"\n", + test_x_result="MISS", command_v=None) + ahc.check_wrapper_integrity() + out = capsys.readouterr().out + assert ahc.FAIL == [] + assert "wrapper resolves creds without infisical" in out + + +# ── item 4: prose-lint enforces report provenance (real consumer) ───── + +GOOD_CONTRACT = textwrap.dedent("""\ + --- + kind: function + name: good + description: fixture with provenance + --- + + ## Parameters + + - x: y + + ## Returns + + ok + + ### check-health + + ```bash + pwd -P + ``` + + **Report format**: Begin with the absolute path the probe executed from. + """) + +DECOY_CONTRACT = textwrap.dedent("""\ + --- + kind: function + name: decoy + description: fixture with provenance only outside the report format + --- + + ## Parameters + + - x: y + + ## Returns + + ok + + The absolute path of the config is /etc/foo. + + ### check-health + + ```bash + true + ``` + + **Report format**: Summarize actual results from each probe. + """) + +MISSING_CONTRACT = textwrap.dedent("""\ + --- + kind: function + name: missing + description: check-health contract with no report format + --- + + ## Parameters + + - x: y + + ## Returns + + ok + + ### check-health + + ```bash + pwd -P + ``` + """) + + +def _run_lint(tmp_path, text, name): + (tmp_path / name).write_text(text) + return subprocess.run(["bash", str(LINT)], cwd=tmp_path, + capture_output=True, text=True) + + +def test_prose_lint_accepts_report_format_with_provenance(tmp_path): + result = _run_lint(tmp_path, GOOD_CONTRACT, "good.prose.md") + assert result.returncode == 0, result.stdout + result.stderr + + +def test_prose_lint_rejects_report_format_without_provenance(tmp_path): + result = _run_lint(tmp_path, DECOY_CONTRACT, "decoy.prose.md") + assert result.returncode == 1, result.stdout + assert "lacks execution provenance" in result.stdout + + +def test_prose_lint_requires_report_format_on_check_health_contract(tmp_path): + result = _run_lint(tmp_path, MISSING_CONTRACT, "missing.prose.md") + assert result.returncode == 1, result.stdout + assert "no **Report format** paragraph" in result.stdout + + +# ── contracts: parse the executable check-health probe block ───────── + +def _check_health_block(contract): + """Extract the bash probe block under ### check-health (the probe interface).""" + text = contract.read_text() + marker = "### check-health" + assert marker in text, f"{contract.name} has no {marker}" + after = text.split(marker, 1)[1] + match = re.search(r"```bash\n(.*?)```", after, re.S) + assert match, f"{contract.name} check-health has no bash probe block" + return match.group(1) + + +def _loop_nodes(block): + nodes = [] + for line in block.splitlines(): + match = re.match(r"\s*for\s+\w+\s+in\s+(.+?);?\s*do\b", line) + if match: + nodes = match.group(1).split() + return nodes + + +def _record(url): + """Normalize a URL into a probe record: host, port, path, expected status.""" + match = re.match(r"https?://([^/\s\"')]+)(/[^\s\"')]*)?", url) + assert match, f"unparseable probe URL: {url}" + hostport = match.group(1) + if "@" in hostport: + hostport = hostport.split("@", 1)[1] + if hostport.startswith("["): + host, port = hostport[1:hostport.index("]")], None + elif ":" in hostport: + host, raw_port = hostport.rsplit(":", 1) + port = int(raw_port) if raw_port.isdigit() else None + else: + host, port = hostport, None + return {"host": host, "port": port, "path": match.group(2) or "/", + "expected": None} + + +def _probes(block): + """Parse the executable check-health bash block into a normalized probe model. + + Comments are not probes; an `# Expected: ` comment annotates the + preceding probe. URLs using the block's shell-loop variable `$node` are + expanded over the loop's node list. + """ + loop_nodes = _loop_nodes(block) + probes = [] + last = None + for raw in block.splitlines(): + stripped = raw.strip() + if stripped.startswith("#"): + expected = re.search(r"Expected:\s*(\d{3})", stripped, re.I) + if expected and last is not None: + last["expected"] = int(expected.group(1)) + continue + for url in re.findall(r"https?://[^\s\"')]+", raw): + hosts = loop_nodes if "$node" in url else [None] + for node in hosts: + record = _record(url.replace("$node", node) if node else url) + probes.append(record) + last = record + return probes + + +def test_gpu_monitor_probes_every_gpu_health_on_8080(): + probes = _probes(_check_health_block(GPU)) + targets = {(p["host"], p["port"], p["path"]) for p in probes} + assert ("192.168.68.8", 8080, "/health") in targets + assert ("192.168.68.110", 8080, "/health") in targets + + +def test_gpu_monitor_never_probes_bare_port_80_on_gpu_hosts(): + probes = _probes(_check_health_block(GPU)) + gpu_hosts = {"192.168.68.8", "192.168.68.110", "192.168.68.15"} + offenders = [p for p in probes + if p["host"] in gpu_hosts and p["port"] in (None, 80)] + assert offenders == [] + + +def test_probe_model_flags_explicit_port_80_on_gpu_host(): + # Regression: a bare-port probe may be spelled with an explicit :80. + block = ("curl -s -o /dev/null -w '%{http_code}' " + "http://192.168.68.8:80/health\n") + gpu_hosts = {"192.168.68.8", "192.168.68.110", "192.168.68.15"} + offenders = [p for p in _probes(block) + if p["host"] in gpu_hosts and p["port"] in (None, 80)] + assert offenders and offenders[0]["port"] == 80 + + +def test_gpu_monitor_treats_router_301_as_alive(): + probes = _probes(_check_health_block(GPU)) + unified = [p for p in probes + if p["host"] == "192.168.68.116" and p["path"] == "/health/unified"] + assert unified, "router /health/unified probe missing" + assert unified[0]["expected"] == 301 + + +def test_infra_monitoring_probes_every_real_pve_node(): + probes = _probes(_check_health_block(INFRA)) + pve = {(p["host"], p["port"], p["path"]) for p in probes if p["port"] == 8006} + assert {host for host, _, _ in pve} == PVE_NODE_IPS + assert {path for _, _, path in pve} == {"/api2/json/version"} + + +def test_infra_monitoring_does_not_probe_ct116_for_pve_api(): + probes = _probes(_check_health_block(INFRA)) + assert not any(p["host"] == "192.168.68.116" and p["port"] == 8006 + for p in probes) diff --git a/tests/zulip-monitor-abiba.sh b/tests/zulip-monitor-abiba.sh new file mode 100755 index 0000000..8f8ae25 --- /dev/null +++ b/tests/zulip-monitor-abiba.sh @@ -0,0 +1,211 @@ +#!/bin/bash +# tests/zulip-monitor-abiba.sh — regression test pinning the producer→consumer +# contract between the pi Zulip extension's :9200/health payload and the Abiba +# leg of scripts/zulip-monitor.sh. +# +# WHY THIS TEST EXISTS: 2026-09-09 live incident. The monitor parsed the health +# payload at the WRONG nesting level (d.get('connected') at top level, while the +# extension serves zulip.connected) so PI_CONNECTED was always False and every +# monitor run restarted a healthy bot: pm2 showed restarts=8 with the process +# created 2026-09-09T09:35:09Z, the monitor log recorded four ❌ Abiba verdicts +# (04:23, 05:35, 06:55, 09:35 UTC) and zero ✅, while the Zulip server answered +# HTTP 200 and the bot logged a clean connect plus continuing heartbeats. The +# watchdog was the fault, not the connection. This test makes that class of +# regression fail loudly instead of silently restarting healthy services. +# +# CONTRACT UNDER TEST (must hold for scripts/zulip-monitor.sh): +# * Connection state is NESTED: zulip.connected (boolean) and zulip.last_error +# live inside the `zulip` object. There is NO top-level `connected` and NO +# retry counter anywhere in the payload (verified against the extension's +# startHealthServer handler) — the old retry_count branch was dropped. +# * zulip.connected=true -> log "✅ Connected", NO pm2 restart. +# * zulip.connected=false -> alert, pm2 restart abiba-zulip. +# * fetch error / non-2xx / empty body / unparseable body / missing or +# non-boolean zulip.connected -> "⚠️ Probe failed" alert with a +# "NOT restarting" label, NO pm2 restart. A parse miss must never kill a +# healthy service. +# * zulip.connected=true with last_error -> degraded 🟡 warning, no restart. +# +# HOW: the Abiba leg of the shipped script sits between the +# `# -- abiba-leg-start` / `# -- abiba-leg-end` marker comments. This runner +# extracts that block verbatim and executes it with a stubbed curl (fixture body +# + HTTP code), recorded notify()/pm2 shims, and a temp $LOG. If the markers +# disappear (fix reverted or renamed) extraction yields nothing and the suite +# fails — the bug cannot return silently. +# +# Usage: bash tests/zulip-monitor-abiba.sh [path/to/zulip-monitor.sh] +# Exit 0 iff every check passes. +# +# shellcheck disable=SC2034,SC2329,SC1090 +# LOG/ISSUES and the notify/pm2/curl stubs below are consumed at runtime by +# the leg extracted between the marker comments and `source`d in each case; +# the static analyzer cannot see across that dynamic source, so it flags them. +set -uo pipefail + +ROOT=$(cd "$(dirname "$0")/.." && pwd) +SCRIPT=${1:-"$ROOT/scripts/zulip-monitor.sh"} +FIXTURES="$ROOT/tests/fixtures" +TMP=$(mktemp -d) +trap 'rm -rf "$TMP"' EXIT + +PASS=0 +FAIL=0 +ok() { PASS=$((PASS + 1)); printf ' \033[32m✔\033[0m %s\n' "$1"; } +bad() { FAIL=$((FAIL + 1)); printf ' \033[31m✘\033[0m %s\n' "$1"; } + +echo "== tests/zulip-monitor-abiba.sh — Abiba leg vs :9200/health producer contract ==" +echo "target script: $SCRIPT" + +# --- structural guards ------------------------------------------------------- +if ! grep -q '^# -- abiba-leg-start' "$SCRIPT"; then + echo "✘ FATAL: $SCRIPT has no '# -- abiba-leg-start' marker — the fix has been reverted or renamed." + exit 1 +fi +if ! grep -q '^# -- abiba-leg-end' "$SCRIPT"; then + echo "✘ FATAL: $SCRIPT has no '# -- abiba-leg-end' marker." + exit 1 +fi + +LEG="$TMP/leg.sh" +awk '/^# -- abiba-leg-start/{f=1; next} + /^# -- abiba-leg-end/{f=0; next} + f' "$SCRIPT" > "$LEG" +if [ ! -s "$LEG" ]; then + echo "✘ FATAL: extracted Abiba leg is empty." + exit 1 +fi +echo "== structural ==" +if bash -n "$SCRIPT"; then ok "syntax: bash -n $SCRIPT"; else bad "syntax: bash -n $SCRIPT failed"; fi +if bash -n "$LEG"; then ok "syntax: extracted leg parses (bash -n)"; else bad "syntax: extracted leg fails bash -n"; fi + +# --- per-case harness --------------------------------------------------------- +CURRENT_NAME="" +CURRENT_DIR="" + +# $1 case name, $2 http-code, $3 body (file path or literal) +run_case() { + local name="$1" http="$2" body_src="$3" body + CURRENT_NAME="$name" + CURRENT_DIR=$(mktemp -d "$TMP/case.XXXXXX") + if [ -f "$body_src" ]; then + body=$(cat "$body_src") + else + body="$body_src" + fi + ( + LOG="$CURRENT_DIR/log"; ISSUES=0 + notify() { printf 'ALERT [%s] %s\n' "$1" "$2" >> "$CURRENT_DIR/alerts"; } + pm2() { printf 'PM2 %s\n' "$*" >> "$CURRENT_DIR/pm2"; } + curl() { + local url="" + for a in "$@"; do case "$a" in http*) url="$a";; esac; done + case "$url" in + *:9200/health*) + case " $* " in + *"-w"*) printf '%s' "$http" ;; # -w '%{http_code}' code probe + *) printf '%s' "$body" ;; # body probe + esac ;; + *) + printf 'UNEXPECTED-CURL %s\n' "$*" >> "$CURRENT_DIR/unexpected-curl" + return 7 ;; + esac + return 0 + } + source "$LEG" + ) +} + +assert_log_has() { + if grep -qF -- "$1" "$CURRENT_DIR/log"; then ok "$CURRENT_NAME — log has: $1"; else bad "$CURRENT_NAME — log MISSING: $1"; fi +} +assert_log_lacks() { + if grep -qF -- "$1" "$CURRENT_DIR/log"; then bad "$CURRENT_NAME — log must NOT contain: $1"; else ok "$CURRENT_NAME — log correctly lacks: $1"; fi +} +assert_alert_has() { + if grep -qF -- "$1" "$CURRENT_DIR/alerts"; then ok "$CURRENT_NAME — alert sent: $1"; else bad "$CURRENT_NAME — alert MISSING: $1"; fi +} +assert_alert_empty() { + if [ ! -s "$CURRENT_DIR/alerts" ]; then ok "$CURRENT_NAME — no alert sent (quiet healthy path)"; else bad "$CURRENT_NAME — unexpected alert: $(cat "$CURRENT_DIR/alerts")"; fi +} +assert_pm2_restarted() { + if grep -qF "PM2 restart abiba-zulip" "$CURRENT_DIR/pm2"; then ok "$CURRENT_NAME — pm2 restart abiba-zulip was called"; else bad "$CURRENT_NAME — expected pm2 restart abiba-zulip, pm2 log: $(cat "$CURRENT_DIR/pm2" 2>/dev/null)"; fi +} +assert_no_restart() { + if [ ! -s "$CURRENT_DIR/pm2" ]; then ok "$CURRENT_NAME — NO pm2 restart (fail-safe holds)"; else bad "$CURRENT_NAME — pm2 was called but must NOT be: $(cat "$CURRENT_DIR/pm2")"; fi +} +assert_no_unexpected_curl() { + if [ ! -s "$CURRENT_DIR/unexpected-curl" ]; then ok "$CURRENT_NAME — only :9200/health was probed"; else bad "$CURRENT_NAME — unexpected curl: $(cat "$CURRENT_DIR/unexpected-curl")"; fi +} + +# --- case 1: real payload shape, zulip.connected=true -> healthy, no restart -- +echo "== case 1: connected (real producer payload: nested zulip.connected=true) ==" +run_case "connected" 200 "$FIXTURES/zulip-health-connected.json" +assert_log_has "Abiba: ✅ Connected (processed=0)" +assert_log_lacks "Disconnected" +assert_alert_empty +assert_no_restart +assert_no_unexpected_curl + +# --- case 2: zulip.connected=false -> disconnected, restart ------------------- +echo "== case 2: disconnected (nested zulip.connected=false triggers restart) ==" +run_case "disconnected" 200 "$FIXTURES/zulip-health-disconnected.json" +assert_log_has "Abiba: ❌ Disconnected — restarted" +assert_alert_has "DISCONNECTED — restarting" +assert_pm2_restarted +assert_no_unexpected_curl + +# --- cases 3-9: probe failures must alert and MUST NOT restart ---------------- +echo "== probe-failure cases: alert 'NOT restarting', zero pm2 restarts ==" + +run_case "empty body" 200 "" +assert_log_has "Abiba: ⚠️ Probe failed" +assert_log_lacks "❌ Disconnected" +assert_alert_has "NOT restarting" +assert_no_restart + +run_case "garbage body" 200 '{not valid json!!' +assert_log_has "Abiba: ⚠️ Probe failed" +assert_alert_has "NOT restarting" +assert_no_restart + +run_case "missing zulip key" 200 '{"status":"ok","platform":"pi","agent":"abiba"}' +assert_log_has "Probe failed" +assert_alert_has "NOT restarting" +assert_no_restart + +run_case "zulip without connected" 200 '{"status":"ok","zulip":{"last_error":null}}' +assert_log_has "Probe failed" +assert_alert_has "NOT restarting" +assert_no_restart + +run_case "non-boolean connected" 200 '{"status":"ok","zulip":{"connected":"true"}}' +assert_log_has "Probe failed" +assert_alert_has "NOT restarting" +assert_no_restart + +run_case "fetch failure http 000" 000 "" +assert_log_has "Probe failed" +assert_alert_has "NOT restarting" +assert_no_restart + +run_case "non-2xx http 500" 500 '{"error":"boom"}' +assert_log_has "Probe failed" +assert_alert_has "NOT restarting" +assert_no_restart + +# --- case 10: connected but last_error set -> degraded 🟡, no restart --------- +echo "== case 10: degraded (connected=true but last_error set) warns, no restart ==" +run_case "degraded" 200 '{"status":"ok","zulip":{"connected":true,"last_error":"transient queue hiccup","messages_processed":3}}' +assert_log_has "Abiba: 🟡 Error: transient queue hiccup" +assert_log_lacks "❌ Disconnected" +assert_no_restart + +# --- summary ------------------------------------------------------------------- +echo "" +if [ "$FAIL" -eq 0 ]; then + echo "✅ ALL CHECKS PASSED ($PASS/$PASS) — tests/zulip-monitor-abiba.sh" + exit 0 +else + echo "❌ $FAIL CHECK(S) FAILED ($PASS passed) — tests/zulip-monitor-abiba.sh" + exit 1 +fi diff --git a/zulip-health.prose.md b/zulip-health.prose.md index 64adba8..6ebafd4 100644 --- a/zulip-health.prose.md +++ b/zulip-health.prose.md @@ -1,9 +1,9 @@ --- kind: responsibility name: zulip-health -description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (Agent Zero Docker), Platform B (Tanko on DSH / Mumuni on Hermes), and the Zulip bridge. Verifies bot registration, DM delivery, and cross-platform connectivity. +description: Multi-platform health monitor for the Zulip messaging mesh spanning Platform A (pi/Abiba Zulip bridge), Platform B (Tanko on DSH), and Platform C (Agent Zero Docker). Verifies bot registration, DM delivery, and cross-platform connectivity. Mumuni is no longer monitored from this host — she runs on her own container (kagentz CT 105 on minipve, .14) and is monitored on her side. title: Zulip Mesh Health Monitor — Multi-Platform -version: 3.0.0 +version: 3.1.0 runtime_contract: 2 agent: abiba report_only_agents: @@ -12,13 +12,23 @@ report_only_agents: # Zulip Mesh Health Monitor -Monitors ALL Zulip-connected agents across platforms (pi, Hermes, DSH, Agent Zero). -Runs every 15 minutes in the background. Also triggers on session start. +Monitors the Zulip-connected agents under this host's operational control (pi, +DSH, Agent Zero). Runs every 15 minutes in the background. Also triggers on +session start. + +> **Mumuni is NOT monitored from this host (captain ruling 2026-09-10).** She +> moved off this host onto her own container — kagentz CT 105 on minipve +> (192.168.68.14), running a dedicated `hermes` user — and is monitored on her +> side. No step in this contract, and no leg of `scripts/zulip-monitor.sh`, may +> ssh to her old deployment, read her `~/.hermes/gateway_state.json`, or alert on +> her state. The former Platform-B-for-Mumuni steps (gateway process, heartbeat, +> response delivery) are retired: they always read "unknown" against the +> decommissioned deployment and produced a false 🔴 alert on every run. ## Requires - **Zulip API key** for `abiba-bot@chat.sysloggh.net` in `$ZULIP_API_KEY` -- **SSH access** to Tanko (192.168.68.122), Mumuni (192.168.68.14, kagentz CT105 on minipve), and Agent Zero Docker host (192.168.68.14) +- **SSH access** to amdpve (192.168.68.15) for Tanko — CT 112 reached via `pct exec` (direct SSH to .122 is not a dependency of this contract: per-worker key availability varies); and the Agent Zero Docker host (192.168.68.14) - **PM2** on localhost for pi process management - **Network access** to `chat.sysloggh.net`, `localhost:9200` - **Write access** to `/root/zulip-health-monitor.log` and `/tmp/zulip-monitor-debounce` @@ -48,10 +58,8 @@ Runs every 15 minutes in the background. Also triggers on session start. }, "tanko": { "platform": "dsh", - "zulip_state": "connected", - "heartbeat_age_seconds": 45, - "gateway_pid": 1234, - "edit_fail_rate_pct": 0, + "service_state": "active", + "http_status": 200, "severity": "healthy" } } @@ -83,13 +91,13 @@ Log as "unreachable" — don't treat as critical unless it persists for 3+ conse ## Streaming Support (2026-07-05) Zulip agents now support progressive message editing during agent generation. -When a Zulip agent (Tanko on DSH, Mumuni on Hermes) processes a message, the response is +When a Zulip agent under this monitor's scope (Tanko on DSH) processes a message, the response is streamed in real-time via Zulip's `PATCH /api/v1/messages/{id}` API: - Adapter implements `edit_message()` using `_api_patch()` helper - Gateway stream consumer progressively edits the Zulip message - User sees real-time agent thinking instead of waiting for full response -- Verified: Tanko (CT 112) and Mumuni (kagentz CT 105) both have streaming active +- Verified: Tanko (CT 112) has streaming active; Mumuni's (kagentz CT 105) is verified on her own host, not from here ### Verification ```bash @@ -185,63 +193,84 @@ grep -a "Finalized\|Failed to finalize" /root/.pm2/logs/abiba-zulip-out.log | ta | Crash loop >10/h | Alert user | -**B1: Gateway State** +### Step 3: Platform B — Tanko (DSH on amdpve CT 112) + +Mumuni is out of scope for this host (see the note above): she runs on her own +container and is monitored on her side. + +Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so +there is no `~/.hermes/gateway_state.json` on CT 112. Tanko's Zulip gateway runs +as the `dsh-web` systemd unit inside **CT 112**, which resides on the **amdpve** +PVE host (**192.168.68.15**). Direct SSH to 192.168.68.122 is not a dependency +of this contract — per-worker key availability varies — so CT 112 probes run +from the amdpve vantage via `pct exec`: ```bash +ssh root@192.168.68.15 "pct exec 112 -- " ``` -Tanko runs on DSH (DeepSeek Harness) — it no longer runs a Hermes gateway, so there is no `~/.hermes/gateway_state.json` on CT 112 (.122). Verify Tanko's Zulip connectivity via the DSH harness bot status instead. +> **By design (verified 2026-09-08):** the `dsh-web` gateway binds +> `127.0.0.1:3080` **loopback-only**. A remote probe against +> `192.168.68.122:3080` gets connection-refused — that is EXPECTED, NOT a fault, +> and must never be raised as Tanko down. Only loopback probes from inside +> CT 112 (or the public-URL fallback below) are valid health signals. -Check `platforms.zulip.state`: `connected` ✅ | `disconnected` ❌ | `error` ❌ | missing → not installed. - -**B2: Agent Process** +**B1: Gateway Service State (Tanko)** ```bash -ssh root@ "ps aux | grep 'gateway run' | grep -v grep" +ssh root@192.168.68.15 "pct exec 112 -- systemctl is-active dsh-web" ``` -Gateway PID should exist with uptime > 60s. **Dual-gateway detection**: if more than one `gateway run` process is found, the gateway has a collision (typically one `--force` and one `--replace` process). Kill the newer/duplicate process, then restart the remaining gateway per-agent (parameterized 2026-08-09, captain ruling): +Expected: `active`. Anything else → gateway service down → apply the Tanko heal +(restart via DSH service, Platform B Actions table below). -| Agent | Restart command | Notes | -|-------|-----------------|-------| - -Check gateway log for "Gateway running with 2 platform(s)" (not 1) to confirm Zulip reloaded. - -**B3: Heartbeat Verification** (Hermes agent Mumuni only — Tanko has no Hermes gateway) +**B2: Gateway HTTP Liveness (Tanko — loopback-only :3080)** ```bash -ssh root@192.168.68.24 "grep Heartbeat ~/.hermes/logs/agent.log | tail -3" +ssh root@192.168.68.15 "pct exec 112 -- curl -s --connect-timeout 5 --max-time 10 -o /dev/null -w '%{http_code}' http://127.0.0.1:3080/" ``` -Expected: recent heartbeat (within 5 min), `polls=N` incrementing. -Silence > 300s → warning. Silence > 600s → critical. +Alive = **ANY** HTTP status response from the endpoint — the expected set is +`200`/`301`/`302`/`307`/`308`/`401`/`403` (the gateway UI is token-gated and +legitimately answers with redirects/auth-challenges, so never require a bare +`200`), and any other status, including `404`/`5xx`, also counts alive: a +process answering `503` is running and self-heal must NOT restart-loop it. +Down = connection refused (`000`) or timeout only. Statuses outside the +expected set are logged/reported as a warning — reported, never healed on. -**B4: Response Delivery** (Hermes agent Mumuni only) +**B3: Public-URL Fallback Probe (Tanko — for nodes without pct/ssh access to amdpve)** ```bash -ssh root@192.168.68.24 "grep -E 'Finalized|Failed to finalize|Replied to' ~/.hermes/logs/agent.log | tail -10" +curl -s --connect-timeout 10 --max-time 15 -o /dev/null -w '%{http_code}' https://tankodhs.sysloggh.net/ ``` -> 50% fail rate → critical. +Fallback only — used when the monitoring node has no pct/SSH path to amdpve. +Alive = **ANY** HTTP status response from the endpoint — healthy signals are +`302` (authentik proxy-auth redirect) and `401` (auth-gated), and any other +status, including `404`/`5xx`, also counts alive: the endpoint is up and +answering and must NOT be restart-looped. Down = connection refused (`000`) or +timeout only. Never expect a bare `200` — the public URL terminates in the +token-gated authentik chain. Statuses outside the healthy set are +logged/reported as a warning — reported, never healed on. **Platform B Actions** | Condition | Action | |-----------|--------| -| `zulip.state != "connected"` | `ssh root@ "pkill -f 'gateway run'; sleep 2; hermes gateway restart"` (Mumuni) / restart Tanko via DSH service | -| No heartbeat in 10min | Same as above | -| `Failed to finalize` > 50% | Check PATCH API, Zulip server | -| Response empty/short | Check A2A endpoint / LiteLLM model | +| `dsh-web` service not `active` | Restart Tanko via DSH service | +| HTTP `:3080` connection refused/timeout (`000`) | Same as above | +| HTTP status outside the expected set | Log/report as a warning — reported, never healed on | ### Step 4: Platform C — Agent Zero (kagentz, CT 105 via Docker host .14) **C1: A2A Server Health** ```bash -ssh root@192.168.68.14 "docker exec agent-zero curl -s --connect-timeout 5 http://127.0.0.1:8001/.well-known/agent.json" +# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated) +ssh root@192.168.68.14 "curl -s --connect-timeout 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:50080/a2a/" ``` -Expected: `{"name":"kagentz",...}`. Connection refused → A2A server down. +Expected: `401` (auth-gated, A2A server is up and responding) or `200` (if no auth required). Connection refused (000) → A2A server down. **C2: Adapter Process** @@ -262,12 +291,14 @@ Check: `processed=N` incrementing, `silence < 600s`, `reconnects` ≈ 0. **C4: A2A Response Verification** ```bash -ssh root@192.168.68.14 "docker exec agent-zero curl -s -X POST http://127.0.0.1:8001/a2a \ +# A2A server is on :50080 (not :8001) and is auth-gated (401 expected for unauthenticated) +ssh root@192.168.68.14 "curl -s -X POST http://127.0.0.1:50080/a2a \ -H 'Content-Type: application/json' \ + -H 'Authorization: Bearer $LITELLM_KEY' \ -d '{\"jsonrpc\":\"2.0\",\"method\":\"tasks/send\",\"params\":{\"message\":{\"role\":\"user\",\"parts\":[{\"text\":\"ping\"}]}},\"id\":1}'" ``` -Expected: task ID with "working" status. Poll for completion with `tasks/get`. +Expected: task ID with "working" status. Poll for completion with `tasks/get`. If 401, check LITELLM_KEY is set. **Platform C Actions** @@ -284,7 +315,7 @@ Expected: task ID with "working" status. Poll for completion with `tasks/get`. Check each agent's log for excessive bot-to-bot chatter: - Abiba: `Skipped.*bot msgs` count -- Tanko/Mumuni: Repeated DM exchanges between bots +- Tanko: Repeated DM exchanges between bots - kagentz: Adapter log for bot DMs being processed If any bot processes >50 bot-originated messages in 15min → warning.