diff --git a/docs/AUTHORING-GUIDE.md b/docs/AUTHORING-GUIDE.md index bfe4c8b..f73129e 100644 --- a/docs/AUTHORING-GUIDE.md +++ b/docs/AUTHORING-GUIDE.md @@ -211,6 +211,40 @@ Errors tell the operator what went wrong and what to do about it. Be specific: Check router /health/unified at http://192.168.68.116/health/unified instead." ``` +### Report provenance + +Every report a contract produces must lead with the **absolute path the probe +executed from** — `pwd -P`, or the running script's absolute path. A live-state +report without provenance is unactionable: a report from a stale copy (a worktree +clone, a retired cron entry, a diverged consumer) looks identical to a live +fault, and the team burns rounds repairing healthy infrastructure. This is not +optional. The 2026-09-09 probe-drift rounds cost three false `DEGRADED` reports +because a stale consumer probed the wrong port and nothing in the report said +where it ran. + +Pair it with the **scoped any-HTTP-response liveness rule**: for unauthenticated +or auth-gated endpoints — where any HTTP answer proves a listener is up (the +PVE API's `401`, LiteLLM health's `301` redirect) — a probe is ALIVE on ANY HTTP +status, including `301` redirects and `401`/`403` auth challenges. **DOWN = +connection refused (`000`) or timeout only.** + +Probes whose success condition is specifically a bare `200` are NOT covered by +the any-HTTP rule. On those — authenticated probes such as the Zulip message +POST and the router `/health` — an unexpected status (`401`/`403` from a bad or +missing credential, `5xx`, or anything other than the expected `200`) is an +**ALERT**, not "alive". + +```markdown +**Report format**: Begin every report with the absolute execution path +(`pwd -P` / script path). On auth-gated endpoints, alive = ANY HTTP status and +DOWN = `000`/timeout only; on probes whose expected result is `200`, any other +status is an alert. +``` + +The lint pipeline enforces the provenance clause: any contract with a +`**Report format**` line must state an absolute path (`pwd -P`, `absolute path`, +or `executed from`). + ### Comments Comments in contracts explain WHY, not WHAT. The execution steps say what to diff --git a/docs/probe-drift-round2-evidence.md b/docs/probe-drift-round2-evidence.md new file mode 100644 index 0000000..c6a5b2d --- /dev/null +++ b/docs/probe-drift-round2-evidence.md @@ -0,0 +1,278 @@ +# Probe-drift round 2 — per-leg before/after evidence + +**Date:** 2026-09-10 +**Worktree (absolute execution path):** `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts` +**Branch:** `fm/probe-drift-round2-20260909` + +Every command below was run from the absolute path above; output is pasted +verbatim. This is the evidence trail for the four scoped corrections; it is not +a contract (never `prose run` it). + +--- + +## Leg 1 — agent-health-check (item 1) + +**Before** — from `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts`, +`python3 scripts/agent-health-check.py --no-deploy` (v2, base of this branch): + +``` +🏥 Agent Health Check v2 — 2026-09-10 01:16 UTC + +🔑 LiteLLM Keys: + ✅ tanko: key valid → syslog-auto + ✅ abiba: key valid → syslog-auto + ✅ koby: key valid → syslog-auto + ✅ koonimo: key valid → syslog-auto + +🎮 GPU Port Health: + ✅ gpu-rtx3090 (.8): healthy (pid=472206) + ✅ gpu-rtx5070 (.110): healthy (pid=207601) + ✅ gpu-strixhalo (.15): healthy (pid=4098872) + +🤖 Agent Gateways: + ✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK) + ⚠️ abiba: gw=no-state-file zulip=? streaming=no errors_10m=0 pid=? + ✅ koby: gw=running zulip=connected streaming=no errors_10m=0 pid=360900 + ✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125 + +🖥️ CT Liveness: + ✅ tanko (CT 112 on amdpve): running + ✅ abiba (CT 100 on minipve): running + ❌ koby (CT 111 on amdpve): PVE UNREACHABLE + ✅ koonimo (CT 113 on amdpve): running + +📝 Config Integrity: + ⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27 + ✅ abiba: config.yaml valid YAML + ✅ koby: config.yaml valid YAML + ✅ koonimo: config.yaml valid YAML + +🔌 Wrapper/CLI Integrity: + ⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27 + ⚠️ abiba: wrapper infisical path may be wrong (infisical at /usr/bin/infisical) + ❌ abiba: hermes-real NOT FOUND (wrapper broken) + ⚠️ abiba: .env may be missing LITELLM_API_KEY entry + ⚠️ koby: wrapper infisical path may be wrong (infisical at /usr/bin/infisical) + ❌ koby: hermes-real NOT FOUND (wrapper broken) + ✅ koby: wrapper + .env key present + ⚠️ koonimo: wrapper infisical path may be wrong (infisical at /usr/bin/infisical) + ✅ koonimo: wrapper + .env key present + +🔐 Vault Secrets: + ✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw + ✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg + ✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ + +❌ 6 FAILURE(S): ct-unreachable:koby:192.168.68.15 | wrapper-infisical-path:abiba | wrapper-no-hermes-real:abiba | wrapper-infisical-path:koby | wrapper-no-hermes-real:koby | wrapper-infisical-path:koonimo +``` + +Root causes (all stale expectations; no live fault): + +| Failure | Why it was stale | +|---|---| +| `ct-unreachable:koby:192.168.68.15` | CT 111 (tdunna/koby) runs on **storepve (.6)**, not amdpve (.15). | +| `wrapper-*:abiba` | Abiba is pi-only since the harness purge. `/root/.local/bin/hermes` is a dangling symlink; no `hermes-real`, no `~/.hermes/.env`. | +| `wrapper-*:koby` | Koby is **report-only** (captain ruling 2026-08-17, Rule 17): detect and report, never repair — its legs must not count as fleet failures. Koby's wrapper is also the genuine no-infisical case: it sources `~/.hermes/.env` rather than `/usr/bin/infisical`, which the check now accepts. | +| `wrapper-infisical-path:koonimo` | Koonimo's wrapper **does** reference `/usr/bin/infisical` — but past the old check's `head -20` window, so the check looked for the path in the wrong slice and false-failed. The fix that mattered was reading the full wrapper body (and then verifying any absolute infisical path it finds actually exists). | + +**After** — same absolute path, `python3 scripts/agent-health-check.py --no-deploy` (v4): + +``` +$ pwd -P +/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts +$ python3 scripts/agent-health-check.py --no-deploy +🏥 Agent Health Check v4 — 2026-09-10 01:22 UTC +📍 executed from: script=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts/scripts/agent-health-check.py cwd=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts + +🔑 LiteLLM Keys: + ✅ tanko: key valid → syslog-auto + ✅ abiba: key valid → syslog-auto + ✅ koby: key valid → syslog-auto + ✅ koonimo: key valid → syslog-auto + +🎮 GPU Port Health: + ✅ gpu-rtx3090 (.8): healthy (pid=472206) + ✅ gpu-rtx5070 (.110): healthy (pid=207601) + ✅ gpu-strixhalo (.15): healthy (pid=4098872) + +🤖 Agent Gateways: + ✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK) + ✅ abiba: pi-only runtime — no Hermes gateway since the harness purge (CT 100, SSH OK) + 🔍 koby: REPORT-ONLY mode (diagnostic only, no repairs on .129) + ✅ koby: gateway running (pid=360900, report-only mode) + ✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125 + +🖥️ CT Liveness: + ✅ tanko (CT 112 on amdpve): running + ✅ abiba (CT 100 on minipve): running + ✅ koby (CT 111 on storepve): running + ✅ koonimo (CT 113 on amdpve): running + +📝 Config Integrity: + ⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27 + ⏭️ abiba: pi-only runtime — no Hermes config.yaml since the harness purge + ✅ koby: config.yaml valid YAML + ✅ koonimo: config.yaml valid YAML + +🔌 Wrapper/CLI Integrity: + ⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27 + ⏭️ abiba: pi-only runtime — no hermes CLI wrapper since the harness purge + ℹ️ koby: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK + ❌ koby: hermes-real NOT FOUND (wrapper broken) + 🔍 report-only (koby): wrapper-no-hermes-real:koby — reported, not counted/repaired + ✅ koby: wrapper + .env key present + ✅ koonimo: wrapper infisical path OK + ✅ koonimo: wrapper + .env key present + +🔐 Vault Secrets: + ✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw + ✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg + ✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ + +✅ All checks passed +exit=0 +``` + +**Live vantage proof** (same worktree): + +``` +$ ssh root@192.168.68.15 "pct status 111" +Configuration file 'nodes/amdpve/lxc/111.conf' does not exist +$ ssh root@192.168.68.6 "pct status 111; pct list | grep '^ *111'" +status: running +111 running tdunna +$ ssh root@192.168.68.129 "hostname" +tdunna +``` + +--- + +## Leg 2 — infrastructure-monitoring PVE API (item 2) + +**Before** — the contract's probe, aimed at the monitoring host CT 116: + +``` +$ curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json +000 +``` + +CT 116 runs no `pveproxy`, so it never answers on `:8006`. The probe target was +wrong, which is what read as PVE-API `000`. + +**After** — probing the five real cluster nodes (`:8006/api2/json/version`), +alive under the any-HTTP-response rule (`401` = up, unauthenticated): + +``` +$ for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do + printf '%s:8006 -> %s\n' "$node" "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")" + done +192.168.68.9:8006 -> 401 +192.168.68.5:8006 -> 401 +192.168.68.15:8006 -> 401 +192.168.68.6:8006 -> 401 +192.168.68.12:8006 -> 401 +``` + +`401` on every node = alive by design. `DOWN` is `000`/timeout only. (The +contract's LiteLLM probe was the same class: `/litellm/health` answers `301` → +`/litellm/health/liveliness`, so it is now specified as any-HTTP too.) + +--- + +## Leg 3 — gpu-monitor GPU probes (item 3) + +**Before** — the false alarm came from probing bare port 80 on GPU hosts: + +``` +$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8/health +000 +$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110/health +000 +``` + +Nothing listens on GPU port 80, so the monitor reported +`DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` three times on 2026-09-09. + +**After** — the real endpoints answer: + +``` +$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health +200 +$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health +200 +$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.15:8080/health +200 +$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified +301 # Location: http://192.168.68.116/gpu/gpu-data — the same payload +$ curl -s -o /dev/null -w '%{http_code}' -L http://192.168.68.116/health/unified +200 +``` + +`301` is healthy under the any-HTTP-response rule. The contract now requires GPU +health on `:8080` (or router `/health/unified`) and forbids bare port 80 on a +GPU host. + +--- + +## Leg 4 — report provenance (item 4) + +Every contract report must now lead with the absolute path it executed from. +`docs/AUTHORING-GUIDE.md` documents the rule and `scripts/prose-lint.sh` +enforces it: + +``` +$ pwd -P +/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts +$ bash scripts/prose-lint.sh + ✅ Report provenance present in all report-format contracts +... +✅ LINT PASSED (12 warning(s)) +``` + +The health script prints `📍 executed from: script=… cwd=…` and includes +`execution_path`/`cwd` in `--json` output. + +--- + +## Full suite + +``` +$ pwd -P +/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts +$ python3 -m pytest -q +24 passed +$ shellcheck scripts/prose-lint.sh +(clean) +``` + +--- + +## Follow-up findings (observed, intentionally NOT changed here) + +These are adjacent stale expectations discovered while verifying the four +scoped legs. Each touches a CRITICAL/HIGH-sensitivity artifact or an unrelated +script, so it is recorded for the captain/verify mate rather than silently +repaired. + +1. **`infrastructure-control.prose.md` (CRITICAL) CT 111 node assignment.** + Lines ~109 and ~615 place `tdunna` (CT 111, koby) on **amdpve**. Live + verification on 2026-09-10 shows `pct status 111` = `running` on + **storepve (.6)** and `Configuration file 'nodes/amdpve/lxc/111.conf' does + not exist` on .15. `agent-health-check.py` now carries the live-verified + `storepve` mapping (the script is not the topology source of truth); the + CRITICAL contract itself needs an authorized correction. +2. **Strix Halo `:8080` firewall claim is stale.** `prose-ai-review.sh` + ground-truth rule #4 and `gpu-monitor.prose.md` say `:8080` is firewalled to + `.116` only and `.24` cannot probe it. Live on .15: + `-A INPUT -s 192.168.68.24/32 -p tcp --dport 8080 -j ACCEPT`, and a probe + from .24 returns `200`. The contract keeps routing Strix via the router + (safe), but the claim no longer matches iptables. +3. **`contract-registry.yaml` references `agent-health-check.prose.md`**, which + does not exist in the repo. The registry entry (with `koby_action: skip_heal`) + is aspirational/stale. +4. **Pre-existing script defects, untouched:** `scripts/pm2-self-heal.sh` has a + bash syntax error at lines 19–20 (`bash -n` fails), and `shellcheck` fails on + five untouched scripts (`netbird-add-domain.sh`, `pct-run.sh`, + `pm2-self-heal.sh`, `prose-ai-review.sh`, `swap-gpu-dense-model.sh`). + `scripts/prose-lint.sh` — the one shell file touched here — is now + shellcheck-clean. diff --git a/gpu-monitor.prose.md b/gpu-monitor.prose.md index 82bd0ee..881bac4 100644 --- a/gpu-monitor.prose.md +++ b/gpu-monitor.prose.md @@ -31,20 +31,31 @@ agent: abiba └──────┘ │qwen27B│ │LiteLLM │ └──────┘ │dashboard│ └────────┘ +``` Note: JSON sidecar exporters at :8090 were never deployed on any GPU host. Router falls back to GPU /health direct probe. Monitor should use router /health/unified as source of truth for GPU status. Strix Halo :8080 is firewalled to .116 only — monitor on .24 cannot poll .15:8080 directly; must go through router on .116. -``` + +**PORT RULE (verified 2026-09-10):** GPU per-host health lives on **:8080** +(`http://:8080/health`); Prometheus GPU exporters live on **:9400**. +There is NO listener on bare port 80 for any GPU host — `http://192.168.68.8/health` +and `http://192.168.68.110/health` answer `000`. Never use a bare-port-80 probe +as a GPU liveness signal: on 2026-09-09 that produced three false +`DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` rounds while +`http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health` +answered `200`. Port 80 is valid only on the router (.116), never on a GPU host. ### Subsystems Polled | Subsystem | Endpoint | Frequency | Metrics | |-----------|----------|-----------|---------| -| GPU Status (all, via router) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (router probes each GPU /health directly) | -| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status | +| GPU Status (all, via router) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (router probes each GPU /health directly) — `301` → `/gpu/gpu-data` is **alive** | +| GPU .8 (RTX 3090) health | `http://192.168.68.8:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** | +| GPU .110 (RTX 5070) health | `http://192.168.68.110:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** | +| Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (`301` → `/gpu/gpu-data` = alive) | | Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness | | LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count | | Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) | @@ -61,6 +72,22 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and ## Alert Thresholds +### Liveness rule (scoped) + +The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness +endpoints, where any HTTP answer proves a listener is up. Applied here: the +router's `/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data` +(the same payload) and LiteLLM's `/litellm/health` answers `301` → +`/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on +**ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included — +and **DOWN = connection refused (`000`) or timeout only**. Same scoped rule as +zulip-health (Tanko) and infrastructure-monitoring. + +Probes whose success condition is specifically a bare `200` are NOT covered by +the any-HTTP rule. On those — the GPU `:8080/health` endpoints, the router +`/health`, and the dashboard — an unexpected status (`401`/`403`, `5xx`, or +anything other than the expected `200`) is an **ALERT**, not "alive". + | Metric | Warning | Critical | |--------|---------|----------| | GPU Temp | >80°C | >90°C | @@ -101,24 +128,43 @@ This replaces the previous DM-only delivery. All agents on the mesh can see and **RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** ```bash +# Provenance — run first; paste the absolute path into the report +pwd -P + # GPU Monitor health curl http://localhost:9100/health | jq # Expected: 200 with {"status": "healthy", "cache_age_seconds": } -# Router health (via nginx on port 80) +# GPU host health — DIRECT on :8080. NEVER probe bare port 80 on a GPU host: +# http://192.168.68.8/health has no listener and returns 000 → false DEGRADED. +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health +# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health +# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) + +# Router unified health (source of truth; 301 → /gpu/gpu-data is HEALTHY) +curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified +# Expected: 301 (or 200 after following the redirect) — any HTTP status = alive + +# Router basic health (via nginx on port 80 — router .116 only, never a GPU host) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health -# Expected: 200 (Router is up and responding) +# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) # LiteLLM health (via nginx on port 80) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health -# Expected: 200 (LiteLLM is up and responding) +# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive # Dashboard curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/ -# Expected: 200 (Dashboard is up and responding) +# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) ``` -**Report format**: Summarize actual results from each probe. If any probe returns non-200, flag as alert. +**Report format**: Begin every report with the **absolute path the probe executed +from** (`pwd -P`, or the monitor script's absolute path) so a stale-consumer +report is distinguishable from a real fault at read time. Summarize actual +results from each probe. Apply the scoped liveness rule above: on auth-gated +endpoints only connection-refused (`000`) or timeout is DOWN; on bare-200 probes +any other status is an alert. Never probe a GPU host on bare port 80. ### view-dashboard Open `http://localhost:9100/` in browser — Live HTML dashboard @@ -131,9 +177,11 @@ python3 /root/scripts/gpu-monitor-server.py & Or via PM2: `pm2 restart gpu-monitor` ### check-router -The router health is accessed through nginx on port 80 (NOT port 9000 directly). -`curl http://192.168.68.116/health/unified` — Router unified health via nginx proxy +The router health is accessed through nginx on port 80 on the **router** +(.116) — NOT port 9000 directly, and NOT bare port 80 on a GPU host. +`curl http://192.168.68.116/health/unified` — Router unified health via nginx proxy; answers `301` → `/gpu/gpu-data` (same payload) = alive `curl http://192.168.68.116:9000/health/unified` — ❌ WILL FAIL (port bound to 127.0.0.1 only) +`curl http://192.168.68.8/health` — ❌ NEVER USE (GPU host, no port-80 listener → false `000`/DEGRADED) ## Configuration Files @@ -145,13 +193,18 @@ The router health is accessed through nginx on port 80 (NOT port 9000 directly). ## Execution -1. **Poll router** (every 15s): GET .116/health/unified — single source of truth for all GPU status (router probes each GPU /health directly via sidecar fallback) -2. **Poll router** (every 15s): GET .116/health via nginx:80 -3. **Poll LiteLLM** (every 15s): GET .116/litellm/health via nginx:80 -4. **Poll Strix** (every 15s): via router /health/unified (cannot poll .15:8080 directly — firewalled to .116 only) -5. **Poll dashboard** (every 15s): GET .116/dashboard/ -6. **Check alerts**: Compare metrics against thresholds -7. **Compute summary**: Fleet-wide health aggregation -8. **Render dashboard**: Generate HTML at /root/dashboard/gpu-fleet.html -9. **Serve API**: HTTP server on port 9100 -10. **Repeat** every 15 seconds +**Port discipline:** probe GPU hosts on `:8080` (or the router's +`/health/unified`); probe port 80 only on the router (.116). Never bare port 80 +on a GPU host. + +1. **Poll router** (every 15s): GET .116/health/unified — single source of truth for all GPU status (router probes each GPU /health directly via sidecar fallback). `301` → `/gpu/gpu-data` counts as alive. +2. **Fallback direct GPU probe** (only if router /health/unified is DOWN): GET `http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health` — **:8080 only, never bare port 80**. +3. **Poll router** (every 15s): GET .116/health via nginx:80 +4. **Poll LiteLLM** (every 15s): GET .116/litellm/health via nginx:80 +5. **Poll Strix** (every 15s): via router /health/unified (cannot poll .15:8080 directly — firewalled to .116 only) +6. **Poll dashboard** (every 15s): GET .116/dashboard/ +7. **Check alerts**: Compare metrics against thresholds +8. **Compute summary**: Fleet-wide health aggregation +9. **Render dashboard**: Generate HTML at /root/dashboard/gpu-fleet.html +10. **Serve API**: HTTP server on port 9100 +11. **Repeat** every 15 seconds diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index 205aa03..2969ac6 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -102,11 +102,32 @@ GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo) - Stack persists across reboots (systemd for exporters, Docker restart policy) ## Execution + +### Liveness rule (scoped) + +The any-HTTP-response rule applies ONLY to unauthenticated/auth-gated endpoints, +where any HTTP answer proves a listener is up: the PVE API +(`https://:8006/api2/json/version`) and LiteLLM health +(`/litellm/health`, `301` → `/litellm/health/liveliness`). For those endpoints a +probe is **ALIVE** on **ANY** HTTP status — `401`/`403` auth challenges and `3xx` +redirects included — and **DOWN = connection refused (`000`) or timeout only**. +The PVE API legitimately answers `401` to an unauthenticated probe — that is the +healthy signal, not a failure. Same scoped rule as zulip-health (Tanko) and +gpu-monitor. + +Probes whose success condition is specifically a bare `200` are NOT covered by +the any-HTTP rule. On those — the authenticated Zulip POST and the router +`/health` — an unexpected status (`401`/`403` from a bad or missing credential, +`5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive". + ### check-health **RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** ```bash +# Provenance — run first; paste the absolute path into the report +pwd -P + # Zulip API health (POST ping) source /etc/litellm-monitor.env ZULIP_USER="abiba-bot@chat.sysloggh.net" @@ -128,11 +149,20 @@ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health # LiteLLM health (via nginx on port 80) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health -# Expected: 200 (LiteLLM is up and responding) +# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive -# PVE API (401 expected for unauthenticated probe — API is up over https) -curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json -# Expected: 401 (unauthorized — API is up; 000 = unreachable, 500 = API down) +# PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring +# host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 — +# that was the stale-vantage bug this replaces (CT 116 is the monitoring host, +# not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE +# by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response); +# DOWN = connection refused (000) or timeout only. +for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do + printf '%s:8006 -> %s\n' "$node" \ + "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")" +done +# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12) +# A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault. # Prometheus targets curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets' @@ -147,7 +177,16 @@ curl -s http://192.168.68.116:4001/metrics | head -20 # Expected: Prometheus-formatted metrics output ``` -**Report format**: Summarize actual results from each probe. If any probe returns non-200 or empty output, flag as alert. +**Report format**: Begin every report with the **absolute path the probe executed +from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is +distinguishable from a real fault at read time. Summarize actual results from +each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE +API and LiteLLM endpoints above: only connection-refused (`000`) or timeout is +DOWN; empty output is a warning. For probes whose expected result is a bare `200` +(the authenticated Zulip POST, router `/health`), flag an alert on any unexpected +status (`401`/`403`/`5xx`) — do not summarize it as alive. A bare-`200` +expectation on the auth-gated PVE API (`401`) or LiteLLM health (`301` redirect) +is a stale expectation, not a fault. ### Phase 1: GPU Exporters diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 66972b1..624d3c7 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -113,7 +113,7 @@ Preferred implementation: uncap shared pool, add capped alias for crew-only. - **Health-check script** (`/opt/inference-harness/scripts/litellm-health-check.sh` on CT 116): `gpu-fleet` check fails only on **critical** alerts (warnings are informational). Tests `strix-moe` (not `ornith-1.0-35b`). - **GPU monitor** (`/root/scripts/gpu-monitor-server.py` on pi .24): runs as **systemd unit `gpu-monitor.service`** (was bare `&` process). `gpu_count` includes Strix Halo (was 2, now 3). VRAM alert thresholds: warning 93%, critical 97% (raised from 90/95 — 128K context steady-state is ~70% on RTX 3090, not a fault). -- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v3 (2026-09-08) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Covers: LiteLLM keys, GPU ports, agent gateways (all 5 agents now SSHa ble), CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames. +- **Agent key monitor** (`/root/scripts/agent-health-check.py` on pi .24, cron `*/10`): v4 (2026-09-10) — vault-backed agents (tanko/koby/koonimo) read their **agent-specific** `{NAME}_LITELLM_API_KEY` from Infisical vault (not the shared master key); abiba (pi agent) reads `LITELLM_API_KEY` from its local `/root/.pi/agent/env.sh` (#735 — moved out of shared `/root/.bashrc`), not from the vault. Abiba is pi-only since the harness purge, so its Hermes config/wrapper/gateway legs are skipped rather than reported as faults; koby is **report-only** (captain's 2026-08-17 ruling) — its findings go to the `--json` `report_only` array and are never counted as fleet failures or repaired, and its CT 111 liveness is probed on storepve (.6). Covers: LiteLLM keys, GPU ports, agent gateways, CT liveness (pct status on PVE nodes), config.yaml YAML integrity, wrapper/CLI integrity, vault secret non-emptiness checks. Every run/report carries the absolute execution path (`script=` + `cwd=`). Fleet roster: tanko (.122), mumuni (.24, inside abiba CT100), koby (.129), koonimo (.114), abiba (.24). Legacy `tdunna`/`baggy` replaced with canonical agent hostnames. - **Stale keys cleaned**: `daily-infra-report.py` SYNTHETIC_API_KEY was stale (`sk-U_ydi3B` → 401); now reads `LITELLM_MASTER_KEY` from env. Deprecated scripts (`router-original.py`, `router-phase0-backup.py`, `apply-fixes.py`) still reference `sk-syslog-local-master-key` but do not actively poll LiteLLM. ## Maintains diff --git a/proxmox-monitor.prose.md b/proxmox-monitor.prose.md index 80ecd31..39e81d6 100644 --- a/proxmox-monitor.prose.md +++ b/proxmox-monitor.prose.md @@ -122,7 +122,10 @@ ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1 # Expected: 200 (pve-exporter is up and responding) ``` -**Report format**: Summarize actual results from each probe. If any probe returns non-200, flag as alert. +**Report format**: Begin every report with the **absolute path the probe executed +from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is +distinguishable from a real fault at read time. Summarize actual results from +each probe. If any probe returns non-200, flag as alert. **Note**: Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only) so they must be probed from .116 via SSH. Prometheus and Grafana are bound to 0.0.0.0 so they can be probed from the LAN. diff --git a/scripts/agent-health-check.py b/scripts/agent-health-check.py index ae94aa1..37966ed 100755 --- a/scripts/agent-health-check.py +++ b/scripts/agent-health-check.py @@ -1,6 +1,6 @@ #!/usr/bin/env python3 """ -/root/scripts/agent-health-check.py — Consolidated Agent Health Verification v2 +/root/scripts/agent-health-check.py — Consolidated Agent Health Verification v4 Verifies: LiteLLM keys (agent-specific), GPU port conflicts, agent Zulip streaming, gateway liveness, gateway log health, CT liveness, config YAML integrity, @@ -26,9 +26,27 @@ Changelog: Fixed UnboundLocalError on the abiba/koonimo gateway leg (pid unbound in the summary f-string). Abiba's LiteLLM key now comes from /root/.pi/agent/env.sh (#735 agent separation; creds moved out of shared /root/.bashrc). + v4 (2026-09-10): probe-drift round 2 (prose-contracts follow-up to #65/#66/#68). + abiba declared pi-only runtime — Hermes-era config/wrapper/gateway checks are + skipped (harness purge). koby declared report_only per the captain's + 2026-08-17 ruling: every koby leg is detected and reported, never counted as a + fleet failure and never repaired. koby's PVE mapping corrected to storepve + (CT 111 tdunna lives on .6 — the old amdpve mapping produced a false + ct-unreachable). The wrapper infisical-path check had two stale-expectation + bugs: it read only the first 20 lines of the wrapper, so koonimo (whose + wrapper does reference /usr/bin/infisical, just past line 20) was falsely + FAILed as "path may be wrong"; and it treated the absence of any infisical + reference as a fault, though koby's wrapper sources the key from + ~/.hermes/.env and never invokes infisical. The check now reads the full + wrapper body, accepts a no-infisical wrapper, and verifies that any absolute + infisical path the wrapper references actually exists. Report-only findings + are surfaced in a machine-readable `report_only` array in --json output, + separate from `failures`. Every run prints absolute execution provenance + (script + cwd) in the header, in the cron ALERT line, and in --json output so + a stale-consumer report is distinguishable from a fault at read time. """ -import subprocess, json, sys, os, time +import subprocess, json, sys, os, time, re, io, contextlib from datetime import datetime LITELLM = "http://192.168.68.116:80" @@ -49,10 +67,18 @@ AGENTS = { "tanko": {"ct": 112, "host": "192.168.68.122", "user": "jerome", "pve": "amdpve", "vault_key": "TANKO_LITELLM_API_KEY", "runtime": "dsh"}, # abiba = pi agent (.24) — no vault key; its LiteLLM key is read from its # local env file (key_env below), not from the shared vault or .bashrc. + # runtime=pi: abiba has run pi-only since the harness purge. There is no + # Hermes gateway, no ~/.hermes/config.yaml and no hermes CLI wrapper on .24 + # (the /root/.local/bin/hermes symlink is dangling), so the Hermes-era + # config/wrapper/gateway legs are skipped rather than reported as faults. "abiba": {"ct": 100, "host": "192.168.68.24", "user": "root", "pve": "minipve", - "vault_key": None, + "vault_key": None, "runtime": "pi", "key_env": {"file": "/root/.pi/agent/env.sh", "var": "LITELLM_API_KEY"}}, - "koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "amdpve", "vault_key": "KOBY_LITELLM_API_KEY"}, + # koby = report-only (captain's 2026-08-17 ruling, Rule 17): detect and + # report, NEVER repair, and never count against fleet failures. CT 111 + # (tdunna) lives on storepve (.6) — verified live 2026-09-10; the previous + # amdpve mapping made `pct status 111` fail and read as ct-unreachable. + "koby": {"ct": 111, "host": "192.168.68.129", "user": "root", "pve": "storepve", "vault_key": "KOBY_LITELLM_API_KEY", "report_only": True}, "koonimo": {"ct": 113, "host": "192.168.68.114", "user": "root", "pve": "amdpve", "vault_key": "KOONIMO_LITELLM_API_KEY"}, } @@ -69,6 +95,25 @@ GPU_HOSTS = { } FAIL = [] +REPORT_ONLY = [] + + +def _fail(key, agent_name=None): + """Record a failure, except for report-only agents. + + Koby is report-only per the captain's 2026-08-17 ruling (Rule 17): its legs + are detected and reported, never repaired and never counted as fleet + failures. A red fleet alert on a known report-only leg is a false alarm. + Report-only findings are tracked separately so --json consumers can still + see them without them counting as fleet failures. Any non-report-only agent + (or a leg with no agent, e.g. GPU hosts) records normally. + """ + if agent_name and AGENTS.get(agent_name, {}).get("report_only"): + REPORT_ONLY.append(key) + print(f" 🔍 report-only ({agent_name}): {key} — reported, not counted/repaired") + return + FAIL.append(key) + INFISICAL_TOKEN = os.environ.get("INFISICAL_TOKEN") INFISICAL_API_URL = os.environ.get("INFISICAL_API_URL", "https://vault.sysloggh.net") @@ -203,17 +248,21 @@ def _read_env_export(path, var): return None -# Inject keys for each agent: -# - vault-backed agents (tanko/koby/koonimo): {NAME}_LITELLM_API_KEY from -# Infisical (project 322fceab-39da-4854-a55a-568e76c0f13f, env prod). -# - abiba (pi agent, no vault key): LITELLM_API_KEY from its local env file -# /root/.pi/agent/env.sh (moved there from /root/.bashrc in #735). -for agent_name in AGENTS: - info = AGENTS[agent_name] - key = _get_agent_key(agent_name, info.get("vault_key")) - if not key and info.get("key_env"): - key = _read_env_export(info["key_env"]["file"], info["key_env"]["var"]) - AGENTS[agent_name]["key"] = key +def load_agent_keys(): + """Populate AGENTS[*]["key"] from the vault or the agent's local env file. + + Called from main(), not at import: keeping this out of module scope lets the + module be imported (and unit tested) without live vault/SSH access. Vault + format is {NAME}_LITELLM_API_KEY (project 322fceab-39da-4854-a55a-568e76c0f13f, + env prod); abiba has no vault key and reads LITELLM_API_KEY from its local + /root/.pi/agent/env.sh (moved there from /root/.bashrc in #735). + """ + for agent_name in AGENTS: + info = AGENTS[agent_name] + key = _get_agent_key(agent_name, info.get("vault_key")) + if not key and info.get("key_env"): + key = _read_env_export(info["key_env"]["file"], info["key_env"]["var"]) + AGENTS[agent_name]["key"] = key # ═══════════════════════════════════════════════════════════════════ @@ -225,7 +274,7 @@ def check_keys(): key = agent.get("key") if not key: print(f" ❌ {name}: NO KEY FOUND (vault/env empty or unreachable)") - FAIL.append(f"key:{name}:no-key") + _fail(f"key:{name}:no-key", name) continue data = http_json(f"{LITELLM}/v1/models", headers={"Authorization": f"Bearer {key}"}) @@ -234,7 +283,7 @@ def check_keys(): print(f" ✅ {name}: key valid → {model}") else: print(f" ❌ {name}: KEY FAILURE — auth rejected or unreachable") - FAIL.append(f"key:{name}") + _fail(f"key:{name}", name) # ═══════════════════════════════════════════════════════════════════ @@ -294,12 +343,17 @@ def check_agents(): # Tanko runs on DSH (DeepSeek Harness) since 2026-08-27 — it no longer runs a # Hermes gateway, so skip the Hermes gateway/state/streaming/journal checks. - if agent.get("runtime") == "dsh": + # Non-Hermes runtimes have no gateway to probe. dsh = Tanko since + # 2026-08-27; pi = abiba since the harness purge (.24 is pi-only). + if agent.get("runtime") in ("dsh", "pi"): + is_dsh = agent.get("runtime") == "dsh" + label = "DSH (DeepSeek Harness)" if is_dsh else "pi-only runtime" + since = "since 2026-08-27" if is_dsh else "since the harness purge" live = ssh(host, "true", user=user) - print(f" {'✅' if live is not None else '❌'} {name}: DSH (DeepSeek Harness) — " - f"no Hermes gateway since 2026-08-27 (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})") + print(f" {'✅' if live is not None else '❌'} {name}: {label} — " + f"no Hermes gateway {since} (CT {ct}, SSH {'OK' if live is not None else 'FAIL'})") if live is None: - FAIL.append(f"unreachable:{name}") + _fail(f"unreachable:{name}", name) continue if not host or not user: @@ -323,7 +377,7 @@ def check_agents(): # Still check gateway status for reporting purposes if pid == "?": print(f" ⚠️ {name}: GATEWAY NOT RUNNING (reported only)") - FAIL.append(f"gateway-down:{name}") + _fail(f"gateway-down:{name}", name) continue else: print(f" ✅ {name}: gateway running (pid={pid}, report-only mode)") @@ -386,12 +440,12 @@ def check_ct_liveness(): status = ssh(pve_ip, f"pct status {ct} 2>/dev/null", user="root") if not status: print(f" ❌ {name} (CT {ct} on {pve_node}): PVE UNREACHABLE") - FAIL.append(f"ct-unreachable:{name}:{pve_ip}") + _fail(f"ct-unreachable:{name}:{pve_ip}", name) elif "running" in status: print(f" ✅ {name} (CT {ct} on {pve_node}): running") elif "stopped" in status: print(f" ❌ {name} (CT {ct} on {pve_node}): STOPPED") - FAIL.append(f"ct-stopped:{name}") + _fail(f"ct-stopped:{name}", name) else: print(f" ⚠️ {name} (CT {ct} on {pve_node}): {status.strip()}") @@ -407,6 +461,9 @@ def check_config_integrity(): if agent.get("runtime") == "dsh": print(f" ⏭️ {name}: DSH — no Hermes config.yaml since 2026-08-27") continue + if agent.get("runtime") == "pi": + print(f" ⏭️ {name}: pi-only runtime — no Hermes config.yaml since the harness purge") + continue host = agent.get("host") user = agent.get("user") if not host or not user: @@ -421,18 +478,38 @@ def check_config_integrity(): user=user) if not yaml_ok: print(f" ❌ {name}: SSH UNREACHABLE (config check skipped)") - FAIL.append(f"config-unreachable:{name}") + _fail(f"config-unreachable:{name}", name) elif "OK" in yaml_ok: print(f" ✅ {name}: config.yaml valid YAML") else: print(f" ❌ {name}: config.yaml YAML ERROR — {yaml_ok[:120]}") - FAIL.append(f"config-yaml-error:{name}") + _fail(f"config-yaml-error:{name}", name) # ═══════════════════════════════════════════════════════════════════ # CHECK 6: Wrapper/CLI Integrity (NEW) # ═══════════════════════════════════════════════════════════════════ +def _infisical_invocation_paths(wrapper_body): + """Absolute infisical paths the wrapper actually invokes. + + Only executed (non-comment) lines count, and only a path followed by a real + infisical subcommand (e.g. `/usr/bin/infisical run`) is treated as an + invocation. A note such as `# migrated from /usr/local/bin/infisical` is + prose, not a call, so it must not manufacture a dangling-path false alarm. + """ + paths = [] + for line in wrapper_body.splitlines(): + code = line.split("#", 1)[0] + for _m in re.finditer( + r"(/[A-Za-z0-9._/-]*infisical)\s+(?:run|export|secrets|login|logout)\b", + code, + ): + if _m.group(1) not in paths: + paths.append(_m.group(1)) + return paths + + def check_wrapper_integrity(): """Verify the hermes CLI wrapper exists and can reach hermes-real.""" for name, agent in AGENTS.items(): @@ -440,6 +517,9 @@ def check_wrapper_integrity(): if agent.get("runtime") == "dsh": print(f" ⏭️ {name}: DSH — no hermes CLI wrapper since 2026-08-27") continue + if agent.get("runtime") == "pi": + print(f" ⏭️ {name}: pi-only runtime — no hermes CLI wrapper since the harness purge") + continue host = agent.get("host") user = agent.get("user") if not host or not user: @@ -453,24 +533,56 @@ def check_wrapper_integrity(): wrapper = ssh(host, "which hermes 2>/dev/null; command -v hermes 2>/dev/null", user=user) if not wrapper: print(f" ❌ {name}: NO HERMES CLI WRAPPER FOUND") - FAIL.append(f"wrapper-missing:{name}") + _fail(f"wrapper-missing:{name}", name) continue else: print(f" ⚠️ {name}: hermes at {wrapper.strip()} (not ~/.local/bin/hermes)") - # Check wrapper has correct infisical path - infisical_path_valid = ssh(host, - "head -20 /root/.local/bin/hermes 2>/dev/null | grep -q '/usr/bin/infisical' && echo OK || echo MISS", - user=user) - if infisical_path_valid == "MISS": - # Check if infisical exists on path - inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user) - if not inf_actual: - print(f" ❌ {name}: INFISICAL NOT INSTALLED (wrapper broken)") - FAIL.append(f"wrapper-no-infisical:{name}") + # Credential-injection mechanism. The Hermes-era wrapper injected creds + # with `/usr/bin/infisical run`, but the mechanism is not required to be + # infisical at all: koby's wrapper sources the key from ~/.hermes/.env + # and never mentions infisical, which is valid. The old check read only + # the first 20 lines, so koonimo's wrapper — which DOES reference + # /usr/bin/infisical, just past line 20 — false-failed as "path may be + # wrong". Read the full body, accept a no-infisical wrapper, and verify + # the absolute infisical path(s) the wrapper actually invokes. Only + # executed (non-comment) lines count: a comment or dead prose mentioning + # a removed path (litellm-api-keys.prose.md documents + # `rm -f /usr/local/bin/infisical`) must neither produce a dangling path + # nor trigger the PATH check — it is not an invocation. + wrapper_body = ssh(host, "cat /root/.local/bin/hermes 2>/dev/null", user=user) or "" + wrapper_code = "\n".join(line.split("#", 1)[0] for line in wrapper_body.splitlines()) + invoked_paths = _infisical_invocation_paths(wrapper_body) + if "infisical" in wrapper_code: + if invoked_paths: + missing = [] + for _p in invoked_paths: + _exists = ssh(host, f"test -x {_p} && echo OK || echo MISS", user=user) + if not _exists or _exists.strip().splitlines()[-1] != "OK": + missing.append(_p) + if len(missing) == len(invoked_paths): + inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user) + suffix = f" (infisical at {inf_actual})" if inf_actual else "" + print(f" ❌ {name}: wrapper invokes infisical via missing path(s) " + f"{', '.join(missing)}{suffix}") + _fail(f"wrapper-infisical-path:{name}", name) + elif missing: + print(f" ⚠️ {name}: wrapper has an unused/missing infisical path " + f"({', '.join(missing)}) but a working invocation — informational") + elif "/usr/bin/infisical" not in invoked_paths: + print(f" ⚠️ {name}: wrapper infisical path differs " + f"({', '.join(invoked_paths)}) — informational") + else: + print(f" ✅ {name}: wrapper infisical path OK") else: - print(f" ⚠️ {name}: wrapper infisical path may be wrong (infisical at {inf_actual})") - FAIL.append(f"wrapper-infisical-path:{name}") + inf_actual = ssh(host, "command -v infisical 2>/dev/null", user=user) + if not inf_actual: + print(f" ❌ {name}: wrapper invokes infisical but the binary is MISSING") + _fail(f"wrapper-no-infisical:{name}", name) + else: + print(f" ✅ {name}: wrapper infisical resolves via PATH ({inf_actual})") + else: + print(f" ℹ️ {name}: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK") # Check hermes-real exists hermes_real = ssh(host, @@ -483,7 +595,7 @@ def check_wrapper_integrity(): user=user) if not hermes_real or hermes_real.strip() == "MISS": print(f" ❌ {name}: hermes-real NOT FOUND (wrapper broken)") - FAIL.append(f"wrapper-no-hermes-real:{name}") + _fail(f"wrapper-no-hermes-real:{name}", name) else: print(f" ✅ {name}: hermes-real at alt path") @@ -511,10 +623,10 @@ def check_vault_secrets(): key = agent.get("key") if not key: print(f" ❌ {name}: vault secret {vault_key_name} MISSING or EMPTY") - FAIL.append(f"vault-empty:{name}:{vault_key_name}") + _fail(f"vault-empty:{name}:{vault_key_name}", name) elif not key.startswith("sk-"): print(f" ❌ {name}: vault secret {vault_key_name} WRONG FORMAT (starts '{key[:8]}...')") - FAIL.append(f"vault-bad-format:{name}:{vault_key_name}") + _fail(f"vault-bad-format:{name}:{vault_key_name}", name) else: print(f" ✅ {name}: vault {vault_key_name}=sk-...{key[-4:]}") @@ -547,18 +659,7 @@ def deploy_self(): # MAIN # ═══════════════════════════════════════════════════════════════════ -def main(): - quiet = "--quiet" in sys.argv - as_json = "--json" in sys.argv - - # Self-deploy to canonical location - if not quiet and "--no-deploy" not in sys.argv: - deploy_self() - - if not quiet: - print(f"🏥 Agent Health Check v2 — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}") - print() - +def _run_checks(): print("🔑 LiteLLM Keys:") check_keys() print() @@ -586,16 +687,49 @@ def main(): print("🔐 Vault Secrets:") check_vault_secrets() + +def main(): + quiet = "--quiet" in sys.argv + as_json = "--json" in sys.argv + + # Self-deploy to canonical location + if not quiet and "--no-deploy" not in sys.argv: + deploy_self() + + # Provenance: a report is only actionable if the reader can tell WHICH copy + # of this script produced it. A normal run carries it in the header, --json + # carries it for machine consumers, and the cron ALERT line carries it on + # failure. --quiet is documented as "only output on failure", so the header + # is emitted only when not quiet and a healthy quiet run stays silent. + script_path = os.path.abspath(__file__) + cwd = os.getcwd() + + if quiet: + captured = io.StringIO() + with contextlib.redirect_stdout(captured): + load_agent_keys() + _run_checks() + if FAIL: + sys.stdout.write(captured.getvalue()) + else: + print(f"🏥 Agent Health Check v4 — {datetime.now().strftime('%Y-%m-%d %H:%M UTC')}") + print(f"📍 executed from: script={script_path} cwd={cwd}") + print() + load_agent_keys() + _run_checks() + if FAIL: print(f"\n❌ {len(FAIL)} FAILURE(S): {' | '.join(FAIL)}") if quiet: - print(f"ALERT agent-health:{','.join(FAIL)}") + print(f"ALERT agent-health:{','.join(FAIL)} script={script_path} cwd={cwd}") elif not quiet: print("\n✅ All checks passed") if as_json: print(json.dumps({"timestamp": datetime.now().isoformat(), - "failures": FAIL, "healthy": len(FAIL) == 0})) + "execution_path": script_path, "cwd": cwd, + "failures": FAIL, "report_only": REPORT_ONLY, + "healthy": len(FAIL) == 0})) sys.exit(1 if FAIL else 0) diff --git a/scripts/prose-lint.sh b/scripts/prose-lint.sh index 6874267..1756856 100755 --- a/scripts/prose-lint.sh +++ b/scripts/prose-lint.sh @@ -57,7 +57,7 @@ echo "── 2. Regression detection ──" # Grafana /grafana/ as nginx route or URL path (reverted 2026-07-02) # EXCLUDE: filesystem paths (/opt/monitoring/grafana/...), directory creation, revert docs -GRAFANA_HITS=$(grep -rn '/grafana/' *.prose.md 2>/dev/null \ +GRAFANA_HITS=$(grep -rn '/grafana/' ./*.prose.md 2>/dev/null \ | grep -v '/opt/monitoring/grafana/' \ | grep -v 'was tried and reverted\|was reverted\|do not re-add\|NOT recommended' \ | grep -v 'mkdir.*grafana\|Create.*grafana' \ @@ -71,7 +71,7 @@ else fi # Stale CT IDs (CT 122, CT 123 as CT IDs — not IPs .122, .123) -CT_STALE=$(grep -rn '\bCT 122\b' *.prose.md 2>/dev/null || true) +CT_STALE=$(grep -rn '\bCT 122\b' ./*.prose.md 2>/dev/null || true) if [ -n "$CT_STALE" ]; then echo " ❌ REGRESSION: CT 122 used as CT ID — Tanko is CT 112" echo "$CT_STALE" @@ -84,6 +84,35 @@ fi # .122/.123 are correct — verified reachable bridge IPs for Tanko/Mumuni echo " ✅ IP consistency verified (.19=.122=.123 all reachable)" +# Report provenance — every contract report must state the absolute path it +# executed from, so a stale-consumer report is distinguishable from a real fault +# at read time (2026-09-09 probe-drift incident: three false DEGRADED rounds). +# Enforced only inside the **Report format** paragraph, and a check-health +# contract with no Report format paragraph FAILs rather than being skipped. +PROV_FILES=$(grep -rlE '^### check-health|\*\*Report format\*\*' ./*.prose.md 2>/dev/null || true) +if [ -z "$PROV_FILES" ]; then + echo " ❌ No check-health/report-format contracts found — provenance not enforced" + FAILED=1 +else + PROV_BAD=0 + while IFS= read -r f; do + [ -n "$f" ] || continue + REPORT_PARA=$(awk '/\*\*Report format\*\*/{found=1} found{print} found && /^[[:space:]]*$/{exit}' "$f") + if [ -z "$REPORT_PARA" ]; then + echo " ❌ $f: check-health contract has no **Report format** paragraph" + PROV_BAD=1 + elif ! printf '%s\n' "$REPORT_PARA" | grep -qE 'absolute path|pwd -P|executed from'; then + echo " ❌ $f: **Report format** lacks execution provenance (absolute path / pwd -P)" + PROV_BAD=1 + fi + done <<< "$PROV_FILES" + if [ "$PROV_BAD" -eq 1 ]; then + FAILED=1 + else + echo " ✅ Report provenance present in all report-format contracts" + fi +fi + # ── 3. Cross-contract consistency ── echo "" echo "── 3. Cross-contract consistency ──" diff --git a/tests/test_probe_drift.py b/tests/test_probe_drift.py new file mode 100644 index 0000000..28a4e47 --- /dev/null +++ b/tests/test_probe_drift.py @@ -0,0 +1,466 @@ +"""Regression tests for the 2026-09-09/10 probe-drift corrections. + +WHY THIS FILE EXISTS: the monitoring contracts kept emitting false alarms from +stale expectations rather than live faults. + + * agent-health-check v3 reported 6 failures that were all stale expectations: + abiba (pi-only since the harness purge) was tested as a Hermes host, koby + (report-only per the captain's 2026-08-17 ruling) was counted as repairable, + koby's CT 111 was probed on amdpve where it does not exist (it runs on + storepve .6), and the wrapper infisical check had two bugs — it read only + the first 20 lines, so koonimo's wrapper (which references /usr/bin/infisical + past line 20) false-failed, and it treated koby's genuine no-infisical + (~/.hermes/.env) wrapper as broken. + * gpu-monitor emitted "DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000" three + times from probing bare port 80 on GPU hosts while :8080 answered 200. + * infrastructure-monitoring probed CT 116 for the PVE API (no pveproxy -> + 000) instead of the five real cluster nodes, which answer 401 = alive. + +These tests execute the health script (with SSH/vault stubbed) and the real +provenance consumer (scripts/prose-lint.sh), and parse the contracts' executable +check-health probe blocks into normalized probe sets. No live network, vault, or +SSH access is required. +""" +from __future__ import annotations + +import importlib.util +import json +import os +import pathlib +import re +import subprocess +import sys +import textwrap + +import pytest + +ROOT = pathlib.Path(__file__).resolve().parents[1] +AHC = ROOT / "scripts" / "agent-health-check.py" +LINT = ROOT / "scripts" / "prose-lint.sh" +GPU = ROOT / "gpu-monitor.prose.md" +INFRA = ROOT / "infrastructure-monitoring.prose.md" + +PVE_NODE_IPS = { + "192.168.68.9", + "192.168.68.5", + "192.168.68.15", + "192.168.68.6", + "192.168.68.12", +} + + +@pytest.fixture(scope="module") +def ahc(): + """Import agent-health-check.py without live network/SSH side effects.""" + spec = importlib.util.spec_from_file_location("agent_health_check", AHC) + assert spec and spec.loader + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +# ── helpers: execute the health script with SSH/vault stubbed ───────── + +def _run_main(ahc, monkeypatch, capsys, argv, ssh_result=None): + ahc.FAIL.clear() + ahc.REPORT_ONLY.clear() + monkeypatch.setattr(ahc, "load_agent_keys", lambda: None) + monkeypatch.setattr(ahc, "ssh", lambda *a, **k: ssh_result) + monkeypatch.setattr(sys, "argv", ["agent-health-check.py", "--no-deploy", *argv]) + with pytest.raises(SystemExit) as exc: + ahc.main() + return exc.value.code, capsys.readouterr().out + + +def _json_payload(out): + for line in reversed(out.splitlines()): + if line.startswith('{"timestamp"'): + return json.loads(line) + raise AssertionError(f"no JSON payload in output:\n{out}") + + +# ── agent-health-check: stale-expectation legs ─────────────────────── + +def test_import_does_not_contact_vault(ahc): + # Keys are loaded in main() via load_agent_keys(); importing must stay inert. + assert callable(ahc.load_agent_keys) + assert all(agent.get("key") is None for agent in ahc.AGENTS.values()) + + +def test_abiba_is_pi_only_runtime(ahc): + # .24 has run pi-only since the harness purge: no Hermes gateway, config, or + # wrapper. Probing those legs produced false failures. + assert ahc.AGENTS["abiba"]["runtime"] == "pi" + + +def test_koby_is_report_only(ahc): + # Captain's 2026-08-17 ruling (Rule 17): detect and report, never repair. + assert ahc.AGENTS["koby"]["report_only"] is True + + +def test_koby_ct111_is_on_storepve(ahc): + # Live-verified 2026-09-10: `pct status 111` = running on storepve (.6); + # amdpve has no lxc/111.conf, which is what false-failed before. + assert ahc.AGENTS["koby"]["pve"] == "storepve" + + +def test_report_only_legs_never_count_as_failures(ahc): + for agent, report_only in (("koby", True), ("koonimo", False), ("tanko", False)): + ahc.FAIL.clear() + ahc.REPORT_ONLY.clear() + ahc._fail(f"probe:{agent}", agent) + if report_only: + assert ahc.FAIL == [] + assert ahc.REPORT_ONLY == [f"probe:{agent}"] + else: + assert ahc.FAIL == [f"probe:{agent}"] + assert ahc.REPORT_ONLY == [] + ahc.FAIL.clear() + ahc.REPORT_ONLY.clear() + + +def test_failure_recording_accepts_agentless_keys(ahc): + ahc.FAIL.clear() + try: + ahc._fail("gpu-no-port:gpu-rtx3090 (.8)") + assert ahc.FAIL == ["gpu-no-port:gpu-rtx3090 (.8)"] + finally: + ahc.FAIL.clear() + + +def test_json_reports_absolute_execution_provenance(ahc, monkeypatch, capsys): + code, out = _run_main(ahc, monkeypatch, capsys, ["--json"]) + payload = _json_payload(out) + assert payload["execution_path"] == os.path.abspath(str(AHC)) + assert payload["cwd"] == os.getcwd() + assert code == 1 # stubbed SSH fails every leg, but provenance is still emitted + + +def test_quiet_run_still_carries_provenance_on_the_alert_path(ahc, monkeypatch, capsys): + # The cron runs --quiet; a failure report must still carry provenance. The + # header line is suppressed in quiet mode, so the ALERT line is the carrier. + code, out = _run_main(ahc, monkeypatch, capsys, ["--quiet"]) + assert code == 1 + assert "📍 executed from:" not in out + alerts = [ln for ln in out.splitlines() if ln.startswith("ALERT agent-health:")] + assert alerts, out + assert f"script={os.path.abspath(str(AHC))}" in alerts[0] + assert f"cwd={os.getcwd()}" in alerts[0] + + +def test_quiet_healthy_run_emits_no_stdout(ahc, monkeypatch, capsys): + # --quiet is documented as "only output on failure": a run with no fleet + # failures must produce no stdout at all (the production cron runs --quiet). + for name in ("check_keys", "check_gpu_ports", "check_agents", "check_ct_liveness", + "check_config_integrity", "check_wrapper_integrity", "check_vault_secrets"): + monkeypatch.setattr(ahc, name, lambda: None) + code, out = _run_main(ahc, monkeypatch, capsys, ["--quiet"]) + assert code == 0 + assert out == "" + + +def test_json_surfaces_report_only_findings_separately(ahc, monkeypatch, capsys): + # Koby's down legs are reported but must not count as fleet failures; the + # --json payload exposes them in their own array (item 1 + f8). + _, out = _run_main(ahc, monkeypatch, capsys, ["--json"]) + payload = _json_payload(out) + assert isinstance(payload["report_only"], list) + assert any(key.startswith(("gateway-down:koby", "ct-unreachable:koby")) + for key in payload["report_only"]) + assert not any("koby" in key for key in payload["failures"]) + + +# ── agent-health-check: wrapper infisical behavior (f3) ─────────────── + +def _stub_wrapper_ssh(ahc, monkeypatch, wrapper_body, test_x_result="OK", command_v="/usr/local/bin/infisical"): + def fake_ssh(host, cmd, user="root"): + if cmd.startswith("cat /root/.local/bin/hermes"): + return wrapper_body + if cmd.startswith("ls -la /root/.local/bin/hermes "): + return "-rwxr-xr-x 1 root root 0 Jan 1 00:00 /root/.local/bin/hermes" + if cmd.startswith("ls -la /root/.local/bin/hermes-real") or "venv/bin/hermes" in cmd: + return "-rwxr-xr-x 1 root root 0 Jan 1 00:00 /root/.local/bin/hermes-real" + if cmd.startswith("grep -c 'LITELLM_API_KEY'"): + return "1" + if cmd.startswith("test -x "): + path = cmd[len("test -x "):].split()[0] + if isinstance(test_x_result, dict): + return test_x_result.get(path, "MISS") + return test_x_result + if cmd.startswith("command -v infisical"): + return command_v + return None + + monkeypatch.setattr(ahc, "ssh", fake_ssh) + monkeypatch.setattr(ahc, "AGENTS", {"koonimo": dict(ahc.AGENTS["koonimo"])}) + ahc.FAIL.clear() + ahc.REPORT_ONLY.clear() + + +def test_env_based_wrapper_without_infisical_is_not_failed(ahc, monkeypatch, capsys): + _stub_wrapper_ssh(ahc, monkeypatch, + "#!/bin/bash\nsource ~/.hermes/.env\nexec hermes-real \"$@\"\n") + ahc.check_wrapper_integrity() + out = capsys.readouterr().out + assert ahc.FAIL == [] + assert "wrapper resolves creds without infisical" in out + + +def test_dangling_absolute_infisical_path_is_failed(ahc, monkeypatch, capsys): + # Wrapper hardcodes /usr/bin/infisical, which is absent, while PATH resolves + # infisical to /usr/local/bin/infisical. The literal path must be verified, + # not inferred from PATH resolution. + _stub_wrapper_ssh(ahc, monkeypatch, + "#!/bin/bash\n/usr/bin/infisical run -- hermes-real \"$@\"\n", + test_x_result="MISS", command_v="/usr/local/bin/infisical") + ahc.check_wrapper_integrity() + assert "wrapper-infisical-path:koonimo" in ahc.FAIL + + +def test_existing_absolute_infisical_path_passes(ahc, monkeypatch, capsys): + _stub_wrapper_ssh(ahc, monkeypatch, + "#!/bin/bash\n/usr/bin/infisical run -- hermes-real \"$@\"\n", + test_x_result="OK") + ahc.check_wrapper_integrity() + out = capsys.readouterr().out + assert ahc.FAIL == [] + assert "wrapper infisical path OK" in out + + +def test_comment_mentioning_removed_infisical_path_is_not_failed(ahc, monkeypatch, capsys): + # litellm-api-keys.prose.md documents `rm -f /usr/local/bin/infisical`; a + # wrapper comment about that migration must not manufacture a dangling path + # when the real invocation (/usr/bin/infisical) is present and executable. + _stub_wrapper_ssh(ahc, monkeypatch, + "#!/bin/bash\n# migrated from /usr/local/bin/infisical\n" + "exec /usr/bin/infisical run -- hermes-real \"$@\"\n", + test_x_result={"/usr/bin/infisical": "OK", + "/usr/local/bin/infisical": "MISS"}) + ahc.check_wrapper_integrity() + out = capsys.readouterr().out + assert ahc.FAIL == [] + assert "wrapper infisical path OK" in out + + +def test_comment_only_infisical_mention_does_not_reach_path_check(ahc, monkeypatch, capsys): + # A comment-only mention of a removed infisical path on a healthy .env-based + # wrapper is not an invocation: it must not fall through to the `command -v` + # PATH check and false-FAIL `wrapper-no-infisical`. + _stub_wrapper_ssh(ahc, monkeypatch, + "#!/bin/bash\n# migrated from /usr/local/bin/infisical\n" + "source ~/.hermes/.env\nexec hermes-real \"$@\"\n", + test_x_result="MISS", command_v=None) + ahc.check_wrapper_integrity() + out = capsys.readouterr().out + assert ahc.FAIL == [] + assert "wrapper resolves creds without infisical" in out + + +# ── item 4: prose-lint enforces report provenance (real consumer) ───── + +GOOD_CONTRACT = textwrap.dedent("""\ + --- + kind: function + name: good + description: fixture with provenance + --- + + ## Parameters + + - x: y + + ## Returns + + ok + + ### check-health + + ```bash + pwd -P + ``` + + **Report format**: Begin with the absolute path the probe executed from. + """) + +DECOY_CONTRACT = textwrap.dedent("""\ + --- + kind: function + name: decoy + description: fixture with provenance only outside the report format + --- + + ## Parameters + + - x: y + + ## Returns + + ok + + The absolute path of the config is /etc/foo. + + ### check-health + + ```bash + true + ``` + + **Report format**: Summarize actual results from each probe. + """) + +MISSING_CONTRACT = textwrap.dedent("""\ + --- + kind: function + name: missing + description: check-health contract with no report format + --- + + ## Parameters + + - x: y + + ## Returns + + ok + + ### check-health + + ```bash + pwd -P + ``` + """) + + +def _run_lint(tmp_path, text, name): + (tmp_path / name).write_text(text) + return subprocess.run(["bash", str(LINT)], cwd=tmp_path, + capture_output=True, text=True) + + +def test_prose_lint_accepts_report_format_with_provenance(tmp_path): + result = _run_lint(tmp_path, GOOD_CONTRACT, "good.prose.md") + assert result.returncode == 0, result.stdout + result.stderr + + +def test_prose_lint_rejects_report_format_without_provenance(tmp_path): + result = _run_lint(tmp_path, DECOY_CONTRACT, "decoy.prose.md") + assert result.returncode == 1, result.stdout + assert "lacks execution provenance" in result.stdout + + +def test_prose_lint_requires_report_format_on_check_health_contract(tmp_path): + result = _run_lint(tmp_path, MISSING_CONTRACT, "missing.prose.md") + assert result.returncode == 1, result.stdout + assert "no **Report format** paragraph" in result.stdout + + +# ── contracts: parse the executable check-health probe block ───────── + +def _check_health_block(contract): + """Extract the bash probe block under ### check-health (the probe interface).""" + text = contract.read_text() + marker = "### check-health" + assert marker in text, f"{contract.name} has no {marker}" + after = text.split(marker, 1)[1] + match = re.search(r"```bash\n(.*?)```", after, re.S) + assert match, f"{contract.name} check-health has no bash probe block" + return match.group(1) + + +def _loop_nodes(block): + nodes = [] + for line in block.splitlines(): + match = re.match(r"\s*for\s+\w+\s+in\s+(.+?);?\s*do\b", line) + if match: + nodes = match.group(1).split() + return nodes + + +def _record(url): + """Normalize a URL into a probe record: host, port, path, expected status.""" + match = re.match(r"https?://([^/\s\"')]+)(/[^\s\"')]*)?", url) + assert match, f"unparseable probe URL: {url}" + hostport = match.group(1) + if "@" in hostport: + hostport = hostport.split("@", 1)[1] + if hostport.startswith("["): + host, port = hostport[1:hostport.index("]")], None + elif ":" in hostport: + host, raw_port = hostport.rsplit(":", 1) + port = int(raw_port) if raw_port.isdigit() else None + else: + host, port = hostport, None + return {"host": host, "port": port, "path": match.group(2) or "/", + "expected": None} + + +def _probes(block): + """Parse the executable check-health bash block into a normalized probe model. + + Comments are not probes; an `# Expected: ` comment annotates the + preceding probe. URLs using the block's shell-loop variable `$node` are + expanded over the loop's node list. + """ + loop_nodes = _loop_nodes(block) + probes = [] + last = None + for raw in block.splitlines(): + stripped = raw.strip() + if stripped.startswith("#"): + expected = re.search(r"Expected:\s*(\d{3})", stripped, re.I) + if expected and last is not None: + last["expected"] = int(expected.group(1)) + continue + for url in re.findall(r"https?://[^\s\"')]+", raw): + hosts = loop_nodes if "$node" in url else [None] + for node in hosts: + record = _record(url.replace("$node", node) if node else url) + probes.append(record) + last = record + return probes + + +def test_gpu_monitor_probes_every_gpu_health_on_8080(): + probes = _probes(_check_health_block(GPU)) + targets = {(p["host"], p["port"], p["path"]) for p in probes} + assert ("192.168.68.8", 8080, "/health") in targets + assert ("192.168.68.110", 8080, "/health") in targets + + +def test_gpu_monitor_never_probes_bare_port_80_on_gpu_hosts(): + probes = _probes(_check_health_block(GPU)) + gpu_hosts = {"192.168.68.8", "192.168.68.110", "192.168.68.15"} + offenders = [p for p in probes + if p["host"] in gpu_hosts and p["port"] in (None, 80)] + assert offenders == [] + + +def test_probe_model_flags_explicit_port_80_on_gpu_host(): + # Regression: a bare-port probe may be spelled with an explicit :80. + block = ("curl -s -o /dev/null -w '%{http_code}' " + "http://192.168.68.8:80/health\n") + gpu_hosts = {"192.168.68.8", "192.168.68.110", "192.168.68.15"} + offenders = [p for p in _probes(block) + if p["host"] in gpu_hosts and p["port"] in (None, 80)] + assert offenders and offenders[0]["port"] == 80 + + +def test_gpu_monitor_treats_router_301_as_alive(): + probes = _probes(_check_health_block(GPU)) + unified = [p for p in probes + if p["host"] == "192.168.68.116" and p["path"] == "/health/unified"] + assert unified, "router /health/unified probe missing" + assert unified[0]["expected"] == 301 + + +def test_infra_monitoring_probes_every_real_pve_node(): + probes = _probes(_check_health_block(INFRA)) + pve = {(p["host"], p["port"], p["path"]) for p in probes if p["port"] == 8006} + assert {host for host, _, _ in pve} == PVE_NODE_IPS + assert {path for _, _, path in pve} == {"/api2/json/version"} + + +def test_infra_monitoring_does_not_probe_ct116_for_pve_api(): + probes = _probes(_check_health_block(INFRA)) + assert not any(p["host"] == "192.168.68.116" and p["port"] == 8006 + for p in probes)