# Probe-drift round 2 — per-leg before/after evidence > **Historical record** — 2026-09-28: The lines below that describe tanko as > "DSH (DeepSeek Harness)" only reflect what the check reported when it was > running. Tanko's runtime was later found to be **hybrid (DSH + Hermes)** — > the check had a `/root/` hardcoding bug that made it probe the wrong home > directory and report `wrapper-missing:tanko` for an agent with a working > wrapper. This document records the observed output, not the underlying > truth; see `fix/agent-health-root-hardcoding-20260928` for the correction. **Date:** 2026-09-10 **Worktree (absolute execution path):** `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts` **Branch:** `fm/probe-drift-round2-20260909` Every command below was run from the absolute path above; output is pasted verbatim. This is the evidence trail for the four scoped corrections; it is not a contract (never `prose run` it). --- ## Leg 1 — agent-health-check (item 1) **Before** — from `/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts`, `python3 scripts/agent-health-check.py --no-deploy` (v2, base of this branch): ``` 🏥 Agent Health Check v2 — 2026-09-10 01:16 UTC 🔑 LiteLLM Keys: ✅ tanko: key valid → syslog-auto ✅ abiba: key valid → syslog-auto ✅ koby: key valid → syslog-auto ✅ koonimo: key valid → syslog-auto 🎮 GPU Port Health: ✅ gpu-rtx3090 (.8): healthy (pid=472206) ✅ gpu-rtx5070 (.110): healthy (pid=207601) ✅ gpu-strixhalo (.15): healthy (pid=4098872) 🤖 Agent Gateways: ✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK) ⚠️ abiba: gw=no-state-file zulip=? streaming=no errors_10m=0 pid=? ✅ koby: gw=running zulip=connected streaming=no errors_10m=0 pid=360900 ✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125 🖥️ CT Liveness: ✅ tanko (CT 112 on amdpve): running ✅ abiba (CT 100 on minipve): running ❌ koby (CT 111 on amdpve): PVE UNREACHABLE ✅ koonimo (CT 113 on amdpve): running 📝 Config Integrity: ⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27 ✅ abiba: config.yaml valid YAML ✅ koby: config.yaml valid YAML ✅ koonimo: config.yaml valid YAML 🔌 Wrapper/CLI Integrity: ⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27 ⚠️ abiba: wrapper infisical path may be wrong (infisical at /usr/bin/infisical) ❌ abiba: hermes-real NOT FOUND (wrapper broken) ⚠️ abiba: .env may be missing LITELLM_API_KEY entry ⚠️ koby: wrapper infisical path may be wrong (infisical at /usr/bin/infisical) ❌ koby: hermes-real NOT FOUND (wrapper broken) ✅ koby: wrapper + .env key present ⚠️ koonimo: wrapper infisical path may be wrong (infisical at /usr/bin/infisical) ✅ koonimo: wrapper + .env key present 🔐 Vault Secrets: ✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw ✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg ✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ ❌ 6 FAILURE(S): ct-unreachable:koby:192.168.68.15 | wrapper-infisical-path:abiba | wrapper-no-hermes-real:abiba | wrapper-infisical-path:koby | wrapper-no-hermes-real:koby | wrapper-infisical-path:koonimo ``` Root causes (all stale expectations; no live fault): | Failure | Why it was stale | |---|---| | `ct-unreachable:koby:192.168.68.15` | CT 111 (tdunna/koby) runs on **storepve (.6)**, not amdpve (.15). | | `wrapper-*:abiba` | Abiba is pi-only since the harness purge. `/root/.local/bin/hermes` is a dangling symlink; no `hermes-real`, no `~/.hermes/.env`. | | `wrapper-*:koby` | Koby is **report-only** (captain ruling 2026-08-17, Rule 17): detect and report, never repair — its legs must not count as fleet failures. Koby's wrapper is also the genuine no-infisical case: it sources `~/.hermes/.env` rather than `/usr/bin/infisical`, which the check now accepts. | | `wrapper-infisical-path:koonimo` | Koonimo's wrapper **does** reference `/usr/bin/infisical` — but past the old check's `head -20` window, so the check looked for the path in the wrong slice and false-failed. The fix that mattered was reading the full wrapper body (and then verifying any absolute infisical path it finds actually exists). | **After** — same absolute path, `python3 scripts/agent-health-check.py --no-deploy` (v4): ``` $ pwd -P /root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts $ python3 scripts/agent-health-check.py --no-deploy 🏥 Agent Health Check v4 — 2026-09-10 01:22 UTC 📍 executed from: script=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts/scripts/agent-health-check.py cwd=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts 🔑 LiteLLM Keys: ✅ tanko: key valid → syslog-auto ✅ abiba: key valid → syslog-auto ✅ koby: key valid → syslog-auto ✅ koonimo: key valid → syslog-auto 🎮 GPU Port Health: ✅ gpu-rtx3090 (.8): healthy (pid=472206) ✅ gpu-rtx5070 (.110): healthy (pid=207601) ✅ gpu-strixhalo (.15): healthy (pid=4098872) 🤖 Agent Gateways: ✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK) ✅ abiba: pi-only runtime — no Hermes gateway since the harness purge (CT 100, SSH OK) 🔍 koby: REPORT-ONLY mode (diagnostic only, no repairs on .129) ✅ koby: gateway running (pid=360900, report-only mode) ✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125 🖥️ CT Liveness: ✅ tanko (CT 112 on amdpve): running ✅ abiba (CT 100 on minipve): running ✅ koby (CT 111 on storepve): running ✅ koonimo (CT 113 on amdpve): running 📝 Config Integrity: ⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27 ⏭️ abiba: pi-only runtime — no Hermes config.yaml since the harness purge ✅ koby: config.yaml valid YAML ✅ koonimo: config.yaml valid YAML 🔌 Wrapper/CLI Integrity: ⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27 ⏭️ abiba: pi-only runtime — no hermes CLI wrapper since the harness purge ℹ️ koby: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK ❌ koby: hermes-real NOT FOUND (wrapper broken) 🔍 report-only (koby): wrapper-no-hermes-real:koby — reported, not counted/repaired ✅ koby: wrapper + .env key present ✅ koonimo: wrapper infisical path OK ✅ koonimo: wrapper + .env key present 🔐 Vault Secrets: ✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw ✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg ✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ ✅ All checks passed exit=0 ``` **Live vantage proof** (same worktree): ``` $ ssh root@192.168.68.15 "pct status 111" Configuration file 'nodes/amdpve/lxc/111.conf' does not exist $ ssh root@192.168.68.6 "pct status 111; pct list | grep '^ *111'" status: running 111 running tdunna $ ssh root@192.168.68.129 "hostname" tdunna ``` --- ## Leg 2 — infrastructure-monitoring PVE API (item 2) **Before** — the contract's probe, aimed at the monitoring host CT 116: ``` $ curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json 000 ``` CT 116 runs no `pveproxy`, so it never answers on `:8006`. The probe target was wrong, which is what read as PVE-API `000`. **After** — probing the five real cluster nodes (`:8006/api2/json/version`), alive under the any-HTTP-response rule (`401` = up, unauthenticated): ``` $ for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do printf '%s:8006 -> %s\n' "$node" "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")" done 192.168.68.9:8006 -> 401 192.168.68.5:8006 -> 401 192.168.68.15:8006 -> 401 192.168.68.6:8006 -> 401 192.168.68.12:8006 -> 401 ``` `401` on every node = alive by design. `DOWN` is `000`/timeout only. (The contract's LiteLLM probe was the same class: `/litellm/health` answers `301` → `/litellm/health/liveliness`, so it is now specified as any-HTTP too.) --- ## Leg 3 — gpu-monitor GPU probes (item 3) **Before** — the false alarm came from probing bare port 80 on GPU hosts: ``` $ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8/health 000 $ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110/health 000 ``` Nothing listens on GPU port 80, so the monitor reported `DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` three times on 2026-09-09. **After** — the real endpoints answer: ``` $ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health 200 $ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health 200 $ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.15:8080/health 200 $ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified 301 # Location: http://192.168.68.116/gpu/gpu-data — the same payload $ curl -s -o /dev/null -w '%{http_code}' -L http://192.168.68.116/health/unified 200 ``` `301` is healthy under the any-HTTP-response rule. The contract now requires GPU health on `:8080` (or router `/health/unified`) and forbids bare port 80 on a GPU host. --- ## Leg 4 — report provenance (item 4) Every contract report must now lead with the absolute path it executed from. `docs/AUTHORING-GUIDE.md` documents the rule and `scripts/prose-lint.sh` enforces it: ``` $ pwd -P /root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts $ bash scripts/prose-lint.sh ✅ Report provenance present in all report-format contracts ... ✅ LINT PASSED (12 warning(s)) ``` The health script prints `📍 executed from: script=… cwd=…` and includes `execution_path`/`cwd` in `--json` output. --- ## Full suite ``` $ pwd -P /root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts $ python3 -m pytest -q 24 passed $ shellcheck scripts/prose-lint.sh (clean) ``` --- ## Follow-up findings (observed, intentionally NOT changed here) These are adjacent stale expectations discovered while verifying the four scoped legs. Each touches a CRITICAL/HIGH-sensitivity artifact or an unrelated script, so it is recorded for the captain/verify mate rather than silently repaired. 1. **`infrastructure-control.prose.md` (CRITICAL) CT 111 node assignment.** Lines ~109 and ~615 place `tdunna` (CT 111, koby) on **amdpve**. Live verification on 2026-09-10 shows `pct status 111` = `running` on **storepve (.6)** and `Configuration file 'nodes/amdpve/lxc/111.conf' does not exist` on .15. `agent-health-check.py` now carries the live-verified `storepve` mapping (the script is not the topology source of truth); the CRITICAL contract itself needs an authorized correction. **✅ Resolved 2026-09-12:** the topology was corrected in its owner, `infrastructure-control.prose.md` (CT 111 → storepve, CT 105 → amdpve), and `scripts/pct-run.sh` now matches. This snapshot is left as observed; treat those owner documents as authoritative. 2. **Strix Halo `:8080` firewall claim is stale.** `prose-ai-review.sh` ground-truth rule #4 and `gpu-monitor.prose.md` say `:8080` is firewalled to `.116` only and `.24` cannot probe it. Live on .15: `-A INPUT -s 192.168.68.24/32 -p tcp --dport 8080 -j ACCEPT`, and a probe from .24 returns `200`. The contract keeps routing Strix via the router (safe), but the claim no longer matches iptables. 3. **`contract-registry.yaml` references `agent-health-check.prose.md`**, which does not exist in the repo. The registry entry (with `koby_action: skip_heal`) is aspirational/stale. 4. **Pre-existing script defects, untouched:** `scripts/pm2-self-heal.sh` has a bash syntax error at lines 19–20 (`bash -n` fails), and `shellcheck` fails on five untouched scripts (`netbird-add-domain.sh`, `pct-run.sh`, `pm2-self-heal.sh`, `prose-ai-review.sh`, `swap-gpu-dense-model.sh`). `scripts/prose-lint.sh` — the one shell file touched here — is now shellcheck-clean.