Second probe-drift correction pass after #65/#66/#68. All four legs were stale
consumer expectations, not live faults.
1. scripts/agent-health-check.py (v4)
- abiba declared pi-only runtime (harness purge): Hermes-era gateway, config
and wrapper legs are skipped instead of failing.
- koby declared report_only (captain ruling 2026-08-17, Rule 17): every koby
leg is detected and reported, never counted as a fleet failure or repaired.
- koby's CT 111 mapping corrected to storepve (.6); the old amdpve mapping
made `pct status 111` fail and read as ct-unreachable.
- wrapper check no longer FAILs .env-based wrappers that legitimately never
invoke infisical (koonimo).
- keys load in main() (load_agent_keys) so the module is importable/testable.
- every run prints absolute execution provenance (script + cwd), in the header
and in --json.
Before: 6 FAILURE(S). After: 0 failures, koby reported read-only.
2. infrastructure-monitoring.prose.md
- PVE API probe repointed from CT 116 (no pveproxy, 000) to the five real
nodes on https://<node>:8006/api2/json/version, all 401 = alive.
- any-HTTP-response liveness rule added (401/3xx alive; 000/timeout = DOWN).
- LiteLLM health documented as 301 -> /litellm/health/liveliness, not bare 200.
3. gpu-monitor.prose.md
- GPU health probes on :8080 (or router /health/unified); bare port 80 on a
GPU host is forbidden (no listener -> false DEGRADED).
- router /health/unified 301 -> /gpu/gpu-data documented as alive.
- port-discipline + liveness rule + direct-fallback execution step.
4. Report provenance (all contracts)
- docs/AUTHORING-GUIDE.md documents the rule; scripts/prose-lint.sh enforces
that any **Report format** contract states an absolute path (pwd -P).
- provenance added to gpu-monitor, infrastructure-monitoring, proxmox-monitor.
Tests: tests/test_probe_drift.py (23 passed); prose-lint.sh clean + shellcheck
clean; CI frontmatter validation passes. Evidence with per-leg before/after and
absolute paths: docs/probe-drift-round2-evidence.md.
11 KiB
Probe-drift round 2 — per-leg before/after evidence
Date: 2026-09-10
Worktree (absolute execution path): /root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
Branch: fm/probe-drift-round2-20260909
Every command below was run from the absolute path above; output is pasted
verbatim. This is the evidence trail for the four scoped corrections; it is not
a contract (never prose run it).
Leg 1 — agent-health-check (item 1)
Before — from /root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts,
python3 scripts/agent-health-check.py --no-deploy (v2, base of this branch):
🏥 Agent Health Check v2 — 2026-09-10 01:16 UTC
🔑 LiteLLM Keys:
✅ tanko: key valid → syslog-auto
✅ abiba: key valid → syslog-auto
✅ koby: key valid → syslog-auto
✅ koonimo: key valid → syslog-auto
🎮 GPU Port Health:
✅ gpu-rtx3090 (.8): healthy (pid=472206)
✅ gpu-rtx5070 (.110): healthy (pid=207601)
✅ gpu-strixhalo (.15): healthy (pid=4098872)
🤖 Agent Gateways:
✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK)
⚠️ abiba: gw=no-state-file zulip=? streaming=no errors_10m=0 pid=?
✅ koby: gw=running zulip=connected streaming=no errors_10m=0 pid=360900
✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125
🖥️ CT Liveness:
✅ tanko (CT 112 on amdpve): running
✅ abiba (CT 100 on minipve): running
❌ koby (CT 111 on amdpve): PVE UNREACHABLE
✅ koonimo (CT 113 on amdpve): running
📝 Config Integrity:
⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27
✅ abiba: config.yaml valid YAML
✅ koby: config.yaml valid YAML
✅ koonimo: config.yaml valid YAML
🔌 Wrapper/CLI Integrity:
⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27
⚠️ abiba: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
❌ abiba: hermes-real NOT FOUND (wrapper broken)
⚠️ abiba: .env may be missing LITELLM_API_KEY entry
⚠️ koby: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
❌ koby: hermes-real NOT FOUND (wrapper broken)
✅ koby: wrapper + .env key present
⚠️ koonimo: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
✅ koonimo: wrapper + .env key present
🔐 Vault Secrets:
✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw
✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg
✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ
❌ 6 FAILURE(S): ct-unreachable:koby:192.168.68.15 | wrapper-infisical-path:abiba | wrapper-no-hermes-real:abiba | wrapper-infisical-path:koby | wrapper-no-hermes-real:koby | wrapper-infisical-path:koonimo
Root causes (all stale expectations; no live fault):
| Failure | Why it was stale |
|---|---|
ct-unreachable:koby:192.168.68.15 |
CT 111 (tdunna/koby) runs on storepve (.6), not amdpve (.15). |
wrapper-*:abiba |
Abiba is pi-only since the harness purge. /root/.local/bin/hermes is a dangling symlink; no hermes-real, no ~/.hermes/.env. |
wrapper-*:koby |
Koby is report-only (captain ruling 2026-08-17, Rule 17): detect and report, never repair — its legs must not count as fleet failures. |
wrapper-infisical-path:koonimo |
Koonimo's wrapper injects KOONIMO_LITELLM_API_KEY from ~/.hermes/.env and never invokes infisical. The check required /usr/bin/infisical, which is only one valid mechanism. |
After — same absolute path, python3 scripts/agent-health-check.py --no-deploy (v4):
$ pwd -P
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
$ python3 scripts/agent-health-check.py --no-deploy
🏥 Agent Health Check v4 — 2026-09-10 01:22 UTC
📍 executed from: script=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts/scripts/agent-health-check.py cwd=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
🔑 LiteLLM Keys:
✅ tanko: key valid → syslog-auto
✅ abiba: key valid → syslog-auto
✅ koby: key valid → syslog-auto
✅ koonimo: key valid → syslog-auto
🎮 GPU Port Health:
✅ gpu-rtx3090 (.8): healthy (pid=472206)
✅ gpu-rtx5070 (.110): healthy (pid=207601)
✅ gpu-strixhalo (.15): healthy (pid=4098872)
🤖 Agent Gateways:
✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK)
✅ abiba: pi-only runtime — no Hermes gateway since the harness purge (CT 100, SSH OK)
🔍 koby: REPORT-ONLY mode (diagnostic only, no repairs on .129)
✅ koby: gateway running (pid=360900, report-only mode)
✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125
🖥️ CT Liveness:
✅ tanko (CT 112 on amdpve): running
✅ abiba (CT 100 on minipve): running
✅ koby (CT 111 on storepve): running
✅ koonimo (CT 113 on amdpve): running
📝 Config Integrity:
⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27
⏭️ abiba: pi-only runtime — no Hermes config.yaml since the harness purge
✅ koby: config.yaml valid YAML
✅ koonimo: config.yaml valid YAML
🔌 Wrapper/CLI Integrity:
⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27
⏭️ abiba: pi-only runtime — no hermes CLI wrapper since the harness purge
ℹ️ koby: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK
❌ koby: hermes-real NOT FOUND (wrapper broken)
🔍 report-only (koby): wrapper-no-hermes-real:koby — reported, not counted/repaired
✅ koby: wrapper + .env key present
✅ koonimo: wrapper infisical path OK
✅ koonimo: wrapper + .env key present
🔐 Vault Secrets:
✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw
✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg
✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ
✅ All checks passed
exit=0
Live vantage proof (same worktree):
$ ssh root@192.168.68.15 "pct status 111"
Configuration file 'nodes/amdpve/lxc/111.conf' does not exist
$ ssh root@192.168.68.6 "pct status 111; pct list | grep '^ *111'"
status: running
111 running tdunna
$ ssh root@192.168.68.129 "hostname"
tdunna
Leg 2 — infrastructure-monitoring PVE API (item 2)
Before — the contract's probe, aimed at the monitoring host CT 116:
$ curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json
000
CT 116 runs no pveproxy, so it never answers on :8006. The probe target was
wrong, which is what read as PVE-API 000.
After — probing the five real cluster nodes (:8006/api2/json/version),
alive under the any-HTTP-response rule (401 = up, unauthenticated):
$ for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
printf '%s:8006 -> %s\n' "$node" "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
done
192.168.68.9:8006 -> 401
192.168.68.5:8006 -> 401
192.168.68.15:8006 -> 401
192.168.68.6:8006 -> 401
192.168.68.12:8006 -> 401
401 on every node = alive by design. DOWN is 000/timeout only. (The
contract's LiteLLM probe was the same class: /litellm/health answers 301 →
/litellm/health/liveliness, so it is now specified as any-HTTP too.)
Leg 3 — gpu-monitor GPU probes (item 3)
Before — the false alarm came from probing bare port 80 on GPU hosts:
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8/health
000
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110/health
000
Nothing listens on GPU port 80, so the monitor reported
DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000 three times on 2026-09-09.
After — the real endpoints answer:
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health
200
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health
200
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.15:8080/health
200
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified
301 # Location: http://192.168.68.116/gpu/gpu-data — the same payload
$ curl -s -o /dev/null -w '%{http_code}' -L http://192.168.68.116/health/unified
200
301 is healthy under the any-HTTP-response rule. The contract now requires GPU
health on :8080 (or router /health/unified) and forbids bare port 80 on a
GPU host.
Leg 4 — report provenance (item 4)
Every contract report must now lead with the absolute path it executed from.
docs/AUTHORING-GUIDE.md documents the rule and scripts/prose-lint.sh
enforces it:
$ pwd -P
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
$ bash scripts/prose-lint.sh
✅ Report provenance present in all report-format contracts
...
✅ LINT PASSED (16 warning(s))
The health script prints 📍 executed from: script=… cwd=… and includes
execution_path/cwd in --json output.
Full suite
$ pwd -P
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
$ python3 -m pytest -q
23 passed
$ shellcheck scripts/prose-lint.sh
(clean)
Follow-up findings (observed, intentionally NOT changed here)
These are adjacent stale expectations discovered while verifying the four scoped legs. Each touches a CRITICAL/HIGH-sensitivity artifact or an unrelated script, so it is recorded for the captain/verify mate rather than silently repaired.
infrastructure-control.prose.md(CRITICAL) CT 111 node assignment. Lines ~109 and ~615 placetdunna(CT 111, koby) on amdpve. Live verification on 2026-09-10 showspct status 111=runningon storepve (.6) andConfiguration file 'nodes/amdpve/lxc/111.conf' does not existon .15.agent-health-check.pynow carries the live-verifiedstorepvemapping (the script is not the topology source of truth); the CRITICAL contract itself needs an authorized correction.- Strix Halo
:8080firewall claim is stale.prose-ai-review.shground-truth rule #4 andgpu-monitor.prose.mdsay:8080is firewalled to.116only and.24cannot probe it. Live on .15:-A INPUT -s 192.168.68.24/32 -p tcp --dport 8080 -j ACCEPT, and a probe from .24 returns200. The contract keeps routing Strix via the router (safe), but the claim no longer matches iptables. contract-registry.yamlreferencesagent-health-check.prose.md, which does not exist in the repo. The registry entry (withkoby_action: skip_heal) is aspirational/stale.- Pre-existing script defects, untouched:
scripts/pm2-self-heal.shhas a bash syntax error at lines 19–20 (bash -nfails), andshellcheckfails on five untouched scripts (netbird-add-domain.sh,pct-run.sh,pm2-self-heal.sh,prose-ai-review.sh,swap-gpu-dense-model.sh).scripts/prose-lint.sh— the one shell file touched here — is now shellcheck-clean.