F1: agent-health-check.py now resolves the home directory from the agent's user field via a shared helper (get_user_home) instead of hardcoding /root/. This fixes the false-positive wrapper-missing:tanko report — tanko has a working wrapper at /home/jerome/.local/bin/hermes, but the check was looking in /root/.local/bin/. F2: Updated stale references that described tanko as DSH-only: - hermes-zulip-restore.prose.md: tanko excluded — hybrid (DSH + Hermes) - hermes-zulip-plugin.prose.md: tanko excluded — hybrid (DSH + Hermes) - infrastructure-control.prose.md: tanko is hybrid (DSH + Hermes) agent - docs/probe-drift-round2-evidence.md: marked as historical record with dated note explaining that the DSH-only observations reflected the /root/ hardcoding bug, not the underlying truth Refs: fix/agent-health-root-hardcoding-20260928
12 KiB
Probe-drift round 2 — per-leg before/after evidence
Historical record — 2026-09-28: The lines below that describe tanko as "DSH (DeepSeek Harness)" only reflect what the check reported when it was running. Tanko's runtime was later found to be hybrid (DSH + Hermes) — the check had a
/root/hardcoding bug that made it probe the wrong home directory and reportwrapper-missing:tankofor an agent with a working wrapper. This document records the observed output, not the underlying truth; seefix/agent-health-root-hardcoding-20260928for the correction.
Date: 2026-09-10
Worktree (absolute execution path): /root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
Branch: fm/probe-drift-round2-20260909
Every command below was run from the absolute path above; output is pasted
verbatim. This is the evidence trail for the four scoped corrections; it is not
a contract (never prose run it).
Leg 1 — agent-health-check (item 1)
Before — from /root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts,
python3 scripts/agent-health-check.py --no-deploy (v2, base of this branch):
🏥 Agent Health Check v2 — 2026-09-10 01:16 UTC
🔑 LiteLLM Keys:
✅ tanko: key valid → syslog-auto
✅ abiba: key valid → syslog-auto
✅ koby: key valid → syslog-auto
✅ koonimo: key valid → syslog-auto
🎮 GPU Port Health:
✅ gpu-rtx3090 (.8): healthy (pid=472206)
✅ gpu-rtx5070 (.110): healthy (pid=207601)
✅ gpu-strixhalo (.15): healthy (pid=4098872)
🤖 Agent Gateways:
✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK)
⚠️ abiba: gw=no-state-file zulip=? streaming=no errors_10m=0 pid=?
✅ koby: gw=running zulip=connected streaming=no errors_10m=0 pid=360900
✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125
🖥️ CT Liveness:
✅ tanko (CT 112 on amdpve): running
✅ abiba (CT 100 on minipve): running
❌ koby (CT 111 on amdpve): PVE UNREACHABLE
✅ koonimo (CT 113 on amdpve): running
📝 Config Integrity:
⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27
✅ abiba: config.yaml valid YAML
✅ koby: config.yaml valid YAML
✅ koonimo: config.yaml valid YAML
🔌 Wrapper/CLI Integrity:
⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27
⚠️ abiba: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
❌ abiba: hermes-real NOT FOUND (wrapper broken)
⚠️ abiba: .env may be missing LITELLM_API_KEY entry
⚠️ koby: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
❌ koby: hermes-real NOT FOUND (wrapper broken)
✅ koby: wrapper + .env key present
⚠️ koonimo: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
✅ koonimo: wrapper + .env key present
🔐 Vault Secrets:
✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw
✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg
✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ
❌ 6 FAILURE(S): ct-unreachable:koby:192.168.68.15 | wrapper-infisical-path:abiba | wrapper-no-hermes-real:abiba | wrapper-infisical-path:koby | wrapper-no-hermes-real:koby | wrapper-infisical-path:koonimo
Root causes (all stale expectations; no live fault):
| Failure | Why it was stale |
|---|---|
ct-unreachable:koby:192.168.68.15 |
CT 111 (tdunna/koby) runs on storepve (.6), not amdpve (.15). |
wrapper-*:abiba |
Abiba is pi-only since the harness purge. /root/.local/bin/hermes is a dangling symlink; no hermes-real, no ~/.hermes/.env. |
wrapper-*:koby |
Koby is report-only (captain ruling 2026-08-17, Rule 17): detect and report, never repair — its legs must not count as fleet failures. Koby's wrapper is also the genuine no-infisical case: it sources ~/.hermes/.env rather than /usr/bin/infisical, which the check now accepts. |
wrapper-infisical-path:koonimo |
Koonimo's wrapper does reference /usr/bin/infisical — but past the old check's head -20 window, so the check looked for the path in the wrong slice and false-failed. The fix that mattered was reading the full wrapper body (and then verifying any absolute infisical path it finds actually exists). |
After — same absolute path, python3 scripts/agent-health-check.py --no-deploy (v4):
$ pwd -P
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
$ python3 scripts/agent-health-check.py --no-deploy
🏥 Agent Health Check v4 — 2026-09-10 01:22 UTC
📍 executed from: script=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts/scripts/agent-health-check.py cwd=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
🔑 LiteLLM Keys:
✅ tanko: key valid → syslog-auto
✅ abiba: key valid → syslog-auto
✅ koby: key valid → syslog-auto
✅ koonimo: key valid → syslog-auto
🎮 GPU Port Health:
✅ gpu-rtx3090 (.8): healthy (pid=472206)
✅ gpu-rtx5070 (.110): healthy (pid=207601)
✅ gpu-strixhalo (.15): healthy (pid=4098872)
🤖 Agent Gateways:
✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK)
✅ abiba: pi-only runtime — no Hermes gateway since the harness purge (CT 100, SSH OK)
🔍 koby: REPORT-ONLY mode (diagnostic only, no repairs on .129)
✅ koby: gateway running (pid=360900, report-only mode)
✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125
🖥️ CT Liveness:
✅ tanko (CT 112 on amdpve): running
✅ abiba (CT 100 on minipve): running
✅ koby (CT 111 on storepve): running
✅ koonimo (CT 113 on amdpve): running
📝 Config Integrity:
⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27
⏭️ abiba: pi-only runtime — no Hermes config.yaml since the harness purge
✅ koby: config.yaml valid YAML
✅ koonimo: config.yaml valid YAML
🔌 Wrapper/CLI Integrity:
⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27
⏭️ abiba: pi-only runtime — no hermes CLI wrapper since the harness purge
ℹ️ koby: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK
❌ koby: hermes-real NOT FOUND (wrapper broken)
🔍 report-only (koby): wrapper-no-hermes-real:koby — reported, not counted/repaired
✅ koby: wrapper + .env key present
✅ koonimo: wrapper infisical path OK
✅ koonimo: wrapper + .env key present
🔐 Vault Secrets:
✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw
✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg
✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ
✅ All checks passed
exit=0
Live vantage proof (same worktree):
$ ssh root@192.168.68.15 "pct status 111"
Configuration file 'nodes/amdpve/lxc/111.conf' does not exist
$ ssh root@192.168.68.6 "pct status 111; pct list | grep '^ *111'"
status: running
111 running tdunna
$ ssh root@192.168.68.129 "hostname"
tdunna
Leg 2 — infrastructure-monitoring PVE API (item 2)
Before — the contract's probe, aimed at the monitoring host CT 116:
$ curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json
000
CT 116 runs no pveproxy, so it never answers on :8006. The probe target was
wrong, which is what read as PVE-API 000.
After — probing the five real cluster nodes (:8006/api2/json/version),
alive under the any-HTTP-response rule (401 = up, unauthenticated):
$ for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
printf '%s:8006 -> %s\n' "$node" "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
done
192.168.68.9:8006 -> 401
192.168.68.5:8006 -> 401
192.168.68.15:8006 -> 401
192.168.68.6:8006 -> 401
192.168.68.12:8006 -> 401
401 on every node = alive by design. DOWN is 000/timeout only. (The
contract's LiteLLM probe was the same class: /litellm/health answers 301 →
/litellm/health/liveliness, so it is now specified as any-HTTP too.)
Leg 3 — gpu-monitor GPU probes (item 3)
Before — the false alarm came from probing bare port 80 on GPU hosts:
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8/health
000
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110/health
000
Nothing listens on GPU port 80, so the monitor reported
DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000 three times on 2026-09-09.
After — the real endpoints answer:
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health
200
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health
200
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.15:8080/health
200
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified
301 # Location: http://192.168.68.116/gpu/gpu-data — the same payload
$ curl -s -o /dev/null -w '%{http_code}' -L http://192.168.68.116/health/unified
200
301 is healthy under the any-HTTP-response rule. The contract now requires GPU
health on :8080 (or router /health/unified) and forbids bare port 80 on a
GPU host.
Leg 4 — report provenance (item 4)
Every contract report must now lead with the absolute path it executed from.
docs/AUTHORING-GUIDE.md documents the rule and scripts/prose-lint.sh
enforces it:
$ pwd -P
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
$ bash scripts/prose-lint.sh
✅ Report provenance present in all report-format contracts
...
✅ LINT PASSED (12 warning(s))
The health script prints 📍 executed from: script=… cwd=… and includes
execution_path/cwd in --json output.
Full suite
$ pwd -P
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
$ python3 -m pytest -q
24 passed
$ shellcheck scripts/prose-lint.sh
(clean)
Follow-up findings (observed, intentionally NOT changed here)
These are adjacent stale expectations discovered while verifying the four scoped legs. Each touches a CRITICAL/HIGH-sensitivity artifact or an unrelated script, so it is recorded for the captain/verify mate rather than silently repaired.
infrastructure-control.prose.md(CRITICAL) CT 111 node assignment. Lines ~109 and ~615 placetdunna(CT 111, koby) on amdpve. Live verification on 2026-09-10 showspct status 111=runningon storepve (.6) andConfiguration file 'nodes/amdpve/lxc/111.conf' does not existon .15.agent-health-check.pynow carries the live-verifiedstorepvemapping (the script is not the topology source of truth); the CRITICAL contract itself needs an authorized correction. ✅ Resolved 2026-09-12: the topology was corrected in its owner,infrastructure-control.prose.md(CT 111 → storepve, CT 105 → amdpve), andscripts/pct-run.shnow matches. This snapshot is left as observed; treat those owner documents as authoritative.- Strix Halo
:8080firewall claim is stale.prose-ai-review.shground-truth rule #4 andgpu-monitor.prose.mdsay:8080is firewalled to.116only and.24cannot probe it. Live on .15:-A INPUT -s 192.168.68.24/32 -p tcp --dport 8080 -j ACCEPT, and a probe from .24 returns200. The contract keeps routing Strix via the router (safe), but the claim no longer matches iptables. contract-registry.yamlreferencesagent-health-check.prose.md, which does not exist in the repo. The registry entry (withkoby_action: skip_heal) is aspirational/stale.- Pre-existing script defects, untouched:
scripts/pm2-self-heal.shhas a bash syntax error at lines 19–20 (bash -nfails), andshellcheckfails on five untouched scripts (netbird-add-domain.sh,pct-run.sh,pm2-self-heal.sh,prose-ai-review.sh,swap-gpu-dense-model.sh).scripts/prose-lint.sh— the one shell file touched here — is now shellcheck-clean.