Files
prose-contracts/docs/probe-drift-round2-evidence.md
T
root 2238777a2f
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Failing after 13m19s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
fix(alignment): resolve /root/ hardcoding and stale tanko-DSH references
F1: agent-health-check.py now resolves the home directory from the agent's
user field via a shared helper (get_user_home) instead of hardcoding /root/.
This fixes the false-positive wrapper-missing:tanko report — tanko has a
working wrapper at /home/jerome/.local/bin/hermes, but the check was looking
in /root/.local/bin/.

F2: Updated stale references that described tanko as DSH-only:
- hermes-zulip-restore.prose.md: tanko excluded — hybrid (DSH + Hermes)
- hermes-zulip-plugin.prose.md: tanko excluded — hybrid (DSH + Hermes)
- infrastructure-control.prose.md: tanko is hybrid (DSH + Hermes) agent
- docs/probe-drift-round2-evidence.md: marked as historical record with
  dated note explaining that the DSH-only observations reflected the
  /root/ hardcoding bug, not the underlying truth

Refs: fix/agent-health-root-hardcoding-20260928
2026-09-28 20:48:49 +00:00

12 KiB
Raw Blame History

Probe-drift round 2 — per-leg before/after evidence

Historical record — 2026-09-28: The lines below that describe tanko as "DSH (DeepSeek Harness)" only reflect what the check reported when it was running. Tanko's runtime was later found to be hybrid (DSH + Hermes) — the check had a /root/ hardcoding bug that made it probe the wrong home directory and report wrapper-missing:tanko for an agent with a working wrapper. This document records the observed output, not the underlying truth; see fix/agent-health-root-hardcoding-20260928 for the correction.

Date: 2026-09-10 Worktree (absolute execution path): /root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts Branch: fm/probe-drift-round2-20260909

Every command below was run from the absolute path above; output is pasted verbatim. This is the evidence trail for the four scoped corrections; it is not a contract (never prose run it).


Leg 1 — agent-health-check (item 1)

Before — from /root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts, python3 scripts/agent-health-check.py --no-deploy (v2, base of this branch):

🏥 Agent Health Check v2 — 2026-09-10 01:16 UTC

🔑 LiteLLM Keys:
  ✅ tanko: key valid → syslog-auto
  ✅ abiba: key valid → syslog-auto
  ✅ koby: key valid → syslog-auto
  ✅ koonimo: key valid → syslog-auto

🎮 GPU Port Health:
  ✅ gpu-rtx3090 (.8): healthy (pid=472206)
  ✅ gpu-rtx5070 (.110): healthy (pid=207601)
  ✅ gpu-strixhalo (.15): healthy (pid=4098872)

🤖 Agent Gateways:
  ✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK)
  ⚠️ abiba: gw=no-state-file zulip=? streaming=no errors_10m=0 pid=?
  ✅ koby: gw=running zulip=connected streaming=no errors_10m=0 pid=360900
  ✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125

🖥️  CT Liveness:
  ✅ tanko (CT 112 on amdpve): running
  ✅ abiba (CT 100 on minipve): running
  ❌ koby (CT 111 on amdpve): PVE UNREACHABLE
  ✅ koonimo (CT 113 on amdpve): running

📝 Config Integrity:
  ⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27
  ✅ abiba: config.yaml valid YAML
  ✅ koby: config.yaml valid YAML
  ✅ koonimo: config.yaml valid YAML

🔌 Wrapper/CLI Integrity:
  ⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27
  ⚠️  abiba: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
  ❌ abiba: hermes-real NOT FOUND (wrapper broken)
  ⚠️  abiba: .env may be missing LITELLM_API_KEY entry
  ⚠️  koby: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
  ❌ koby: hermes-real NOT FOUND (wrapper broken)
  ✅ koby: wrapper + .env key present
  ⚠️  koonimo: wrapper infisical path may be wrong (infisical at /usr/bin/infisical)
  ✅ koonimo: wrapper + .env key present

🔐 Vault Secrets:
  ✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw
  ✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg
  ✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ

❌ 6 FAILURE(S): ct-unreachable:koby:192.168.68.15 | wrapper-infisical-path:abiba | wrapper-no-hermes-real:abiba | wrapper-infisical-path:koby | wrapper-no-hermes-real:koby | wrapper-infisical-path:koonimo

Root causes (all stale expectations; no live fault):

Failure Why it was stale
ct-unreachable:koby:192.168.68.15 CT 111 (tdunna/koby) runs on storepve (.6), not amdpve (.15).
wrapper-*:abiba Abiba is pi-only since the harness purge. /root/.local/bin/hermes is a dangling symlink; no hermes-real, no ~/.hermes/.env.
wrapper-*:koby Koby is report-only (captain ruling 2026-08-17, Rule 17): detect and report, never repair — its legs must not count as fleet failures. Koby's wrapper is also the genuine no-infisical case: it sources ~/.hermes/.env rather than /usr/bin/infisical, which the check now accepts.
wrapper-infisical-path:koonimo Koonimo's wrapper does reference /usr/bin/infisical — but past the old check's head -20 window, so the check looked for the path in the wrong slice and false-failed. The fix that mattered was reading the full wrapper body (and then verifying any absolute infisical path it finds actually exists).

After — same absolute path, python3 scripts/agent-health-check.py --no-deploy (v4):

$ pwd -P
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
$ python3 scripts/agent-health-check.py --no-deploy
🏥 Agent Health Check v4 — 2026-09-10 01:22 UTC
📍 executed from: script=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts/scripts/agent-health-check.py cwd=/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts

🔑 LiteLLM Keys:
  ✅ tanko: key valid → syslog-auto
  ✅ abiba: key valid → syslog-auto
  ✅ koby: key valid → syslog-auto
  ✅ koonimo: key valid → syslog-auto

🎮 GPU Port Health:
  ✅ gpu-rtx3090 (.8): healthy (pid=472206)
  ✅ gpu-rtx5070 (.110): healthy (pid=207601)
  ✅ gpu-strixhalo (.15): healthy (pid=4098872)

🤖 Agent Gateways:
  ✅ tanko: DSH (DeepSeek Harness) — no Hermes gateway since 2026-08-27 (CT 112, SSH OK)
  ✅ abiba: pi-only runtime — no Hermes gateway since the harness purge (CT 100, SSH OK)
  🔍 koby: REPORT-ONLY mode (diagnostic only, no repairs on .129)
  ✅ koby: gateway running (pid=360900, report-only mode)
  ✅ koonimo: gw=running zulip=connected streaming=no errors_10m=0 pid=155125

🖥️  CT Liveness:
  ✅ tanko (CT 112 on amdpve): running
  ✅ abiba (CT 100 on minipve): running
  ✅ koby (CT 111 on storepve): running
  ✅ koonimo (CT 113 on amdpve): running

📝 Config Integrity:
  ⏭️ tanko: DSH — no Hermes config.yaml since 2026-08-27
  ⏭️ abiba: pi-only runtime — no Hermes config.yaml since the harness purge
  ✅ koby: config.yaml valid YAML
  ✅ koonimo: config.yaml valid YAML

🔌 Wrapper/CLI Integrity:
  ⏭️ tanko: DSH — no hermes CLI wrapper since 2026-08-27
  ⏭️ abiba: pi-only runtime — no hermes CLI wrapper since the harness purge
  ℹ️  koby: wrapper resolves creds without infisical (e.g. ~/.hermes/.env) — OK
  ❌ koby: hermes-real NOT FOUND (wrapper broken)
  🔍 report-only (koby): wrapper-no-hermes-real:koby — reported, not counted/repaired
  ✅ koby: wrapper + .env key present
  ✅ koonimo: wrapper infisical path OK
  ✅ koonimo: wrapper + .env key present

🔐 Vault Secrets:
  ✅ tanko: vault TANKO_LITELLM_API_KEY=sk-...x6uw
  ✅ koby: vault KOBY_LITELLM_API_KEY=sk-...jxlg
  ✅ koonimo: vault KOONIMO_LITELLM_API_KEY=sk-...Y0KQ

✅ All checks passed
exit=0

Live vantage proof (same worktree):

$ ssh root@192.168.68.15 "pct status 111"
Configuration file 'nodes/amdpve/lxc/111.conf' does not exist
$ ssh root@192.168.68.6 "pct status 111; pct list | grep '^ *111'"
status: running
111        running                 tdunna
$ ssh root@192.168.68.129 "hostname"
tdunna

Leg 2 — infrastructure-monitoring PVE API (item 2)

Before — the contract's probe, aimed at the monitoring host CT 116:

$ curl -s -o /dev/null -w '%{http_code}' https://192.168.68.116:8006/api2/json
000

CT 116 runs no pveproxy, so it never answers on :8006. The probe target was wrong, which is what read as PVE-API 000.

After — probing the five real cluster nodes (:8006/api2/json/version), alive under the any-HTTP-response rule (401 = up, unauthenticated):

$ for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
    printf '%s:8006 -> %s\n' "$node" "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
  done
192.168.68.9:8006 -> 401
192.168.68.5:8006 -> 401
192.168.68.15:8006 -> 401
192.168.68.6:8006 -> 401
192.168.68.12:8006 -> 401

401 on every node = alive by design. DOWN is 000/timeout only. (The contract's LiteLLM probe was the same class: /litellm/health answers 301 → /litellm/health/liveliness, so it is now specified as any-HTTP too.)


Leg 3 — gpu-monitor GPU probes (item 3)

Before — the false alarm came from probing bare port 80 on GPU hosts:

$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8/health
000
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110/health
000

Nothing listens on GPU port 80, so the monitor reported DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000 three times on 2026-09-09.

After — the real endpoints answer:

$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health
200
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health
200
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.15:8080/health
200
$ curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified
301        # Location: http://192.168.68.116/gpu/gpu-data — the same payload
$ curl -s -o /dev/null -w '%{http_code}' -L http://192.168.68.116/health/unified
200

301 is healthy under the any-HTTP-response rule. The contract now requires GPU health on :8080 (or router /health/unified) and forbids bare port 80 on a GPU host.


Leg 4 — report provenance (item 4)

Every contract report must now lead with the absolute path it executed from. docs/AUTHORING-GUIDE.md documents the rule and scripts/prose-lint.sh enforces it:

$ pwd -P
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
$ bash scripts/prose-lint.sh
  ✅ Report provenance present in all report-format contracts
...
✅ LINT PASSED (12 warning(s))

The health script prints 📍 executed from: script=… cwd=… and includes execution_path/cwd in --json output.


Full suite

$ pwd -P
/root/.treehouse/prose-contracts-9ce5f3/3/prose-contracts
$ python3 -m pytest -q
24 passed
$ shellcheck scripts/prose-lint.sh
(clean)

Follow-up findings (observed, intentionally NOT changed here)

These are adjacent stale expectations discovered while verifying the four scoped legs. Each touches a CRITICAL/HIGH-sensitivity artifact or an unrelated script, so it is recorded for the captain/verify mate rather than silently repaired.

  1. infrastructure-control.prose.md (CRITICAL) CT 111 node assignment. Lines ~109 and ~615 place tdunna (CT 111, koby) on amdpve. Live verification on 2026-09-10 shows pct status 111 = running on storepve (.6) and Configuration file 'nodes/amdpve/lxc/111.conf' does not exist on .15. agent-health-check.py now carries the live-verified storepve mapping (the script is not the topology source of truth); the CRITICAL contract itself needs an authorized correction. ✅ Resolved 2026-09-12: the topology was corrected in its owner, infrastructure-control.prose.md (CT 111 → storepve, CT 105 → amdpve), and scripts/pct-run.sh now matches. This snapshot is left as observed; treat those owner documents as authoritative.
  2. Strix Halo :8080 firewall claim is stale. prose-ai-review.sh ground-truth rule #4 and gpu-monitor.prose.md say :8080 is firewalled to .116 only and .24 cannot probe it. Live on .15: -A INPUT -s 192.168.68.24/32 -p tcp --dport 8080 -j ACCEPT, and a probe from .24 returns 200. The contract keeps routing Strix via the router (safe), but the claim no longer matches iptables.
  3. contract-registry.yaml references agent-health-check.prose.md, which does not exist in the repo. The registry entry (with koby_action: skip_heal) is aspirational/stale.
  4. Pre-existing script defects, untouched: scripts/pm2-self-heal.sh has a bash syntax error at lines 19–20 (bash -n fails), and shellcheck fails on five untouched scripts (netbird-add-domain.sh, pct-run.sh, pm2-self-heal.sh, prose-ai-review.sh, swap-gpu-dense-model.sh). scripts/prose-lint.sh — the one shell file touched here — is now shellcheck-clean.