Per defect report 1154.msg (agent-health-gateway-leg-flap-20260913):
1. Add _ssh_retry() helper with one retry at longer timeout (25s)
2. Name probe target explicitly: ssh {user}@{host} <command>
3. Print probe-failed when first attempt fails, then retry
4. Only declare gateway-down after retry fails
5. Make Koby report-only explicit in output
Changes:
- check_agents(): all gateway probes now use _ssh_retry()
- All output lines name the probe target (ssh host:port)
- Koby's report-only status is explicit in output
- Never print bare "gateway down" — always name target and failure kind
Verified: koonimo shows "probe-failed" on first attempt (transient SSH),
retries at 25s, succeeds, reports ✅ koonimo: gw=running
The helper was correct but dead code - nothing called it. This commit:
1. Makes the helper runnable standalone: scripts/hermes-reachability-check.sh <host> <pattern> <path>
2. Wires all THREE contracts to it:
- hermes-key-enforcement.prose.md
- hermes-config-template.prose.md
- hermes-agent-baseline.prose.md
3. Each contract now explicitly instructs to run the helper and interpret the three outcomes
4. States that the bug this replaces was deriving reachability from the remote grep's exit code
Files changed (4):
- scripts/hermes-reachability-check.sh (standalone mode added)
- hermes-key-enforcement.prose.md (reachability section added)
- hermes-config-template.prose.md (reachability section added)
- hermes-agent-baseline.prose.md (reachability section added)
Bug: The reachability check used 'ssh ... grep ... || echo unreachable',
which conflated grep's 'no matches found' (exit 1) with SSH failure.
This caused clean hosts to be reported as unreachable.
Fix: Add scripts/hermes-reachability-check.sh with the pattern:
out=$(ssh -o BatchMode=yes root@HOST "grep ... 2>/dev/null; true")
if [ $? -ne 0 ]; then verdict="unreachable"
elif [ -n "$out" ]; then verdict="violation: $out"
else verdict="compliant"
fi
This correctly distinguishes:
- SSH failure (connection/auth/route) -> unreachable
- SSH success + grep found matches -> violation
- SSH success + grep found nothing -> compliant
Evidence: 3 of 4 hosts (Tanko, Mumuni, Koonimo) were reported as
'unreachable' when they were actually compliant. Only Koby (.129)
has a real finding (plaintext key in state snapshot).
Used by: hermes-key-enforcement, hermes-config-template,
hermes-agent-baseline contracts.
1. Timeout fix for pool alias (syslog-auto):
- Single-host aliases (gpu-dense, gpu-vision, strix-moe): 30s timeout
- Pool alias (syslog-auto): 60s timeout, retry once on 000 before failing
- Cold first request to pool alias can take ~13s; 10s was too short
2. Remove DEBUG prints from output:
- Removed 'DEBUG: keylen=...' and 'DEBUG: response=...' lines
- These leaked key inventory to status logs
- Success output now shows only counts (e.g., 'Admin Key List: 10 keys')
Implements all 11 checks from litellm-health contract:
- Liveliness, Containers, Prometheus, Grafana health probes
- Model probes for gpu-dense, gpu-vision, strix-moe, syslog-auto
- Admin API key list (10 keys), GitHub status, Docker Stats metrics
Fixed quoting for SSH commands and response parsing (dict with 'keys' field).
Backend edge uses internal IP 192.168.68.116, not public URL.
Docker Stats fetched from CT 116 host itself (127.0.0.1:9324/metrics).
All 11 checks passing consistently.
CT 111 (tdunna, 192.168.68.129) is Theo's box and is report-only per the captain
(2026-08-17, re-confirmed 2026-09-10). The contract defined AMBER as 'GC scheduled
for next run' and its Execution loop called gc-executor for EVERY threat with no
guest-level exclusion - so a single AMBER reading there would have scheduled apt
clean / journal vacuum / log+tmp deletion against someone else's box. The only
marker was frontmatter report_only_agents, which names an AGENT while the scan unit
is a GUEST.
- Gate the Execution loop on a guest/host-keyed report_only_guests block (guest id,
hostname and IP all match); an excluded guest is alerted and skipped, so no
gc-executor call is constructed for it at any level.
- Carry the ruling in the contract body next to the loop, not only in frontmatter.
- scripts/disk-gc-plan.py: executable planner that reads the contract's
authoritative exclusion block and emits the action plan; tests/ covers it.
- Correct the stale fleet map against pvesh /cluster/resources: CT 105 -> amdpve
(was minipve), CT 111 -> storepve (was amdpve), add guests 118/119/120, and fix
the '15 CTs' counts (20 guests: 17 LXC + 3 QEMU VMs).
- scripts/pct-run.sh: same stale map (105/111 wrong node, 120 missing) - this is
why pct-run 111/105 failed.
- Fold in disk-gc-ct100-probe-gap-20260911: CT 100 verifiably works through
pct-run now that the map is correct; documented that the scanner must probe it
like any other guest, never via a local-only path (the scanner runs inside CT 100).
- Remove adapter process check and restart logic
- Keep A2A probe (port 80, HTTP code check)
- The adapter code at /a0/usr/kagentz-zulip/ no longer exists
- Captain's ruling: Zulip communication with agent zero is not priority
- Load API key from durable file /root/.abiba-workspace/secrets/litellm-key.txt (works in cron)
- Fix http_get to use Bearer token instead of Basic Auth for API endpoint check
- All 6 LiteLLM checks now pass (was 5/6)
- Changed A2A probe from http://127.0.0.1:8001/.well-known/agent.json to
http://127.0.0.1:80/a2a/ inside agent-zero container
- Port 8001 does not exist inside container (nothing listens there)
- Port 80 maps to external port 50080; returns 401 (auth-gated, alive by design)
- Updated A2A_URL in adapter.py restart command to use port 80 instead of 8001
- Verified: probe now returns 401 (auth-gated) instead of 000 (connection refused)
- Before: kagentz A2A reported DOWN on every run (false positive due to stale port)
- After: kagentz A2A reports ✅ A2A alive (auth-gated 401 = healthy)
Correct the dsh-web authentication fix after parent correction:
- Remove the unauthenticated :8081 endpoint (0.0.0.0 bind with no
auth_request = full Authentik bypass for the LAN). The script now removes
/etc/nginx/sites-enabled/dsh.token automatically if it reappears.
- Put the login path inside the Authentik-gated :80 server block as
location = /dsh-web-login; proxy to dsh-web with Host =
tankodhs.sysloggh.net so the 30-day cookie is bound to the public
authority, never to 127.0.0.1:3080.
- Isolate the rotating token in a generated include
/etc/dsh-web/nginx-login.conf; reload nginx only when it changes.
- Replace the disruptive capture (systemctl stop/start dsh-web) with a
non-disruptive read of the running service's journal, scoped to the
current systemd invocation so a restarted process's stale token is never
reused while the new banner is still pending.
- Keep x-dsh-task-board-proxy-token and Host $ak_origin_host intact in '/'.
- Document the corrected design (B4) in zulip-health.prose.md, v3.2.0.
Live-verified 2026-09-11: no auth bypass (302), :8081 refused (000), a
cookie minted before two dsh-web restarts still returns 200, the refreshed
token mints a fresh cookie, and the systemd ExecStartPost/timer refreshes
the token automatically without touching dsh-web.
- Add capture-dsh-token.sh script that captures the dsh-web launch token
- Add login endpoint (/dsh-web-login on :8081) that mints 30-day auth cookie
- Document the authentication flow in zulip-health.prose.md (Platform B4)
- Cookie is authority-bound to 127.0.0.1:3080 with 30-day expiry
- After first login, subsequent requests use the cookie — no token required
Captain ruling 2026-09-10: Mumuni moved off this host onto her own container
(kagentz CT 105 on minipve, 192.168.68.14, dedicated `hermes` user) and is
monitored from her side. This host must not monitor anything Mumuni.
The stale probes fired false alerts repeatedly:
* scripts/zulip-monitor.sh ssh'd to root@192.168.68.24 for the
decommissioned deployment's ~/.hermes/gateway_state.json, read "unknown"
on every run, and posted a 🔴 "Mumuni (Hermes) Zulip state: unknown" DM +
#agent-hub stream alert each cycle.
* scripts/daily-infra-report.py published a matching "mumuni:unknown" row in
every digest.
Changes:
* zulip-monitor.sh: delete the "Platform B: Hermes (Mumuni)" leg and its
notify; keep the Zulip-server, Platform A pi/Abiba (bridge), Platform B
Tanko and Platform C Agent Zero legs. A comment records why the leg is
retired so it is not re-added. Also fixes SC2155 so shellcheck is clean.
* daily-infra-report.py: delete the .24 ~/.hermes/gateway_state.json agent
probe, its hermes --version probe, and the now-dead mumuni render branch.
Abiba (CT 100, its own .24 address) and Tanko legs unchanged; the PVE API
token name is untouched.
* zulip-health.prose.md (v3.1.0): drop the Mumuni-only B4 gateway-process,
B5 heartbeat and B6 response-delivery steps and the stale 192.168.68.24
references; state explicitly that Mumuni is not monitored from this host.
Tanko/Agent-Zero/bridge steps retained.
* agent-health-check.py: correct the v2 changelog roster comment that still
placed mumuni at .24/CT100. No behavior change — the mumuni probe was
already absent from the AGENTS dict; v5 changelog notes the correction.
Tests: tests/test_mumuni_monitor_removal.py pins the removal structurally and
behaviorally — the shipped zulip-monitor.sh is run in a sandbox (only its LOG
constant rewritten) with stub ssh/curl on PATH; the ssh stub records every
target host, so "never reaches .24" and "no Mumuni notify even on the alert
path" are asserted from observed behavior. A mutation check (re-inject the old
leg) fails the suite, so the guarantee is not vacuous.
Second probe-drift correction pass after #65/#66/#68. All four legs were stale
consumer expectations, not live faults.
1. scripts/agent-health-check.py (v4)
- abiba declared pi-only runtime (harness purge): Hermes-era gateway, config
and wrapper legs are skipped instead of failing.
- koby declared report_only (captain ruling 2026-08-17, Rule 17): every koby
leg is detected and reported, never counted as a fleet failure or repaired.
- koby's CT 111 mapping corrected to storepve (.6); the old amdpve mapping
made `pct status 111` fail and read as ct-unreachable.
- wrapper check no longer FAILs .env-based wrappers that legitimately never
invoke infisical (koonimo).
- keys load in main() (load_agent_keys) so the module is importable/testable.
- every run prints absolute execution provenance (script + cwd), in the header
and in --json.
Before: 6 FAILURE(S). After: 0 failures, koby reported read-only.
2. infrastructure-monitoring.prose.md
- PVE API probe repointed from CT 116 (no pveproxy, 000) to the five real
nodes on https://<node>:8006/api2/json/version, all 401 = alive.
- any-HTTP-response liveness rule added (401/3xx alive; 000/timeout = DOWN).
- LiteLLM health documented as 301 -> /litellm/health/liveliness, not bare 200.
3. gpu-monitor.prose.md
- GPU health probes on :8080 (or router /health/unified); bare port 80 on a
GPU host is forbidden (no listener -> false DEGRADED).
- router /health/unified 301 -> /gpu/gpu-data documented as alive.
- port-discipline + liveness rule + direct-fallback execution step.
4. Report provenance (all contracts)
- docs/AUTHORING-GUIDE.md documents the rule; scripts/prose-lint.sh enforces
that any **Report format** contract states an absolute path (pwd -P).
- provenance added to gpu-monitor, infrastructure-monitoring, proxmox-monitor.
Tests: tests/test_probe_drift.py (23 passed); prose-lint.sh clean + shellcheck
clean; CI frontmatter validation passes. Evidence with per-leg before/after and
absolute paths: docs/probe-drift-round2-evidence.md.
The monitor's Abiba leg parsed the :9200/health payload at the top level
(d.get('connected',False)) while the pi Zulip extension serves connection
state NESTED at zulip.connected / zulip.last_error (verified against the
extension's startHealthServer handler and the live endpoint). PI_CONNECTED
was therefore always False and every run took the DISCONNECTED path,
restarting a healthy bot: pm2 restarts=8 with the process created
2026-09-09T09:35:09Z, four ❌ Abiba verdicts today (04:23/05:35/06:55/09:35
UTC) and zero ✅, while the Zulip server answered HTTP 200 and the bot kept
heartbeating. The watchdog was the fault, not the connection.
Changes (Abiba leg only; every other leg byte-identical):
- Read the real nested shape: zulip.connected, zulip.last_error and
zulip.messages_processed. The old retry_count branch is DROPPED — the
payload exposes no retry counter (the extension keeps retryCount internal
and never serialises it), so the branch is fabricated and cannot stay.
- Fail-safe restart decision: a fetch error (HTTP 000), non-2xx response,
empty/unparseable body, or payload missing a boolean zulip.connected is a
clearly-labelled PROBE FAILURE (🟠 alert + ⚠️ log line naming reason and
HTTP code) and NEVER calls pm2 restart. pm2 restart runs only on
affirmative zulip.connected=false (🔴/❌ path unchanged in wording).
connected=true with last_error keeps the degraded 🟡 warn-no-restart path.
Regression test (new tests/zulip-monitor-abiba.sh, 36 checks): extracts the
real Abiba leg from between the # -- abiba-leg-start/-end markers in the
shipped script and executes it verbatim with stubbed curl/notify/pm2 against
fixtures of the real payload shape — asserts connected=true → no restart,
connected=false → restart, and empty/garbage/missing-key/non-boolean/HTTP
000/HTTP 500 bodies → probe-failure alert with zero restarts. The suite
fails loudly on the pre-fix base and on a marker-intact top-level-parse
variant, so this class of bug cannot return silently.
The #agent-hub / zulip-health stream post in notify() encoded content with
python quote(str()) (always empty), so every stream alert posted empty
content and Zulip rejected it silently behind '|| true'. The merged body
also used backslash-escaped ampersands inside double quotes, which curl
transmits literally (type=stream\ -> Zulip 400 'Invalid type').
Stream post now percent-encodes the real message text by piping it through
the encoder (locale-proof: quote_from_bytes on stdin.buffer), uses plain &
separators, and on failure appends one WARN line to the monitor log with
the curl exit status instead of silently swallowing it. Delivery failure
stays non-fatal and unretried. DM path, probes, alive rule, exit codes and
log format unchanged.