The Scot Murray Hermes playbook package did not belong in this repo. It was
added directly to master in c7af7c0 (and extended in 64ccf65) in violation of
this repo's own rule - "No agent pushes directly to main; all changes go
through PRs with automated validation" - and prose-contracts is publicly
readable (private: false), so client engagement material was exposed beyond
the intended audience.
Both commits are unwound by removing the path here, through a PR this time.
The deliverable and its provenance are preserved outside the repo at:
/home/hermes/syslog/drafts/scot-hermes-playbook/
/home/hermes/syslog/projects/murray-capital/deliverables/
No contract files are touched by this change.
- Remove adapter process check and restart logic
- Keep A2A probe (port 80, HTTP code check)
- The adapter code at /a0/usr/kagentz-zulip/ no longer exists
- Captain's ruling: Zulip communication with agent zero is not priority
- Load API key from durable file /root/.abiba-workspace/secrets/litellm-key.txt (works in cron)
- Fix http_get to use Bearer token instead of Basic Auth for API endpoint check
- All 6 LiteLLM checks now pass (was 5/6)
- Changed A2A probe from http://127.0.0.1:8001/.well-known/agent.json to
http://127.0.0.1:80/a2a/ inside agent-zero container
- Port 8001 does not exist inside container (nothing listens there)
- Port 80 maps to external port 50080; returns 401 (auth-gated, alive by design)
- Updated A2A_URL in adapter.py restart command to use port 80 instead of 8001
- Verified: probe now returns 401 (auth-gated) instead of 000 (connection refused)
- Before: kagentz A2A reported DOWN on every run (false positive due to stale port)
- After: kagentz A2A reports ✅ A2A alive (auth-gated 401 = healthy)
Provenance: kanban t_02148213 (research) -> t_4265c369 (author) -> t_fefdf30b (review).
Review verdict: APPROVED-WITH-FIXES, 7/7 checks PASS, 17/17 YouTube links oEmbed-verified,
26+ CLI commands re-run against v0.21.1, no Murray/JDS client data present.
Artifact: deliverables/scot-hermes-playbook/HERMES-PLAYBOOK-FOR-SCOT.md
377 lines, sha256 b0966a649fec96da4975ec00e627fbaba3a92a62c4a92b33bc06589c47d25f7f
Includes the research dossier and its raw verification evidence under research/.
STATUS: written and reviewed, NOT delivered to the client. Delivery is gated on
Kwame's approval and tracked as a separate blocked kanban card.
Correct the dsh-web authentication fix after parent correction:
- Remove the unauthenticated :8081 endpoint (0.0.0.0 bind with no
auth_request = full Authentik bypass for the LAN). The script now removes
/etc/nginx/sites-enabled/dsh.token automatically if it reappears.
- Put the login path inside the Authentik-gated :80 server block as
location = /dsh-web-login; proxy to dsh-web with Host =
tankodhs.sysloggh.net so the 30-day cookie is bound to the public
authority, never to 127.0.0.1:3080.
- Isolate the rotating token in a generated include
/etc/dsh-web/nginx-login.conf; reload nginx only when it changes.
- Replace the disruptive capture (systemctl stop/start dsh-web) with a
non-disruptive read of the running service's journal, scoped to the
current systemd invocation so a restarted process's stale token is never
reused while the new banner is still pending.
- Keep x-dsh-task-board-proxy-token and Host $ak_origin_host intact in '/'.
- Document the corrected design (B4) in zulip-health.prose.md, v3.2.0.
Live-verified 2026-09-11: no auth bypass (302), :8081 refused (000), a
cookie minted before two dsh-web restarts still returns 200, the refreshed
token mints a fresh cookie, and the systemd ExecStartPost/timer refreshes
the token automatically without touching dsh-web.
- Add capture-dsh-token.sh script that captures the dsh-web launch token
- Add login endpoint (/dsh-web-login on :8081) that mints 30-day auth cookie
- Document the authentication flow in zulip-health.prose.md (Platform B4)
- Cookie is authority-bound to 127.0.0.1:3080 with 30-day expiry
- After first login, subsequent requests use the cookie — no token required
- New Level 1 fix 4: archive-suggested stale nodes are archived outright via one
updateNode call (description -> '[ARCHIVED] ' + metadata state=archived).
The [REVIEW: archive] tag is retired.
- Fix 3 now tags refresh-suggested (living) nodes only; those are still escalated.
- Corrected the SSH claim: updateNode DOES accept state transitions and bumps
updated_at (verified 2026-09-11 archiving 7 nodes: #61#373#388#465#475#526#1476).
- Level 2 escalation 1 rescoped to refresh-suggested nodes.
- Checks section: any remaining [REVIEW: archive] means fix 4 was skipped.
Digest-pinned images (image: repo@sha256:...) are invisible to 'docker compose
pull' — the pin re-pulls the same digest forever, so new releases never appear.
Verified 2026-09-10: audiobookshelf ran 2.34.0 for 7 weeks despite weekly pulls;
un-pinned to :latest, now 2.36.0 (HTTP 200). Dockhand stack (also pinned) was
removed 2026-09-10 as unused (user decision). Weekly task qSOOVzsU now sweeps
for digest pins every run and flags them for user-approved un-pinning.
Captain ruling 2026-09-10: Mumuni moved off this host onto her own container
(kagentz CT 105 on minipve, 192.168.68.14, dedicated `hermes` user) and is
monitored from her side. This host must not monitor anything Mumuni.
The stale probes fired false alerts repeatedly:
* scripts/zulip-monitor.sh ssh'd to root@192.168.68.24 for the
decommissioned deployment's ~/.hermes/gateway_state.json, read "unknown"
on every run, and posted a 🔴 "Mumuni (Hermes) Zulip state: unknown" DM +
#agent-hub stream alert each cycle.
* scripts/daily-infra-report.py published a matching "mumuni:unknown" row in
every digest.
Changes:
* zulip-monitor.sh: delete the "Platform B: Hermes (Mumuni)" leg and its
notify; keep the Zulip-server, Platform A pi/Abiba (bridge), Platform B
Tanko and Platform C Agent Zero legs. A comment records why the leg is
retired so it is not re-added. Also fixes SC2155 so shellcheck is clean.
* daily-infra-report.py: delete the .24 ~/.hermes/gateway_state.json agent
probe, its hermes --version probe, and the now-dead mumuni render branch.
Abiba (CT 100, its own .24 address) and Tanko legs unchanged; the PVE API
token name is untouched.
* zulip-health.prose.md (v3.1.0): drop the Mumuni-only B4 gateway-process,
B5 heartbeat and B6 response-delivery steps and the stale 192.168.68.24
references; state explicitly that Mumuni is not monitored from this host.
Tanko/Agent-Zero/bridge steps retained.
* agent-health-check.py: correct the v2 changelog roster comment that still
placed mumuni at .24/CT100. No behavior change — the mumuni probe was
already absent from the AGENTS dict; v5 changelog notes the correction.
Tests: tests/test_mumuni_monitor_removal.py pins the removal structurally and
behaviorally — the shipped zulip-monitor.sh is run in a sandbox (only its LOG
constant rewritten) with stub ssh/curl on PATH; the ssh stub records every
target host, so "never reaches .24" and "no Mumuni notify even on the alert
path" are asserted from observed behavior. A mutation check (re-inject the old
leg) fails the suite, so the guarantee is not vacuous.