The audit assumed fallback_providers was always a dict (single provider).
Two live agents (koby, koonimo) carry it as a LIST of dicts (one entry per
fallback), so the script crashed with:
File "audit-hermes-config.py", line 211, in audit
fb.get("provider") == "deepseek",
AttributeError: 'list' object has no attribute 'get'
Both are REAL agent configs, so this is not a malformed-input case — the
script simply could not audit two of the four agents it exists to audit.
Fix:
- Normalize fallback_providers to a list of entries (dict → [dict], list → list)
- Apply the existing checks to each entry
- A malformed entry (not a mapping) produces a reported VIOLATION naming the
offending entry, NOT an uncaught exception
Adds regression test using the real failing shape (list of dicts) and proves
it bites against the pre-fix revision.
Real audit results after fix:
- mumuni: FAIL — 7 violations
- tanko: FAIL — 21 violations
- koby: FAIL — 16 violations (previously crashed)
- koonimo: FAIL — 10 violations (previously crashed)
No agent configs were changed. No existing rules were relaxed.
Bounded correction round on PR #134 after a PASS-WITH-FINDINGS review whose
Finding 4 is High. The guard's purpose and its fail-closed fix stand; the
problem was that with 'enforce' as the default it gates EVERY scheduled
contract, and three legitimate states produce a refusal - a clone legitimately
ahead of origin/master mid-review, a detached HEAD, and an offline or failed
fetch - so any of them would turn the fleet's monitoring into withheld
verdicts. That risk outweighs the staleness the guard catches.
1. DEFAULT IS NOW 'warn'. 'enforce' remains available and documented. The
criteria for flipping the default later are written into the doc as a
decision with evidence - a sustained window (30 days / 200+ runs) with zero
mismatch:* and zero cannot-verify:* refusals, no fetch blips, and a pinned
clone demonstrably kept current - explicitly as its own change, not a silent
flip.
2. 'COULD NOT CHECK' IS NOW DISTINGUISHABLE FROM 'THIS COPY IS WRONG'. Every
non-zero exit prints a machine-readable REASON=<class> line:
cannot-verify:fetch-failed | cannot-verify:ref-unresolvable (exit 2)
mismatch:path-absent | mismatch:content
mismatch:detached-head | mismatch:clone-ahead (exit 1)
detached-head and clone-ahead are named separately because they are
legitimate states, far less alarming than a hand-edited file. clone-ahead
requires HEAD to be STRICTLY ahead; an uncommitted edit on a commit that IS
the ref is a plain content mismatch (my own first cut got this wrong and the
new test 7d caught it).
3. THE DEFAULT FETCH IS BOUNDED: --fetch-timeout, default 20s, 0 = unbounded,
and a missing 'timeout' binary is itself a cannot-verify rather than an
unbounded fetch inside a scheduled contract.
4. TEST COVERAGE ADDED for every new class: fetch failure, fetch timeout
(asserted to return promptly under a 1s bound), unresolvable ref, detached
HEAD, clone-ahead, genuine content mismatch, and the contract-run.sh default.
The pre-fix draft fixture comparisons are kept: 31 passed, 0 failed.
5. MERGE-TIME SEQUENCE documented: fast-forward /opt/contract-runner, confirm
clean, prove a contract runs and reports. Baseline recorded as of today -
firstmate has already fast-forwarded it to 9faffe4 - with the note that an
untracked file blocks a fast-forward even when byte-identical.
Live behaviour re-verified on the real runner path:
default: REASON=mismatch:clone-ahead -> 'continuing because ...=warn' -> VERDICT: PASS, exit 0
enforce: REASON=mismatch:clone-ahead -> 'VERDICT WITHHELD: mismatch:clone-ahead', exit 2
MANDATORY CHECKS (master went red once from a credential-SHAPED string, so
these are now run on every shippable branch):
bash scripts/prose-lint.sh -> LINT PASSED (18 warning(s))
secret scan -> secret scan clean (tree; 34 allowlisted,
24 inert value(s) ignored); No committed credentials
shellcheck revision-preflight.sh -> clean
shellcheck test_revision_preflight.sh -> clean
shellcheck contract-run.sh -> SC2034 x1, SC2086 x2 - byte-identical on
master, i.e. pre-existing, none introduced
tests/test_probe_drift.py::test_prose_lint_accepts_report_format_with_provenance
fails both before and after this branch (it runs prose-lint from a temp CWD and
cannot find its sibling secret-scan.sh). Pre-existing, unrelated, not fixed here.
Master's lint is RED right now, and it is my doing.
at 483a66b (before #133): prose-lint -> LINT PASSED
at 9d64b0b (after #133): prose-lint -> LINT FAILED, 4 credential-shaped strings
The regression came from #133's new pve_auth() work. The secret scanner flags
the literal header shape PVEAPIToken=<value> (rule proxmox-token), and three
occurrences landed in the tree:
scripts/daily-infra-report.py the f-string building the auth header
tests/test_daily_infra_report.py a docstring quoting the old placeholder
tests/test_daily_infra_report.py a synthetic token in an assertion
Fixed without allowlisting anything, because none of these is a credential:
* the header prefix becomes PVE_AUTH_HEADER = "PVEAPIToken=", a constant ending
at '=' so the scanner's pattern (which needs a character after '=') cannot
match, and the f-string no longer contains the literal;
* the test docstring no longer reproduces the old placeholder verbatim;
* the test builds its expected value from the constant plus a local sample
variable instead of embedding a credential-shaped literal.
Behaviour is unchanged and re-verified: with PVE_TOKEN unset the script still
exits 1 with the probe failure, and with the vault token it still reports
pve_probe_status ok / node_count 5 / nodes_online 5.
before: LINT FAILED — 4 credential-shaped strings
after: LINT PASSED (18 warnings)
A contract verdict is only meaningful if it came from the merged copy. The
fleet has been bitten three times on 2026-09-25 (a clone parked on a merged
feature branch while executing from another clone; a script copied into the
runner clone by hand; a stale local origin/master making an ancestry check
report unlanded work). The control for this existed as an untracked draft and
protected nobody, because it was entirely fail-open.
Defect in the draft, preserved verbatim as tests/fixtures/revision-preflight.prefix.sh:
git -C "$CLONE" show "origin/master:$(basename "$SCRIPT")"
basename drops the scripts/ prefix, so for any script under scripts/ it queried
the repo root, failed, took the "warn but don't block" branch and exited 0 -
passing a script that exists in no revision at all. Reproduced:
pre-fix + scripts/demo.sh under scripts/ -> 'could not resolve', EXIT=0
pre-fix + a script in no revision -> EXIT=0
Fixed guard (scripts/revision-preflight.sh):
* resolves the repo-relative path inside the clone, so scripts/ paths resolve;
* FAILS CLOSED - a path absent from the ref, an unresolvable ref, or a failed
fetch is a failure, never a warning;
* fetches the remote by default, because a stale local ref would otherwise
pass a stale script as current; --no-fetch states the assumption instead of
hiding it.
Wiring (scripts/contract-run.sh): before executing, the wrapper runs the guard
against the clone it lives in. Default CONTRACT_REVISION_PREFLIGHT=enforce
withholds the verdict, alerts and exits 2 on mismatch; =warn logs and
continues; =off skips. Verified live: match -> contract proceeds and PASSes;
mismatch -> 'VERDICT WITHHELD', exit 2; =warn -> continues.
Pinning (docs/contract-execution-pinning.md): every contract pins the clone
contract-run.sh lives in - the deployed runner being /opt/contract-runner on
CT 100. Documented that daily-health-digest has no contract file at all, which
is why its execution copy was silently operator-chosen.
Tests: tests/test_revision_preflight.sh, 15 assertions over a throwaway clone
with a real bare remote. It runs the pre-fix draft against the same cases and
shows it passing a ghost script, so the tests provably bite.
shellcheck: scripts/revision-preflight.sh and the new test are clean. The three
findings remaining in contract-run.sh (SC2086 x2, SC2034) are pre-existing and
byte-identical on master.
The Proxmox leg of the daily digest has been reporting NOTHING while exiting 0.
Root cause: the auth header was a literal placeholder string,
AUTH = "Authorization: PVEAPIToken=«vault: infrastructure/production PVE_API_TOKEN»"
which was sent verbatim. The API rejected it, pve_get() returned None, and the
report rendered node_count=0 / nodes_online=0 with pve_probe_status='unreachable'
while still exiting 0. A monitoring gap that looks like a healthy run.
Also: the vault key is PVE_TOKEN, not PVE_API_TOKEN, so even reading os.environ
by the old name would not have found it.
Fixes:
* pve_auth() resolves the token at call time from PVE_TOKEN (injected by
'infisical run --env=prod'). Nothing is hardcoded; a missing token raises.
* pve_get() builds the command inside its try block, so a missing token degrades
to None instead of escaping as an unhandled exception.
* PROBE_FAILURES records an unreachable node/resources probe. Probe failures are
deliberately separate from DEGRADED_LEGS: a missing credential stays exit 0
(existing intent), but a probe with no data now exits 1 in both the report and
--json paths, so it cannot pass unnoticed.
Measured effect on the live host: pve_probe_status unreachable -> ok,
node_count 0 -> 5, nodes_online 0 -> 5, total_vms 0 -> 22, running_vms 0 -> 22.
Tests: 4 new regression tests; all 4 fail against the pre-fix script and pass
after, and the 3 pre-existing tests still pass (7/7).
Not fixed here (needs the captain): the email leg fails with
'534 5.7.9 Application-specific password required' - EMAIL_PASSWORD in the vault
is not a valid Gmail app password for jtabiri@gmail.com. That is a credential
action, not a code change.
- Implement four-state PBS GC logic in proxmox-monitor.sh:
- probe-failed: unparseable JSON, store not found, or empty body → FAIL
- running: collection in progress (last-run-endtime absent, upid present) → DO NOT FAIL
- stale: no completed run within 48h → FAIL, naming last completed run age
- healthy: completed within 48h → PASS, naming endtime and pending bytes
- Add tests/test_pbs_gc_states.sh covering all four states
- Proves the test bites on the pre-fix version (5/6 tests fail)
- All 6 tests pass against the fixed version
- Update contract-run.sh to map disk-gc-threat-response -> scripts/disk-gc-scan.py
- Update Execution sections of host-scheduled contracts:
infrastructure-monitoring, zulip-health, litellm-health, agent-health-check,
disk-gc-threat-response, pm2-self-heal
Adding note that execution is host-scheduled via cron, not agent session ack.
- Add scripts/contract-run.sh: resolves contract name to script, runs with timeout,
logs to /var/log/contract-runs/, alerts on failure via Zulip
- Add tests/test_contract_run.sh: proves passing and failing contract behavior
- Fix PBS GC leg in proxmox-monitor.sh: simplify logic to check last-run-endtime,
use absolute paths for pct and proxmox-backup-manager to avoid PATH issues
Part of Task: contract-execution-host-scheduler-20260924
The 2026-09-17 purge removed six live credentials that had sat in this repo
for weeks, several in .md prose. Nothing blocked that class of commit, so a
warning in a stream nobody reads was the only signal. This adds a guard that
fails the build instead of warning.
Guard
- scripts/secret-scan.sh: bash + coreutils + grep/sed/awk + git only (the Gitea
Actions runner executes job steps inside the runner container — BusyBox grep,
no node/python). Modes: --tree (git-tracked, default), --path DIR (no git),
--staged (pre-commit), --diff REF. Exit 1 on a finding, 2 on config error.
- scripts/secret-patterns.tsv: checked-in pattern list — sk-, sk-or-v1-,
sk_live_, literal Bearer tokens, PVEAPIToken=, raw Authorization values, PEM
private-key blocks, prose credential lines, and password/api_key/secret/token
assignments carrying a literal value. Prose is scanned exactly like code.
- scripts/secret-allowlist.tsv: one entry per deliberate synthetic example, each
with a reason. A missing reason is a hard error (fail closed). The 2026-09-17
purge's `«vault: ...»` placeholders are listed explicitly rather than filtered
by a general "vault"/"synthetic" rule, so a new occurrence still needs a
reviewed, reasoned entry.
- A small inert-value classifier drops env refs, paths, dotted code access,
variable names and right-truncated redactions; it does not know the words
"synthetic"/"example", so a fabrication is always an explicit exception.
- Findings are printed with the credential masked; a scan never echoes a full
secret into the log.
Wiring
- .gitea/workflows/pr-pipeline.yaml lint job: explicit "Committed-credential
scan" step plus the self-test. A finding fails the required
`pr-pipeline / lint` context, which the merge gate depends on.
- scripts/prose-lint.sh (the local gate): a "Secret scan" section, so
`bash scripts/prose-lint.sh` before pushing is equivalent to CI.
Tests
- tests/test_secret_scan.sh: 20 cases. Plants pattern-matching fixtures in temp
trees (outside every allowlisted path) and asserts the guard FAILS, including
the --staged commit-time path; asserts the tree is quiet; asserts allowlisted
text at an unlisted path still fails (path-explicit, not word-based); asserts
a reasonless allowlist entry exits 2.
Verified: guard run against 8245716^ (the pre-fix revision, before the purge)
fails on the real OpenRouter/LiteLLM/Zulip/Proxmox/Stirling credentials; guard
run over the current tree is clean.
- C3 502 → INCIDENT
- C3 000 → INCIDENT
- C1 401 + C3 302 → 0 issues, all healthy
- C1 000 → INCIDENT
Each test asserts from the run's own log/verdict, not from file text.
Prose assertions kept as secondary.
Proven to bite: run against pre-fix script (origin/master) shows all 4
behavioural cases fail because C3 leg doesn't exist and Result line
doesn't say 'INCIDENT'.
- (b) Changed Result line to say 'INCIDENT' when ISSUES > 0, '0 issues (all healthy)' when ISSUES = 0
- (c) Documented C1 (no credential needed), C2 (requires LITELLM_KEY) distinction
- (d) Added C3 public access path leg for https://kagentz.sysloggh.net/
- C3 treats 200/302/401 as alive, 502/000 as incident
- Added tests/test_zulip_kagentz_legs.py to verify all changes
Rule 5 now treats internal http://192.168.68.116/v1 as non-canonical
but working (authenticated via nginx), producing a WARNING instead of a
FAILURE. The canonical internal path /litellm/v1 and the public host
https://litellm.sysloggh.net/v1 both PASS. Everything else FAILS.
Prose aligned: hermes-key-enforcement.prose.md now states the canonical
internal form, notes that internal /v1 still works but is flagged as
non-canonical (WARN not FAIL), and clarifies that the public host serves
/v1 ONLY (404 on /litellm/v1). Corrected the 'unauthenticated path'
wording at line 116, which was factually wrong.
Tests updated: BASE template uses canonical internal path; new test
cases prove canonical /litellm/v1 PASSES, wrong path FAILS, public host
PASSES, and internal /v1 WARNS (not FAILS). Fixed backwards comment in
test_old_rule5_check_would_fail_canonical.
Rule 5 in audit-hermes-config.py had an inverted check: it expected
base_url=http://192.168.68.116/v1, but the contract hermes-key-enforcement.prose.md
names http://192.168.68.116/litellm/v1 as CORRECT/CANONICAL in multiple places.
The audit script would FAIL a config using the contract's canonical internal path
and PASS one using a path the contract does not name.
Fix: Rule 5 now accepts the canonical internal base (http://192.168.68.116/litellm/v1)
AND the public base (https://litellm.sysloggh.net/v1), and FAILS anything else.
The internal nginx serves both /litellm/v1 and /v1; the public host serves /v1 only
(per 2026-09-19 probe from CT 116).
Tests aligned: BASE template updated to use the canonical internal path, and new
test cases added to prove the canonical internal path PASSES, a wrong path FAILS,
and the public host path PASSES.
- Remove retired names (qwen3.6-27B-code, qwen3.6-35B-udq4) from live alias claims
- Update Strix Halo model to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe, 256K ctx)
- Fix litellm-health step 7 probe to gpu-vision (monitor key scoped)
- Move qwen3.6-27B-code/35B-udq4 from raw-but-live to non-resolving in audit
- Fold in pm2-self-heal: remove spoton-service (live PM2 set is 4/4)
- Update hermes templates, key enforcement, timeout tables to live names
CT 111 (tdunna, 192.168.68.129) is Theo's box and is report-only per the captain
(2026-08-17, re-confirmed 2026-09-10). The contract defined AMBER as 'GC scheduled
for next run' and its Execution loop called gc-executor for EVERY threat with no
guest-level exclusion - so a single AMBER reading there would have scheduled apt
clean / journal vacuum / log+tmp deletion against someone else's box. The only
marker was frontmatter report_only_agents, which names an AGENT while the scan unit
is a GUEST.
- Gate the Execution loop on a guest/host-keyed report_only_guests block (guest id,
hostname and IP all match); an excluded guest is alerted and skipped, so no
gc-executor call is constructed for it at any level.
- Carry the ruling in the contract body next to the loop, not only in frontmatter.
- scripts/disk-gc-plan.py: executable planner that reads the contract's
authoritative exclusion block and emits the action plan; tests/ covers it.
- Correct the stale fleet map against pvesh /cluster/resources: CT 105 -> amdpve
(was minipve), CT 111 -> storepve (was amdpve), add guests 118/119/120, and fix
the '15 CTs' counts (20 guests: 17 LXC + 3 QEMU VMs).
- scripts/pct-run.sh: same stale map (105/111 wrong node, 120 missing) - this is
why pct-run 111/105 failed.
- Fold in disk-gc-ct100-probe-gap-20260911: CT 100 verifiably works through
pct-run now that the map is correct; documented that the scanner must probe it
like any other guest, never via a local-only path (the scanner runs inside CT 100).
audit-hermes-config.py Rule 8 required auxiliary.vision.model and
auxiliary.web_extract.model to equal the retired 'gpu-light', so a config
adopting the live canonical 'gpu-vision' FAILED our own audit - the audit was
enforcing a dead alias (400 Invalid model name). Rule 8 now requires
gpu-vision; retired names gpu-light/crew-auto join the raw-name rejection set;
the guidance message names the live aliases.
Sweep of the remaining references: gpu-self-heal stops canonicalizing
gpu-light; hermes-config-template, hermes-agent-baseline, hermes-key-enforcement,
inference-optimization, litellm-client-timeouts and gpu-fleet now use the live
gpu-vision alias. Where a file restated model/rpm/weight/fallback state it now
points at CT 116 /opt/inference-harness/litellm_config.yaml instead of
duplicating it. koby's .129 config is report-only and recorded, not edited.
Adds tests/test_audit_hermes_config_alias.py: executes the audit CLI and asserts
gpu-vision passes while gpu-light and gemma-4-12b fail.
Captain ruling 2026-09-10: Mumuni moved off this host onto her own container
(kagentz CT 105 on minipve, 192.168.68.14, dedicated `hermes` user) and is
monitored from her side. This host must not monitor anything Mumuni.
The stale probes fired false alerts repeatedly:
* scripts/zulip-monitor.sh ssh'd to root@192.168.68.24 for the
decommissioned deployment's ~/.hermes/gateway_state.json, read "unknown"
on every run, and posted a 🔴 "Mumuni (Hermes) Zulip state: unknown" DM +
#agent-hub stream alert each cycle.
* scripts/daily-infra-report.py published a matching "mumuni:unknown" row in
every digest.
Changes:
* zulip-monitor.sh: delete the "Platform B: Hermes (Mumuni)" leg and its
notify; keep the Zulip-server, Platform A pi/Abiba (bridge), Platform B
Tanko and Platform C Agent Zero legs. A comment records why the leg is
retired so it is not re-added. Also fixes SC2155 so shellcheck is clean.
* daily-infra-report.py: delete the .24 ~/.hermes/gateway_state.json agent
probe, its hermes --version probe, and the now-dead mumuni render branch.
Abiba (CT 100, its own .24 address) and Tanko legs unchanged; the PVE API
token name is untouched.
* zulip-health.prose.md (v3.1.0): drop the Mumuni-only B4 gateway-process,
B5 heartbeat and B6 response-delivery steps and the stale 192.168.68.24
references; state explicitly that Mumuni is not monitored from this host.
Tanko/Agent-Zero/bridge steps retained.
* agent-health-check.py: correct the v2 changelog roster comment that still
placed mumuni at .24/CT100. No behavior change — the mumuni probe was
already absent from the AGENTS dict; v5 changelog notes the correction.
Tests: tests/test_mumuni_monitor_removal.py pins the removal structurally and
behaviorally — the shipped zulip-monitor.sh is run in a sandbox (only its LOG
constant rewritten) with stub ssh/curl on PATH; the ssh stub records every
target host, so "never reaches .24" and "no Mumuni notify even on the alert
path" are asserted from observed behavior. A mutation check (re-inject the old
leg) fails the suite, so the guarantee is not vacuous.
Second probe-drift correction pass after #65/#66/#68. All four legs were stale
consumer expectations, not live faults.
1. scripts/agent-health-check.py (v4)
- abiba declared pi-only runtime (harness purge): Hermes-era gateway, config
and wrapper legs are skipped instead of failing.
- koby declared report_only (captain ruling 2026-08-17, Rule 17): every koby
leg is detected and reported, never counted as a fleet failure or repaired.
- koby's CT 111 mapping corrected to storepve (.6); the old amdpve mapping
made `pct status 111` fail and read as ct-unreachable.
- wrapper check no longer FAILs .env-based wrappers that legitimately never
invoke infisical (koonimo).
- keys load in main() (load_agent_keys) so the module is importable/testable.
- every run prints absolute execution provenance (script + cwd), in the header
and in --json.
Before: 6 FAILURE(S). After: 0 failures, koby reported read-only.
2. infrastructure-monitoring.prose.md
- PVE API probe repointed from CT 116 (no pveproxy, 000) to the five real
nodes on https://<node>:8006/api2/json/version, all 401 = alive.
- any-HTTP-response liveness rule added (401/3xx alive; 000/timeout = DOWN).
- LiteLLM health documented as 301 -> /litellm/health/liveliness, not bare 200.
3. gpu-monitor.prose.md
- GPU health probes on :8080 (or router /health/unified); bare port 80 on a
GPU host is forbidden (no listener -> false DEGRADED).
- router /health/unified 301 -> /gpu/gpu-data documented as alive.
- port-discipline + liveness rule + direct-fallback execution step.
4. Report provenance (all contracts)
- docs/AUTHORING-GUIDE.md documents the rule; scripts/prose-lint.sh enforces
that any **Report format** contract states an absolute path (pwd -P).
- provenance added to gpu-monitor, infrastructure-monitoring, proxmox-monitor.
Tests: tests/test_probe_drift.py (23 passed); prose-lint.sh clean + shellcheck
clean; CI frontmatter validation passes. Evidence with per-leg before/after and
absolute paths: docs/probe-drift-round2-evidence.md.
The monitor's Abiba leg parsed the :9200/health payload at the top level
(d.get('connected',False)) while the pi Zulip extension serves connection
state NESTED at zulip.connected / zulip.last_error (verified against the
extension's startHealthServer handler and the live endpoint). PI_CONNECTED
was therefore always False and every run took the DISCONNECTED path,
restarting a healthy bot: pm2 restarts=8 with the process created
2026-09-09T09:35:09Z, four ❌ Abiba verdicts today (04:23/05:35/06:55/09:35
UTC) and zero ✅, while the Zulip server answered HTTP 200 and the bot kept
heartbeating. The watchdog was the fault, not the connection.
Changes (Abiba leg only; every other leg byte-identical):
- Read the real nested shape: zulip.connected, zulip.last_error and
zulip.messages_processed. The old retry_count branch is DROPPED — the
payload exposes no retry counter (the extension keeps retryCount internal
and never serialises it), so the branch is fabricated and cannot stay.
- Fail-safe restart decision: a fetch error (HTTP 000), non-2xx response,
empty/unparseable body, or payload missing a boolean zulip.connected is a
clearly-labelled PROBE FAILURE (🟠 alert + ⚠️ log line naming reason and
HTTP code) and NEVER calls pm2 restart. pm2 restart runs only on
affirmative zulip.connected=false (🔴/❌ path unchanged in wording).
connected=true with last_error keeps the degraded 🟡 warn-no-restart path.
Regression test (new tests/zulip-monitor-abiba.sh, 36 checks): extracts the
real Abiba leg from between the # -- abiba-leg-start/-end markers in the
shipped script and executes it verbatim with stubbed curl/notify/pm2 against
fixtures of the real payload shape — asserts connected=true → no restart,
connected=false → restart, and empty/garbage/missing-key/non-boolean/HTTP
000/HTTP 500 bodies → probe-failure alert with zero restarts. The suite
fails loudly on the pre-fix base and on a marker-intact top-level-parse
variant, so this class of bug cannot return silently.