Per defect report 1150.msg:
- GPU exporters: probe /metrics (Prometheus scrape target), not bare /
- Grafana: correct port 3001 (not 3000)
- All probes: print target name + full URL + HTTP code
- All probes: retry once at 25s on 000/timeout
- Apply standing rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
Verified: all probes now return real HTTP codes (GPU 200, Grafana 200, Router 200, LiteLLM 301)
Contract design defect: 'uses a non-harness provider' (POLICY) and
'cannot authenticate' (FAULT) were printed as the same violation class.
A policy observation must never be phrased as if the agent were broken.
Changes:
1. Added 'Violation Classification' section to all three contracts
2. Separated POLICY (observation only) from FAULT (requires request-level evidence)
3. Rules:
- Do NOT infer runtime credential resolution from config text alone
- Require request-level evidence before calling a FAULT: observed auth failure
or absence of successful calls
- If calls are succeeding, output is 'POLICY: uses <provider> directly; calls
succeeding' - not a violation
- State what you OBSERVED, not what the field implies
Files changed (3):
- hermes-key-enforcement.prose.md
- hermes-config-template.prose.md
- hermes-agent-baseline.prose.md
The helper was correct but dead code - nothing called it. This commit:
1. Makes the helper runnable standalone: scripts/hermes-reachability-check.sh <host> <pattern> <path>
2. Wires all THREE contracts to it:
- hermes-key-enforcement.prose.md
- hermes-config-template.prose.md
- hermes-agent-baseline.prose.md
3. Each contract now explicitly instructs to run the helper and interpret the three outcomes
4. States that the bug this replaces was deriving reachability from the remote grep's exit code
Files changed (4):
- scripts/hermes-reachability-check.sh (standalone mode added)
- hermes-key-enforcement.prose.md (reachability section added)
- hermes-config-template.prose.md (reachability section added)
- hermes-agent-baseline.prose.md (reachability section added)
Bug: The reachability check used 'ssh ... grep ... || echo unreachable',
which conflated grep's 'no matches found' (exit 1) with SSH failure.
This caused clean hosts to be reported as unreachable.
Fix: Add scripts/hermes-reachability-check.sh with the pattern:
out=$(ssh -o BatchMode=yes root@HOST "grep ... 2>/dev/null; true")
if [ $? -ne 0 ]; then verdict="unreachable"
elif [ -n "$out" ]; then verdict="violation: $out"
else verdict="compliant"
fi
This correctly distinguishes:
- SSH failure (connection/auth/route) -> unreachable
- SSH success + grep found matches -> violation
- SSH success + grep found nothing -> compliant
Evidence: 3 of 4 hosts (Tanko, Mumuni, Koonimo) were reported as
'unreachable' when they were actually compliant. Only Koby (.129)
has a real finding (plaintext key in state snapshot).
Used by: hermes-key-enforcement, hermes-config-template,
hermes-agent-baseline contracts.
1. Timeout fix for pool alias (syslog-auto):
- Single-host aliases (gpu-dense, gpu-vision, strix-moe): 30s timeout
- Pool alias (syslog-auto): 60s timeout, retry once on 000 before failing
- Cold first request to pool alias can take ~13s; 10s was too short
2. Remove DEBUG prints from output:
- Removed 'DEBUG: keylen=...' and 'DEBUG: response=...' lines
- These leaked key inventory to status logs
- Success output now shows only counts (e.g., 'Admin Key List: 10 keys')
Add 'Executor Script' section:
- Run scripts/litellm-health-check.py from the clone
- Hand-rolled probes not acceptable substitute
- Backend-edge checks use internal IP 192.168.68.116, not public URL
- Docker Stats fetched from CT 116 host (127.0.0.1:9324/metrics)
- Admin Key List requires proper quoting for SSH commands
Implements all 11 checks from litellm-health contract:
- Liveliness, Containers, Prometheus, Grafana health probes
- Model probes for gpu-dense, gpu-vision, strix-moe, syslog-auto
- Admin API key list (10 keys), GitHub status, Docker Stats metrics
Fixed quoting for SSH commands and response parsing (dict with 'keys' field).
Backend edge uses internal IP 192.168.68.116, not public URL.
Docker Stats fetched from CT 116 host itself (127.0.0.1:9324/metrics).
All 11 checks passing consistently.
The abiba-zulip-restore.prose.md file incorrectly claimed the default model
was 'deepseek-v4-pro', but /root/.pi/agent/settings.json declares
'defaultModel: syslog-auto' and 'defaultProvider: syslog-harness'.
deepseek-v4-pro is not a servable LiteLLM model name for this fleet.
The only live models are: syslog-auto, gpu-dense, gpu-vision, strix-moe.
Note: The firstmate/pi session talks to DeepSeek through its own
provider (auth.json), NOT through LiteLLM, so no LiteLLM change is
needed to keep firstmate's deepseek usage working.
- harness-litellm container has no curl/wget, so docker exec harness-litellm curl returns empty
- Fix: run curl on CT 116 host (ssh root@192.168.68.116 then curl)
- Add admin-call-failed label for empty/unparseable responses
- Verify: 10 keys found (host-side curl), NO-CURL confirmed in container
- infrastructure-monitoring: Correct Zulip key path to /etc/litellm-monitor.env
(not /etc/zulip-bot.env which doesn't exist); add credential-missing check;
remove unused monitor_key variable
- litellm-health: Clarify that syslog-auto is a fallback pool alias, not a
step 7 probe; monitor key must be scoped for all four aliases
- litellm-health.prose.md: Make monitor key retrieval explicit (ssh from CT 100 to CT 116)
with executable commands; add syslog-auto alias; document credential-missing failure
condition (not bare 401 or 0 keys)
- infrastructure-monitoring.prose.md: Fix Zulip POST probe to retrieve keys from CT 116
via ssh instead of sourcing local env file that doesn't exist on executor host
- Verify model inference probes return 200 for gpu-dense, gpu-vision, strix-moe,
syslog-auto with corrected credential retrieval
- Remove retired names (qwen3.6-27B-code, qwen3.6-35B-udq4) from live alias claims
- Update Strix Halo model to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe, 256K ctx)
- Fix litellm-health step 7 probe to gpu-vision (monitor key scoped)
- Move qwen3.6-27B-code/35B-udq4 from raw-but-live to non-resolving in audit
- Fold in pm2-self-heal: remove spoton-service (live PM2 set is 4/4)
- Update hermes templates, key enforcement, timeout tables to live names
Intent requires a production run log, not only the unit-level dry run. This is a real
scan of the live 25-entry fleet through scripts/disk-gc-plan.py: CT 111 (tdunna) at 84%
AMBER is alerted as report-only and no gc-executor row is emitted for it; the owned hosts
acerpve .9 (77%) and amdpve .15 (76%) still receive gc-executor. No GC command was run
against 192.168.68.129.
CT 111 (tdunna, 192.168.68.129) is Theo's box and is report-only per the captain
(2026-08-17, re-confirmed 2026-09-10). The contract defined AMBER as 'GC scheduled
for next run' and its Execution loop called gc-executor for EVERY threat with no
guest-level exclusion - so a single AMBER reading there would have scheduled apt
clean / journal vacuum / log+tmp deletion against someone else's box. The only
marker was frontmatter report_only_agents, which names an AGENT while the scan unit
is a GUEST.
- Gate the Execution loop on a guest/host-keyed report_only_guests block (guest id,
hostname and IP all match); an excluded guest is alerted and skipped, so no
gc-executor call is constructed for it at any level.
- Carry the ruling in the contract body next to the loop, not only in frontmatter.
- scripts/disk-gc-plan.py: executable planner that reads the contract's
authoritative exclusion block and emits the action plan; tests/ covers it.
- Correct the stale fleet map against pvesh /cluster/resources: CT 105 -> amdpve
(was minipve), CT 111 -> storepve (was amdpve), add guests 118/119/120, and fix
the '15 CTs' counts (20 guests: 17 LXC + 3 QEMU VMs).
- scripts/pct-run.sh: same stale map (105/111 wrong node, 120 missing) - this is
why pct-run 111/105 failed.
- Fold in disk-gc-ct100-probe-gap-20260911: CT 100 verifiably works through
pct-run now that the map is correct; documented that the scanner must probe it
like any other guest, never via a local-only path (the scanner runs inside CT 100).
The validate job runs with bash -e -o pipefail. `echo "$FM" | grep -q '^name:'`
lets grep exit on first match, which can SIGPIPE the echo; pipefail then reports
the pipeline non-zero and the || branch raises a false "Missing name/description".
The flagged file set varied run to run (and included files untouched by the PR)
while a fresh clone of the same commit passes the identical check. Reproduced:
the old form failed 3 of 5 local runs under the same shell flags, the herestring
form passed 5 of 5. Use herestrings so no pipe can be broken.
audit-hermes-config.py Rule 8 required auxiliary.vision.model and
auxiliary.web_extract.model to equal the retired 'gpu-light', so a config
adopting the live canonical 'gpu-vision' FAILED our own audit - the audit was
enforcing a dead alias (400 Invalid model name). Rule 8 now requires
gpu-vision; retired names gpu-light/crew-auto join the raw-name rejection set;
the guidance message names the live aliases.
Sweep of the remaining references: gpu-self-heal stops canonicalizing
gpu-light; hermes-config-template, hermes-agent-baseline, hermes-key-enforcement,
inference-optimization, litellm-client-timeouts and gpu-fleet now use the live
gpu-vision alias. Where a file restated model/rpm/weight/fallback state it now
points at CT 116 /opt/inference-harness/litellm_config.yaml instead of
duplicating it. koby's .129 config is report-only and recorded, not edited.
Adds tests/test_audit_hermes_config_alias.py: executes the audit CLI and asserts
gpu-vision passes while gpu-light and gemma-4-12b fail.