The contract said the state file was written after every scan, but the script
had no state-file logic at all. This PR adds:
1. Host filesystem scanning (probe_host_filesystems) - probes df on all PVE nodes
2. State file I/O (read_state_file/write_state_file) - absolute path from script location
3. Band classification (classify_band) - HOST-WARN/AMBER/RED thresholds
4. Transition detection (detect_transitions) - alerts on escalation/recovery
5. CLI flags (--hosts-only, --guests-only) to control which parts run
The contract now specifies the state file path resolves from the script's own
location (not CWD-relative), so two different execution contexts cannot write
to two different places.
A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER,
AMBER->RED) and ONCE when it drops back down (recovery notice). While it stays
in the same band, it is reported in the scan output only — no DM, no channel
alert. This stops the same 96% easystore2 from re-DMing the owner on every 6h
scan.
State lives in a small JSON state file (state/host-disk-bands.json), keyed by
host/volume -> last-seen band. The scanner reads the prior band, compares to the
current band, and DMs only on a transition; the state file is written after
every scan. Chosen over a periodic digest because the scan already runs every
6h and a transition is genuinely new, actionable state.
First-run behavior: when the state file does not yet exist, the current band of
every volume is recorded as baseline WITHOUT alerting — a first run would
otherwise DM every already-elevated volume at once.
Report-only restriction and volume-naming output kept exactly as-is.
Add separate threat bands for HOST filesystems (distinct from guest bands):
- HOST-WARN at 85%: name volume + % + absolute free space in scan output
- HOST-AMBER at 90%: flag for owner attention, Zulip DM
- HOST-RED at 95%: flag for immediate owner attention, Zulip DM + channel alert
Volume naming rule: every host line MUST name the volume and what lives on it.
Action classes by volume type:
- host-root: near full = real risk (backup staging, thin-pool metadata)
- media (/media/*): near full = capacity decision for owner, never auto-delete
- pbs-datastore (tank): near full = breaks Proxmox Backup Server
Report-only restriction: no automatic deletion of media or datastore content ever.
Justification (measured 2026-09-15): storepve /media/easystore2 at 96% was
reported but never banded or acted on. Two incidents this weekend showed the
host filesystem is the thing that breaks, not the guest's.
Added HOST-WARN/AMBER/RED alert templates.
Added report-only execution rule for host filesystems.
pm2-self-heal.prose.md:
- Add AS-BUILT note: gpu-monitor is systemd-managed, NOT PM2
- gpu-watchdog is decommissioned and folded into gpu-monitor.service
- gitea-runner is KEPT; abiba-zulip is KEPT (online for days)
- spoton-service was deleted; live PM2 set is 4 processes
- Preserve historical context for crash-loop guard
litellm-health.prose.md:
- Correct Prometheus node coverage: 6 nodes (.4/.5/.6/.9/.12/.15:9100)
- Note .4:9100 is DEAD target (no route, down for weeks)
- Clarify this does not read as 6 healthy nodes
Add warning that bare PVE hostnames (acerpve, amdpve, etc.) resolve to VPS
via *.dns.sysloggh.net wildcard, not to actual nodes. List IP addresses:
- acerpve 192.168.68.9
- amdpve 192.168.68.15
- storepve 192.168.68.6
- minipve 192.168.68.12
- ocupve 192.168.68.5
Update acerpve example in backup preflight to include address (192.168.68.9).
Fix dmsetup comment to show full field order:
=start =length =thin-pool =transaction-id
=metadata_used/metadata_total =data_used/data_total
remaining fields are flags
Make it clear lvs command is the primary source for percentages, dmsetup is only for error-state check.
Incidents now include addresses: acerpve (192.168.68.9) and amdpve (192.168.68.15).
Fix two errors in the PREFLIGHT section (measured on acerpve 2026-09-15):
1. lvs -o ... pve/data (not pve-data-tpool) - this is the PRIMARY check that yields percentages directly
- Quote the acerpve example: data 29.95% 1.22% <816.21g
2. dmsetup status pve-data-tpool - document fields correctly:
- = transaction ID (99), NOT data_percent
- = metadata used/total blocks
- = data used/total sectors
- Show how to derive percentages if needed
3. Keep the error-state check (grep -q 'Error|Fail') - this is how the incident presented
Everything else stays: 1777 tmpdir requirement with EACCES symptom, incidents as rationale,
GPU-host fact, --output-format json rule, honest note that metadata/snapshot pressure is unproven.
Add documented preflight checks for VM/CT backups on LVM thin-pool hosts:
- dmsetup status pve-data-tpool + lvs to verify data_percent < 90% and metadata_percent < 70%
- Exit 1 if pool shows Error/Fail state (takes down entire VG including host root)
- tmpdir must be mode 1777 (world-traversable) for vzdump archive step
- --output-format json for tasks started from truncating shells
- Document two incidents: acerpve thin-pool VM 101 (twice on 2026-09-13) and amdpve 0700 tmpdir (2026-09-14)
- Note metadata/snapshot-pressure hypothesis is UNPROVEN; preflight is the control
- Document GPU-host fact: VM 101 (llm-gpu) and VM 103 (ocu-llm) have no scheduled backup
- get_response_body() returns first 200 chars of response body (single line)
- On 401/403 model probe: report code + body + key_alias
- Monitor key alias: monitor-20260813 (from /etc/litellm-monitor.env on CT 116)
- Failed connections stay probe-failed, 200 stays plain 200
- Do not turn other statuses into credential faults
Signed-off-by: Abiba
- Replace Infisical retrieval path with proven docker exec + .env note
- State explicitly that master key is NOT in Infisical project=infrastructure
- Keep the live-key check and never-trust-a-literal instruction
- All other corrections from PR #95 preserved
Signed-off-by: Abiba
- probe_http now returns (code, failure_kind) tuple
- Model probes report 'probe-failed: <model> <kind> (Ns timeout)' on 000
- Do not assert a service verdict from a failed probe
- 30s timeout for single-host aliases (RTX 3090 needs long warmup/prefill)
- 60s timeout for syslog-auto pool alias with retry on 000
Signed-off-by: Abiba
- Replace broken --plain flag (prints nothing on CLI 0.43.110) with awk parsing
- Note that --plain is broken so nobody fixes it back
- Replace unproven nginx path with verified direct endpoint http://127.0.0.1:4000/key/list
Signed-off-by: Abiba
- hermes-key-enforcement.prose.md:
- State that expiry must be set EXPLICITLY at creation with duration
- Record that config default is NOT honoured by LiteLLM 1.99.1
- Describe daily audit as AUDIT-ONLY (reports non-expiring and soon-to-expire)
- State that renewal is NOT implemented
- Document exclusions: abiba-pi and all crewmate keys stay WITHOUT expiry
- koby is report-only
- litellm-api-keys.prose.md:
- Replace literal master key with retrieval path (docker exec + infisical)
- State that literal values must never be trusted again (key rotates)
- Add live-key check (200 from /key/list)
Signed-off-by: Abiba
Per defect report 1154.msg (agent-health-gateway-leg-flap-20260913):
1. Add _ssh_retry() helper with one retry at longer timeout (25s)
2. Name probe target explicitly: ssh {user}@{host} <command>
3. Print probe-failed when first attempt fails, then retry
4. Only declare gateway-down after retry fails
5. Make Koby report-only explicit in output
Changes:
- check_agents(): all gateway probes now use _ssh_retry()
- All output lines name the probe target (ssh host:port)
- Koby's report-only status is explicit in output
- Never print bare "gateway down" — always name target and failure kind
Verified: koonimo shows "probe-failed" on first attempt (transient SSH),
retries at 25s, succeeds, reports ✅ koonimo: gw=running
Per defect report 1150.msg:
- Add standing probe rules section (2026-09-14)
- Step 1 (Zulip API): retry once at 25s on 000, print target + code
- Step 2 (Platform A): retry once at 25s on 000, print target + code, note must run on Abiba host
- Apply same shape as infrastructure-monitoring: any HTTP status = ALIVE; only 000/timeout = probe-failed
Verified: all probes now return real HTTP codes (Zulip API 200, Platform A 200, Tanko 401, Agent Zero 401)
Per defect report 1150.msg:
- GPU exporters: probe /metrics (Prometheus scrape target), not bare /
- Grafana: correct port 3001 (not 3000)
- All probes: print target name + full URL + HTTP code
- All probes: retry once at 25s on 000/timeout
- Apply standing rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
Verified: all probes now return real HTTP codes (GPU 200, Grafana 200, Router 200, LiteLLM 301)
Contract design defect: 'uses a non-harness provider' (POLICY) and
'cannot authenticate' (FAULT) were printed as the same violation class.
A policy observation must never be phrased as if the agent were broken.
Changes:
1. Added 'Violation Classification' section to all three contracts
2. Separated POLICY (observation only) from FAULT (requires request-level evidence)
3. Rules:
- Do NOT infer runtime credential resolution from config text alone
- Require request-level evidence before calling a FAULT: observed auth failure
or absence of successful calls
- If calls are succeeding, output is 'POLICY: uses <provider> directly; calls
succeeding' - not a violation
- State what you OBSERVED, not what the field implies
Files changed (3):
- hermes-key-enforcement.prose.md
- hermes-config-template.prose.md
- hermes-agent-baseline.prose.md
The helper was correct but dead code - nothing called it. This commit:
1. Makes the helper runnable standalone: scripts/hermes-reachability-check.sh <host> <pattern> <path>
2. Wires all THREE contracts to it:
- hermes-key-enforcement.prose.md
- hermes-config-template.prose.md
- hermes-agent-baseline.prose.md
3. Each contract now explicitly instructs to run the helper and interpret the three outcomes
4. States that the bug this replaces was deriving reachability from the remote grep's exit code
Files changed (4):
- scripts/hermes-reachability-check.sh (standalone mode added)
- hermes-key-enforcement.prose.md (reachability section added)
- hermes-config-template.prose.md (reachability section added)
- hermes-agent-baseline.prose.md (reachability section added)
Bug: The reachability check used 'ssh ... grep ... || echo unreachable',
which conflated grep's 'no matches found' (exit 1) with SSH failure.
This caused clean hosts to be reported as unreachable.
Fix: Add scripts/hermes-reachability-check.sh with the pattern:
out=$(ssh -o BatchMode=yes root@HOST "grep ... 2>/dev/null; true")
if [ $? -ne 0 ]; then verdict="unreachable"
elif [ -n "$out" ]; then verdict="violation: $out"
else verdict="compliant"
fi
This correctly distinguishes:
- SSH failure (connection/auth/route) -> unreachable
- SSH success + grep found matches -> violation
- SSH success + grep found nothing -> compliant
Evidence: 3 of 4 hosts (Tanko, Mumuni, Koonimo) were reported as
'unreachable' when they were actually compliant. Only Koby (.129)
has a real finding (plaintext key in state snapshot).
Used by: hermes-key-enforcement, hermes-config-template,
hermes-agent-baseline contracts.
1. Timeout fix for pool alias (syslog-auto):
- Single-host aliases (gpu-dense, gpu-vision, strix-moe): 30s timeout
- Pool alias (syslog-auto): 60s timeout, retry once on 000 before failing
- Cold first request to pool alias can take ~13s; 10s was too short
2. Remove DEBUG prints from output:
- Removed 'DEBUG: keylen=...' and 'DEBUG: response=...' lines
- These leaked key inventory to status logs
- Success output now shows only counts (e.g., 'Admin Key List: 10 keys')
Add 'Executor Script' section:
- Run scripts/litellm-health-check.py from the clone
- Hand-rolled probes not acceptable substitute
- Backend-edge checks use internal IP 192.168.68.116, not public URL
- Docker Stats fetched from CT 116 host (127.0.0.1:9324/metrics)
- Admin Key List requires proper quoting for SSH commands
Implements all 11 checks from litellm-health contract:
- Liveliness, Containers, Prometheus, Grafana health probes
- Model probes for gpu-dense, gpu-vision, strix-moe, syslog-auto
- Admin API key list (10 keys), GitHub status, Docker Stats metrics
Fixed quoting for SSH commands and response parsing (dict with 'keys' field).
Backend edge uses internal IP 192.168.68.116, not public URL.
Docker Stats fetched from CT 116 host itself (127.0.0.1:9324/metrics).
All 11 checks passing consistently.
The abiba-zulip-restore.prose.md file incorrectly claimed the default model
was 'deepseek-v4-pro', but /root/.pi/agent/settings.json declares
'defaultModel: syslog-auto' and 'defaultProvider: syslog-harness'.
deepseek-v4-pro is not a servable LiteLLM model name for this fleet.
The only live models are: syslog-auto, gpu-dense, gpu-vision, strix-moe.
Note: The firstmate/pi session talks to DeepSeek through its own
provider (auth.json), NOT through LiteLLM, so no LiteLLM change is
needed to keep firstmate's deepseek usage working.
- harness-litellm container has no curl/wget, so docker exec harness-litellm curl returns empty
- Fix: run curl on CT 116 host (ssh root@192.168.68.116 then curl)
- Add admin-call-failed label for empty/unparseable responses
- Verify: 10 keys found (host-side curl), NO-CURL confirmed in container
- infrastructure-monitoring: Correct Zulip key path to /etc/litellm-monitor.env
(not /etc/zulip-bot.env which doesn't exist); add credential-missing check;
remove unused monitor_key variable
- litellm-health: Clarify that syslog-auto is a fallback pool alias, not a
step 7 probe; monitor key must be scoped for all four aliases
- litellm-health.prose.md: Make monitor key retrieval explicit (ssh from CT 100 to CT 116)
with executable commands; add syslog-auto alias; document credential-missing failure
condition (not bare 401 or 0 keys)
- infrastructure-monitoring.prose.md: Fix Zulip POST probe to retrieve keys from CT 116
via ssh instead of sourcing local env file that doesn't exist on executor host
- Verify model inference probes return 200 for gpu-dense, gpu-vision, strix-moe,
syslog-auto with corrected credential retrieval
- Remove retired names (qwen3.6-27B-code, qwen3.6-35B-udq4) from live alias claims
- Update Strix Halo model to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe, 256K ctx)
- Fix litellm-health step 7 probe to gpu-vision (monitor key scoped)
- Move qwen3.6-27B-code/35B-udq4 from raw-but-live to non-resolving in audit
- Fold in pm2-self-heal: remove spoton-service (live PM2 set is 4/4)
- Update hermes templates, key enforcement, timeout tables to live names
Intent requires a production run log, not only the unit-level dry run. This is a real
scan of the live 25-entry fleet through scripts/disk-gc-plan.py: CT 111 (tdunna) at 84%
AMBER is alerted as report-only and no gc-executor row is emitted for it; the owned hosts
acerpve .9 (77%) and amdpve .15 (76%) still receive gc-executor. No GC command was run
against 192.168.68.129.
CT 111 (tdunna, 192.168.68.129) is Theo's box and is report-only per the captain
(2026-08-17, re-confirmed 2026-09-10). The contract defined AMBER as 'GC scheduled
for next run' and its Execution loop called gc-executor for EVERY threat with no
guest-level exclusion - so a single AMBER reading there would have scheduled apt
clean / journal vacuum / log+tmp deletion against someone else's box. The only
marker was frontmatter report_only_agents, which names an AGENT while the scan unit
is a GUEST.
- Gate the Execution loop on a guest/host-keyed report_only_guests block (guest id,
hostname and IP all match); an excluded guest is alerted and skipped, so no
gc-executor call is constructed for it at any level.
- Carry the ruling in the contract body next to the loop, not only in frontmatter.
- scripts/disk-gc-plan.py: executable planner that reads the contract's
authoritative exclusion block and emits the action plan; tests/ covers it.
- Correct the stale fleet map against pvesh /cluster/resources: CT 105 -> amdpve
(was minipve), CT 111 -> storepve (was amdpve), add guests 118/119/120, and fix
the '15 CTs' counts (20 guests: 17 LXC + 3 QEMU VMs).
- scripts/pct-run.sh: same stale map (105/111 wrong node, 120 missing) - this is
why pct-run 111/105 failed.
- Fold in disk-gc-ct100-probe-gap-20260911: CT 100 verifiably works through
pct-run now that the map is correct; documented that the scanner must probe it
like any other guest, never via a local-only path (the scanner runs inside CT 100).
The validate job runs with bash -e -o pipefail. `echo "$FM" | grep -q '^name:'`
lets grep exit on first match, which can SIGPIPE the echo; pipefail then reports
the pipeline non-zero and the || branch raises a false "Missing name/description".
The flagged file set varied run to run (and included files untouched by the PR)
while a fresh clone of the same commit passes the identical check. Reproduced:
the old form failed 3 of 5 local runs under the same shell flags, the herestring
form passed 5 of 5. Use herestrings so no pipe can be broken.
audit-hermes-config.py Rule 8 required auxiliary.vision.model and
auxiliary.web_extract.model to equal the retired 'gpu-light', so a config
adopting the live canonical 'gpu-vision' FAILED our own audit - the audit was
enforcing a dead alias (400 Invalid model name). Rule 8 now requires
gpu-vision; retired names gpu-light/crew-auto join the raw-name rejection set;
the guidance message names the live aliases.
Sweep of the remaining references: gpu-self-heal stops canonicalizing
gpu-light; hermes-config-template, hermes-agent-baseline, hermes-key-enforcement,
inference-optimization, litellm-client-timeouts and gpu-fleet now use the live
gpu-vision alias. Where a file restated model/rpm/weight/fallback state it now
points at CT 116 /opt/inference-harness/litellm_config.yaml instead of
duplicating it. koby's .129 config is report-only and recorded, not edited.
Adds tests/test_audit_hermes_config_alias.py: executes the audit CLI and asserts
gpu-vision passes while gpu-light and gemma-4-12b fail.
- Execution step 2 now documents the public edge and the backend edge as two
distinct surfaces: public serves /ui/ and /docs (404 on the /litellm/ prefix),
backend http://192.168.68.116 serves /litellm/ui/ and /litellm/docs (with /ui/
and /docs as 301 helpers). Each probe names its surface.
- GPU topology: ocu-llm RTX 5070 now serves gpu-vision (gemma-4-12b retired).
- Fallback/timeout table rewritten to the live router_settings.fallbacks chains.
- Step 7 model list: gemma-4-12b -> gpu-vision, with a key-scoped /v1/models note
and the 2026-09-12 master-key registry snapshot.
- Remove adapter process check and restart logic
- Keep A2A probe (port 80, HTTP code check)
- The adapter code at /a0/usr/kagentz-zulip/ no longer exists
- Captain's ruling: Zulip communication with agent zero is not priority
- Load API key from durable file /root/.abiba-workspace/secrets/litellm-key.txt (works in cron)
- Fix http_get to use Bearer token instead of Basic Auth for API endpoint check
- All 6 LiteLLM checks now pass (was 5/6)
- Changed A2A probe from http://127.0.0.1:8001/.well-known/agent.json to
http://127.0.0.1:80/a2a/ inside agent-zero container
- Port 8001 does not exist inside container (nothing listens there)
- Port 80 maps to external port 50080; returns 401 (auth-gated, alive by design)
- Updated A2A_URL in adapter.py restart command to use port 80 instead of 8001
- Verified: probe now returns 401 (auth-gated) instead of 000 (connection refused)
- Before: kagentz A2A reported DOWN on every run (false positive due to stale port)
- After: kagentz A2A reports ✅ A2A alive (auth-gated 401 = healthy)
The monitor's Abiba leg parsed the :9200/health payload at the top level
(d.get('connected',False)) while the pi Zulip extension serves connection
state NESTED at zulip.connected / zulip.last_error (verified against the
extension's startHealthServer handler and the live endpoint). PI_CONNECTED
was therefore always False and every run took the DISCONNECTED path,
restarting a healthy bot: pm2 restarts=8 with the process created
2026-09-09T09:35:09Z, four ❌ Abiba verdicts today (04:23/05:35/06:55/09:35
UTC) and zero ✅, while the Zulip server answered HTTP 200 and the bot kept
heartbeating. The watchdog was the fault, not the connection.
Changes (Abiba leg only; every other leg byte-identical):
- Read the real nested shape: zulip.connected, zulip.last_error and
zulip.messages_processed. The old retry_count branch is DROPPED — the
payload exposes no retry counter (the extension keeps retryCount internal
and never serialises it), so the branch is fabricated and cannot stay.
- Fail-safe restart decision: a fetch error (HTTP 000), non-2xx response,
empty/unparseable body, or payload missing a boolean zulip.connected is a
clearly-labelled PROBE FAILURE (🟠 alert + ⚠️ log line naming reason and
HTTP code) and NEVER calls pm2 restart. pm2 restart runs only on
affirmative zulip.connected=false (🔴/❌ path unchanged in wording).
connected=true with last_error keeps the degraded 🟡 warn-no-restart path.
Regression test (new tests/zulip-monitor-abiba.sh, 36 checks): extracts the
real Abiba leg from between the # -- abiba-leg-start/-end markers in the
shipped script and executes it verbatim with stubbed curl/notify/pm2 against
fixtures of the real payload shape — asserts connected=true → no restart,
connected=false → restart, and empty/garbage/missing-key/non-boolean/HTTP
000/HTTP 500 bodies → probe-failure alert with zero restarts. The suite
fails loudly on the pre-fix base and on a marker-intact top-level-parse
variant, so this class of bug cannot return silently.
The #agent-hub / zulip-health stream post in notify() encoded content with
python quote(str()) (always empty), so every stream alert posted empty
content and Zulip rejected it silently behind '|| true'. The merged body
also used backslash-escaped ampersands inside double quotes, which curl
transmits literally (type=stream\ -> Zulip 400 'Invalid type').
Stream post now percent-encodes the real message text by piping it through
the encoder (locale-proof: quote_from_bytes on stdin.buffer), uses plain &
separators, and on failure appends one WARN line to the monitor log with
the curl exit status instead of silently swallowing it. Delivery failure
stays non-fatal and unretried. DM path, probes, alive rule, exit codes and
log format unchanged.
The no-mistakes document step had auto-updated the .8 systemd unit name
(llama-server -> llama-chat-api.service) in gpu-fleet.prose.md and
infrastructure-control.prose.md. Those contracts are co-gated with another
agent and cannot be amended inside a script-scope task (firstmate decision
2026-09-08, key prose-doc-step-scope); the unit-name doc sync is separate
follow-up work. This restores both files to their origin/master content.
The litellm-self-heal.prose.md agent-health-check v3/env.sh description
update (same document commit) is intentional and kept.
- GPU unit repoint verified live 2026-09-08: .8 rtx3090 probes
llama-chat-api.service (stale llama-server unit read inactive -> false
UNREACHABLE for a healthy process); .110 keeps llama-server.service
(ocu-llm VM), .15 keeps strix-server.service. is-active no longer
swallowed as SSH failure (|| true).
- check_agents: bind pid before the summary f-string so the non-report-only
path (abiba/koonimo) no longer raises UnboundLocalError (line 305 crash).
- abiba key leg: read LITELLM_API_KEY from /root/.pi/agent/env.sh (#735
moved creds out of shared /root/.bashrc); 'abiba NO KEY' gone on healthy
setup.
Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only), not 0.0.0.0.
They must be probed from .116 via SSH. Prometheus and Grafana remain 0.0.0.0 (LAN-reachable).
Fixes false alarm from remote probe showing connection-refused (by design for localhost binds).
Fixes the canned 'API UNREACHABLE (HTTP 000)' reporting. Adds an explicit
check-health execution section (Zulip POST, pm2, GPU exporters, Prometheus,
Grafana, LiteLLM probes) with a hard RUN LIVE, NEVER ECHO rule, mirroring
the gpu-monitor contract pattern. Captain priority 2026-08-22.