- Implement four-state PBS GC logic in proxmox-monitor.sh:
- probe-failed: unparseable JSON, store not found, or empty body → FAIL
- running: collection in progress (last-run-endtime absent, upid present) → DO NOT FAIL
- stale: no completed run within 48h → FAIL, naming last completed run age
- healthy: completed within 48h → PASS, naming endtime and pending bytes
- Add tests/test_pbs_gc_states.sh covering all four states
- Proves the test bites on the pre-fix version (5/6 tests fail)
- All 6 tests pass against the fixed version
- Update contract-run.sh to map disk-gc-threat-response -> scripts/disk-gc-scan.py
- Update Execution sections of host-scheduled contracts:
infrastructure-monitoring, zulip-health, litellm-health, agent-health-check,
disk-gc-threat-response, pm2-self-heal
Adding note that execution is host-scheduled via cron, not agent session ack.
- Add scripts/contract-run.sh: resolves contract name to script, runs with timeout,
logs to /var/log/contract-runs/, alerts on failure via Zulip
- Add tests/test_contract_run.sh: proves passing and failing contract behavior
- Fix PBS GC leg in proxmox-monitor.sh: simplify logic to check last-run-endtime,
use absolute paths for pct and proxmox-backup-manager to avoid PATH issues
Part of Task: contract-execution-host-scheduler-20260924
The 2026-09-17 purge removed six live credentials that had sat in this repo
for weeks, several in .md prose. Nothing blocked that class of commit, so a
warning in a stream nobody reads was the only signal. This adds a guard that
fails the build instead of warning.
Guard
- scripts/secret-scan.sh: bash + coreutils + grep/sed/awk + git only (the Gitea
Actions runner executes job steps inside the runner container — BusyBox grep,
no node/python). Modes: --tree (git-tracked, default), --path DIR (no git),
--staged (pre-commit), --diff REF. Exit 1 on a finding, 2 on config error.
- scripts/secret-patterns.tsv: checked-in pattern list — sk-, sk-or-v1-,
sk_live_, literal Bearer tokens, PVEAPIToken=, raw Authorization values, PEM
private-key blocks, prose credential lines, and password/api_key/secret/token
assignments carrying a literal value. Prose is scanned exactly like code.
- scripts/secret-allowlist.tsv: one entry per deliberate synthetic example, each
with a reason. A missing reason is a hard error (fail closed). The 2026-09-17
purge's `«vault: ...»` placeholders are listed explicitly rather than filtered
by a general "vault"/"synthetic" rule, so a new occurrence still needs a
reviewed, reasoned entry.
- A small inert-value classifier drops env refs, paths, dotted code access,
variable names and right-truncated redactions; it does not know the words
"synthetic"/"example", so a fabrication is always an explicit exception.
- Findings are printed with the credential masked; a scan never echoes a full
secret into the log.
Wiring
- .gitea/workflows/pr-pipeline.yaml lint job: explicit "Committed-credential
scan" step plus the self-test. A finding fails the required
`pr-pipeline / lint` context, which the merge gate depends on.
- scripts/prose-lint.sh (the local gate): a "Secret scan" section, so
`bash scripts/prose-lint.sh` before pushing is equivalent to CI.
Tests
- tests/test_secret_scan.sh: 20 cases. Plants pattern-matching fixtures in temp
trees (outside every allowlisted path) and asserts the guard FAILS, including
the --staged commit-time path; asserts the tree is quiet; asserts allowlisted
text at an unlisted path still fails (path-explicit, not word-based); asserts
a reasonless allowlist entry exits 2.
Verified: guard run against 8245716^ (the pre-fix revision, before the purge)
fails on the real OpenRouter/LiteLLM/Zulip/Proxmox/Stirling credentials; guard
run over the current tree is clean.
- Add zulip-watchdog to Maintains (it's running, infrastructure-monitoring expects it)
- Remove gpu-monitor from PM2 Maintains (it's systemd-only, not PM2-tracked)
- Add Execution steps for all monitored processes (gitea-runner, zulip-watchdog)
- Update alert channel: Telegram is primary, Zulip DM is secondary
- Update script to check all 4 processes (abiba-telegram, abiba-zulip, gitea-runner, zulip-watchdog)
- Add restart count thresholds for all processes
- Update log output to include all process statuses
1. Remove vestigial ZULIP_API_KEY requirement:
- /api/v1/server_settings is a PUBLIC endpoint (verified HTTP 200 with or without credential)
- No Zulip API key is required for this call
- If a future leg genuinely needs abiba-bot's key, it must prove it with a 200 from
/api/v1/users/me as abiba-bot and label itself degraded when it cannot
- Never fall back to the vault's shared ZULIP_API_KEY
2. Make failed sends exit non-zero:
- A degraded leg (no credential configured) must stay exit 0
- A failed send (attempted and failed) must exit 1
- This distinguishes 'not configured' from 'attempted and failed'
Test evidence:
- No-credential run: exit 0, digest still produced
- Wrong password: exit 1, labelled SMTP error
- grep -n ZULIP_API_KEY: only comment reference remains
- ZULIP_API_KEY: no longer SystemExit, now reports 'credential-missing: ZULIP_API_KEY'
- EMAIL_PASSWORD: no longer sys.exit(1), now appends to DEGRADED_LEGS and returns success
- PVE API: fixed None check in storage section
- Summary: reports degraded legs before summary
This allows the digest to be produced and emailed even when credentials are missing,
while still explicitly logging which legs are degraded.
Test: empty env produces JSON report + degraded leg labels, no SystemExit.
- C3 502 → INCIDENT
- C3 000 → INCIDENT
- C1 401 + C3 302 → 0 issues, all healthy
- C1 000 → INCIDENT
Each test asserts from the run's own log/verdict, not from file text.
Prose assertions kept as secondary.
Proven to bite: run against pre-fix script (origin/master) shows all 4
behavioural cases fail because C3 leg doesn't exist and Result line
doesn't say 'INCIDENT'.
- (b) Changed Result line to say 'INCIDENT' when ISSUES > 0, '0 issues (all healthy)' when ISSUES = 0
- (c) Documented C1 (no credential needed), C2 (requires LITELLM_KEY) distinction
- (d) Added C3 public access path leg for https://kagentz.sysloggh.net/
- C3 treats 200/302/401 as alive, 502/000 as incident
- Added tests/test_zulip_kagentz_legs.py to verify all changes
The busy line was rendering as:
'busy (completion timed out after retry; host healthy host healthy (200))'
because host_detail already contains 'host healthy (200)' and the prefix
also said 'host healthy'. Fixed to:
'busy (completion timed out after retry; host healthy (200))'
F1 cosmetic fix from PR #123 verify.
Rule 5 now treats internal http://192.168.68.116/v1 as non-canonical
but working (authenticated via nginx), producing a WARNING instead of a
FAILURE. The canonical internal path /litellm/v1 and the public host
https://litellm.sysloggh.net/v1 both PASS. Everything else FAILS.
Prose aligned: hermes-key-enforcement.prose.md now states the canonical
internal form, notes that internal /v1 still works but is flagged as
non-canonical (WARN not FAIL), and clarifies that the public host serves
/v1 ONLY (404 on /litellm/v1). Corrected the 'unauthenticated path'
wording at line 116, which was factually wrong.
Tests updated: BASE template uses canonical internal path; new test
cases prove canonical /litellm/v1 PASSES, wrong path FAILS, public host
PASSES, and internal /v1 WARNS (not FAILS). Fixed backwards comment in
test_old_rule5_check_would_fail_canonical.
Three states:
- healthy: passed, exit 0 (unchanged)
- busy (completion timed out after retry AND host /health answered):
⚠️ DEGRADED line, does NOT fail the run, exit 0
- host unreachable or real fault: ❌, exit 1 (unchanged)
Summary now reports degraded count:
- All pass, no degraded: '✅ All checks passed'
- All pass, 1+ degraded: '✅ All checks passed (1 degraded: gpu-dense)'
- Some failed: '❌ Some checks failed' or '❌ Some checks failed (1 degraded: ...)'
Host health mapping verified:
- gpu-dense -> 192.168.68.8:8080/health
- gpu-vision -> 192.168.68.110:8080/health
- strix-moe -> 192.168.68.15:8080/health
1. TIMEOUT KIND FIX: run_command returns (1, '', 'TIMEOUT') when its own
timeout fires. probe_http now checks for this before falling through to
'curl exit <rc>', so a 30s timeout reports 'timeout after 30s' not
'curl exit 1'.
2. BUSY/DEGRADED DETECTION: After both model probes fail, check the
model's host health endpoint (e.g. 192.168.68.8:8080/health for
gpu-dense). If the host answers 200, report 'busy (completion timed
out after retry; host healthy 200)' — do NOT fail the run on that
alone. If the host does not answer, that's a real FAIL.
3. RETRY TIMEOUT INCREASED: Single-host retry timeout raised from 45s to
90s. Worst-case prefill on a single-slot .8 host is ~76s (observed
83K-token prompt at 1078 tok/s), so 90s covers it.
New line shapes:
- Busy: 'gpu-dense: busy (completion timed out after retry; host healthy 200)'
- Real failure: 'probe-failed: gpu-dense timeout after 30s then timeout after 90s (2 attempts)'
Rule 5 in audit-hermes-config.py had an inverted check: it expected
base_url=http://192.168.68.116/v1, but the contract hermes-key-enforcement.prose.md
names http://192.168.68.116/litellm/v1 as CORRECT/CANONICAL in multiple places.
The audit script would FAIL a config using the contract's canonical internal path
and PASS one using a path the contract does not name.
Fix: Rule 5 now accepts the canonical internal base (http://192.168.68.116/litellm/v1)
AND the public base (https://litellm.sysloggh.net/v1), and FAILS anything else.
The internal nginx serves both /litellm/v1 and /v1; the public host serves /v1 only
(per 2026-09-19 probe from CT 116).
Tests aligned: BASE template updated to use the canonical internal path, and new
test cases added to prove the canonical internal path PASSES, a wrong path FAILS,
and the public host path PASSES.
The failure line now preserves both attempts' failure kinds instead of
hardcoding 'timeout after retry, 45s'. If both attempts fail, the report
shows: 'probe-failed: <model> <first kind> then <retry kind> (2 attempts)'.
This fixes the self-contradictory output when the first attempt timed out
but the retry failed with connection refused, and prevents the duration
from appearing twice when both attempts were timeouts.
Example outputs:
- timeout then timeout: 'probe-failed: gpu-dense timeout after 30s then timeout after 45s (2 attempts)'
- timeout then refused: 'probe-failed: gpu-dense timeout after 30s then connection refused (2 attempts)'
- refused then refused: 'probe-failed: gpu-dense connection refused then connection refused (2 attempts)'
Single-host models (gpu-dense, gpu-vision, strix-moe) now retry once at
45s on initial 30s timeout failure before declaring probe-failed. This
prevents a single transient timeout (cold prefill ~13s or concurrent
generation hold) from failing the entire health digest.
Evidence: 2026-09-19 ~06:55Z digest failed gpu-dense at 30s; 06:56Z
direct probe 200 in 1.04s.
The failed-probe-fails-the-run property is preserved: if both attempts
fail, the script still exits non-zero with the target and duration named.
Closes: daily-health-digest false negative on single transient timeout
The leg comments were wrong:
- Line 207: Docker Stats showed :9323 (dockerd port) but should be :9324
- Line 215: PVE Exporter showed :9324 (docker-stats port) but should be :9221
These were the exact pairing this PR exists to correct.
Read back the changed lines to verify:
scripts/infra-monitoring.sh:207 shows Docker Stats (CT 116 :9324, 127.0.0.1 via SSH)
scripts/infra-monitoring.sh:215 shows PVE Exporter (CT 116 :9221, 127.0.0.1 via SSH)
Branch: fix/infra-monitoring-probe-ports-20260919
F1: Fixed leg comments to match the actual ports
- Line 207: Docker Stats now shows :9324 (was :9323)
- Line 215: PVE Exporter now shows :9221 (was :9324)
These were the exact pairing this PR exists to correct.
F2: Added per-leg assertions that prove which leg owns which port
The new assertions verify:
1. Docker Stats leg uses $DOCKER_STATS_PORT constant
2. PVE Exporter leg uses $PVE_EXPORTER_PORT constant
3. DOCKER_STATS_PORT constant is set to 9324
4. PVE_EXPORTER_PORT constant is set to 9221
Proof the new assertions bite:
Under the both-constants-swapped mutation (DOCKER_STATS_PORT=9221,
PVE_EXPORTER_PORT=9324), the suite fails with 25 passed / 2 failed
(failing exactly the two constant-value assertions). This proves the
per-leg assertions pin which leg owns which port, not just that both
ports appear somewhere in the SSH log.
Branch: fix/infra-monitoring-probe-ports-20260919
The Liveness Check section now documents ALL SIX verdict shapes exactly
as emitted by proxmox-monitor.sh:
1. ✅ PBS GC: healthy (last run Nh ago, pending-bytes: N B)
2. 🔴 PBS GC: stale (last run Nh ago, pending-bytes: N B)
3. 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)
4. 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)
5. 🔴 PBS GC: never-run (storepve-datastore not found in GC list)
6. 🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)
Fixed the quoted healthy example (line ~114) to include the pending-bytes
suffix the code now appends. Previously the prose only documented the
probe-failed shape, missing the PR's own headline cases (never-run).
Proof: grep -n 'never-run' proxmox-monitor.prose.md now returns two lines
(lines 106 and 108), documenting both never-run variants.
Branch: fix/pbs-gc-liveness-signal-20260919
Previously probed wrong ports:
- Docker Stats was at 9323 (dockerd metrics) but should be 9324
(harness-docker-stats, docker_container_* metrics)
- PVE Exporter was at 9324 (harness-docker-stats) but should be 9221
(harness-pve-exporter, 5 pve_* metrics)
Both exporters bind to 127.0.0.1 on CT 116 and must be probed via SSH.
Updated infrastructure-monitoring.prose.md to document the correct ports.
Added test assertions verifying the exact ports are probed.
Branch: fix/infra-monitoring-probe-ports-20260919
(a) Probe-failure detection: now treats empty OR unparseable JSON as
probe-failed, not never-run. This prevents 'command not found'
outputs from being rendered as service verdicts.
(b) pending-bytes: now extracted from JSON and reported in stale verdict.
(c) Null endtime: use .get() with explicit None check, not 0 fallback.
null values now correctly trigger never-run verdict instead of
arithmetic crash (set -u).
(d) Tests: Added 14-assertion stub-driven suite covering: healthy,
stale (>48h), probe-failed (empty and unparseable), null endtime,
datastore absent. Each test stubs ssh/curl to verify exact
behavior against the pre-fix head.
Branch: fix/pbs-gc-liveness-signal-20260919
Document the PBS GC schedule (00:00 UTC, not 20:00 UTC as PR #116 said),
what actually runs (pbs-gc.sh -> pct exec 107 -- proxmox-backup-manager),
datastore location (CT 107's /mnt/pbs-backup on /tank/pbs-backup, NOT
/media/easystore2), and the new liveness check (48h threshold, reports
age in hours, explicit healthy line).
Branch: fix/pbs-gc-liveness-signal-20260919
Add monitoring leg that checks storepve-datastore GC health:
- Reads GC state from CT 107 via pct exec
- FAILS if last-run-endtime is older than 48h
- Reports age in hours and pending-bytes status
- Uses JSON parsing for reliable data extraction
Test: All 5 legs OK, Exit 0.
Branch: fix/pbs-gc-liveness-signal-20260919
PR #118 finding F1 (low): The deployed LiteLLM on CT 116 uses
allowed_mcp_servers (193 occurrences in installed package), not
bare allowed_mcp. One-word doc fix.
Add the audit-hermes-config.py Rule 15 wording fix from PR #117:
- Violation message now reads 'URL is incorrect: <url> (expected: <expected>)'
- Detection logic unchanged
- Matches URL and not-in-known-list branches remain byte-identical
This consolidates relay #779 into a single PR (#118).
PR #117 follow-up (verify PASS-WITH-FINDINGS):
1. infrastructure-update.prose.md:
- Update access table: agent keys now have per-key MCP grants (2026-09-18)
- Strike-through old limitation: per-key grants now work
- Mark Migration Path as COMPLETED 2026-09-18
2. hermes-config-template.prose.md:
- Remove hedge ('may have been upgraded')
- State fact: per-key MCP grants verified 2026-09-18
This resolves the contradiction where one file asserted per-key
MCP access and the other denied it.
F1: Make kind classification real — append (<kind>) to every failure
line, assign kind=tls on curl exit 35/60, fix :112 where
kind=refused was set on successful retry. Update prose shape.
F3: Move credential placeholder skip from server probe to notify()
only — server is always probed (200 without auth verified live).
F4: notify() logs ALERT SUPPRESSED when credential unusable so
alerts from other legs are not silently dropped.
F7: Restore trailing newline in infra-monitoring.sh.
F5 (DO NOT CHANGE): Verified directly — ssh root@192.168.68.6
'grep -n keep-daily /etc/pve/jobs.cfg' returns five
prune-backups keep-daily=35 lines. Prose is CORRECT.
Test: 18 passed, 0 failed (bash scripts/test_infra_monitoring.sh)
When ZULIP_API_KEY is unset or contains 'placeholder'/'REDACTED', skip
the global Zulip server leg with a ⏭ marker instead of failing the whole
script. The pi/Tanko/kagentz legs do not need the Zulip API key and keep
their verdicts.
Tracked as: zulip-health-credential-placeholder-20260913 (captain-held)
This removes the repeated 'Action required' noise every cycle while
keeping the credential enforcement loud and visible.