1. Replace '(## Execution' + ')' with '## Execution' in 6 contract files
2. Make LOG_DIR honor CONTRACT_RUN_LOG_DIR env var (default: /var/log/contract-runs)
3. Fix Zulip alert: correct URL (https://chat.sysloggh.net/api/v1), user
(abiba-bot@chat.sysloggh.net), take ZULIP_API_KEY from environment, and
write alert failures to run log
- Implement four-state PBS GC logic in proxmox-monitor.sh:
- probe-failed: unparseable JSON, store not found, or empty body → FAIL
- running: collection in progress (last-run-endtime absent, upid present) → DO NOT FAIL
- stale: no completed run within 48h → FAIL, naming last completed run age
- healthy: completed within 48h → PASS, naming endtime and pending bytes
- Add tests/test_pbs_gc_states.sh covering all four states
- Proves the test bites on the pre-fix version (5/6 tests fail)
- All 6 tests pass against the fixed version
- Update contract-run.sh to map disk-gc-threat-response -> scripts/disk-gc-scan.py
- Update Execution sections of host-scheduled contracts:
infrastructure-monitoring, zulip-health, litellm-health, agent-health-check,
disk-gc-threat-response, pm2-self-heal
Adding note that execution is host-scheduled via cron, not agent session ack.
Previously probed wrong ports:
- Docker Stats was at 9323 (dockerd metrics) but should be 9324
(harness-docker-stats, docker_container_* metrics)
- PVE Exporter was at 9324 (harness-docker-stats) but should be 9221
(harness-pve-exporter, 5 pve_* metrics)
Both exporters bind to 127.0.0.1 on CT 116 and must be probed via SSH.
Updated infrastructure-monitoring.prose.md to document the correct ports.
Added test assertions verifying the exact ports are probed.
Branch: fix/infra-monitoring-probe-ports-20260919
F1: Make kind classification real — append (<kind>) to every failure
line, assign kind=tls on curl exit 35/60, fix :112 where
kind=refused was set on successful retry. Update prose shape.
F3: Move credential placeholder skip from server probe to notify()
only — server is always probed (200 without auth verified live).
F4: notify() logs ALERT SUPPRESSED when credential unusable so
alerts from other legs are not silently dropped.
F7: Restore trailing newline in infra-monitoring.sh.
F5 (DO NOT CHANGE): Verified directly — ssh root@192.168.68.6
'grep -n keep-daily /etc/pve/jobs.cfg' returns five
prune-backups keep-daily=35 lines. Prose is CORRECT.
Test: 18 passed, 0 failed (bash scripts/test_infra_monitoring.sh)
A1: Test -k assertion now checks use_k:+-k syntax (actual bash pattern)
A2: PVE_NODES assertions now count expected nodes and verify exact array size
A3: Test now asserts liveness behavior (PVE_API_LIVENESS=1) not source text
B1: disk-gc GC schedule corrected: cron runs pbs-gc.sh (not proxmox-backup-manager),
schedule is 20:00 LOCAL (00:00 UTC, not 20:00 UTC), host timezone America/New_York
B2: PROBE SHAPE now documents actual output shape including TLS flag notes
C1: TLS kind is now printed in PVE API failure output
C2: SSH retry logic clarified - retry is in probe_http function (not unreachable)
The 2026-09-17 false-verdict incident (third recurrence) showed that prose
policy is not a control: the agent probed :9325/:9405 (nonexistent ports),
CT 116 for PVE API (should be real PVE nodes), and rendered TLS failures as
connection-refused. This moves the canonical probe set into
scripts/infra-monitoring.sh (executed verbatim by the contract) and adds
scripts/test_infra_monitoring.sh which asserts every probed port matches the
documented value.
- scripts/infra-monitoring.sh: one script per contract pattern; all targets,
ports, paths, and expected-status rules in code; -k for PVE self-signed
certs; non-zero exit naming every failed target; no OK summary on failure
- scripts/test_infra_monitoring.sh: 20 assertions covering port drift,
monitoring-host-as-PVE-node, and missing -k flag
- infrastructure-monitoring.prose.md: check-health section now references the
script as executable owner; paste its raw output verbatim
Proof: all 13 legs pass (exit 0); deliberately broken Grafana port (9325)
produces 'probe-failed: 192.168.68.116:9325 (expected 200)' and exit 1.
Per defect report 1150.msg:
- GPU exporters: probe /metrics (Prometheus scrape target), not bare /
- Grafana: correct port 3001 (not 3000)
- All probes: print target name + full URL + HTTP code
- All probes: retry once at 25s on 000/timeout
- Apply standing rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
Verified: all probes now return real HTTP codes (GPU 200, Grafana 200, Router 200, LiteLLM 301)
- infrastructure-monitoring: Correct Zulip key path to /etc/litellm-monitor.env
(not /etc/zulip-bot.env which doesn't exist); add credential-missing check;
remove unused monitor_key variable
- litellm-health: Clarify that syslog-auto is a fallback pool alias, not a
step 7 probe; monitor key must be scoped for all four aliases
- litellm-health.prose.md: Make monitor key retrieval explicit (ssh from CT 100 to CT 116)
with executable commands; add syslog-auto alias; document credential-missing failure
condition (not bare 401 or 0 keys)
- infrastructure-monitoring.prose.md: Fix Zulip POST probe to retrieve keys from CT 116
via ssh instead of sourcing local env file that doesn't exist on executor host
- Verify model inference probes return 200 for gpu-dense, gpu-vision, strix-moe,
syslog-auto with corrected credential retrieval
Second probe-drift correction pass after #65/#66/#68. All four legs were stale
consumer expectations, not live faults.
1. scripts/agent-health-check.py (v4)
- abiba declared pi-only runtime (harness purge): Hermes-era gateway, config
and wrapper legs are skipped instead of failing.
- koby declared report_only (captain ruling 2026-08-17, Rule 17): every koby
leg is detected and reported, never counted as a fleet failure or repaired.
- koby's CT 111 mapping corrected to storepve (.6); the old amdpve mapping
made `pct status 111` fail and read as ct-unreachable.
- wrapper check no longer FAILs .env-based wrappers that legitimately never
invoke infisical (koonimo).
- keys load in main() (load_agent_keys) so the module is importable/testable.
- every run prints absolute execution provenance (script + cwd), in the header
and in --json.
Before: 6 FAILURE(S). After: 0 failures, koby reported read-only.
2. infrastructure-monitoring.prose.md
- PVE API probe repointed from CT 116 (no pveproxy, 000) to the five real
nodes on https://<node>:8006/api2/json/version, all 401 = alive.
- any-HTTP-response liveness rule added (401/3xx alive; 000/timeout = DOWN).
- LiteLLM health documented as 301 -> /litellm/health/liveliness, not bare 200.
3. gpu-monitor.prose.md
- GPU health probes on :8080 (or router /health/unified); bare port 80 on a
GPU host is forbidden (no listener -> false DEGRADED).
- router /health/unified 301 -> /gpu/gpu-data documented as alive.
- port-discipline + liveness rule + direct-fallback execution step.
4. Report provenance (all contracts)
- docs/AUTHORING-GUIDE.md documents the rule; scripts/prose-lint.sh enforces
that any **Report format** contract states an absolute path (pwd -P).
- provenance added to gpu-monitor, infrastructure-monitoring, proxmox-monitor.
Tests: tests/test_probe_drift.py (23 passed); prose-lint.sh clean + shellcheck
clean; CI frontmatter validation passes. Evidence with per-leg before/after and
absolute paths: docs/probe-drift-round2-evidence.md.
PR #50's content was squash-merged out-of-band (2dfc3e1, 2026-08-30) while
the PR itself stayed open, double-applying the check-health/Execution
section. This drops the second copy; keeps one canonical
Execution -> check-health -> Phases 1-4 -> Verification Commands flow.
Fixes the canned 'API UNREACHABLE (HTTP 000)' reporting. Adds an explicit
check-health execution section (Zulip POST, pm2, GPU exporters, Prometheus,
Grafana, LiteLLM probes) with a hard RUN LIVE, NEVER ECHO rule, mirroring
the gpu-monitor contract pattern. Captain priority 2026-08-22.