Infrastructure-monitoring produced false service verdicts (probed non-existent ports, probed wrong host for PVE API, mislabelled TLS failures) while declaring the run OK.
Solution
Move all probes into a versioned script (scripts/infra-monitoring.sh) with targets/ports/paths in code, -k for PVE API, SSH probes for localhost bindings, non-zero exit with named failures.
## Problem
Infrastructure-monitoring produced false service verdicts (probed non-existent ports, probed wrong host for PVE API, mislabelled TLS failures) while declaring the run OK.
## Solution
Move all probes into a versioned script (scripts/infra-monitoring.sh) with targets/ports/paths in code, -k for PVE API, SSH probes for localhost bindings, non-zero exit with named failures.
## Tests
- tests/test_infra_monitoring.py
- Proved happy path + broken target
Fixes: infra-monitoring-probe-targets-drift-20260917
The 2026-09-17 false-verdict incident (third recurrence) showed that prose
policy is not a control: the agent probed :9325/:9405 (nonexistent ports),
CT 116 for PVE API (should be real PVE nodes), and rendered TLS failures as
connection-refused. This moves the canonical probe set into
scripts/infra-monitoring.sh (executed verbatim by the contract) and adds
scripts/test_infra_monitoring.sh which asserts every probed port matches the
documented value.
- scripts/infra-monitoring.sh: one script per contract pattern; all targets,
ports, paths, and expected-status rules in code; -k for PVE self-signed
certs; non-zero exit naming every failed target; no OK summary on failure
- scripts/test_infra_monitoring.sh: 20 assertions covering port drift,
monitoring-host-as-PVE-node, and missing -k flag
- infrastructure-monitoring.prose.md: check-health section now references the
script as executable owner; paste its raw output verbatim
Proof: all 13 legs pass (exit 0); deliberately broken Grafana port (9325)
produces 'probe-failed: 192.168.68.116:9325 (expected 200)' and exit 1.
The 20:00 UTC PBS GC cron (proxmox-backup-manager datastore prune)
applies only to /tank/pbs-backup (pbs-datastore). It does NOT touch
media volumes (/media/*) which are report-only at all threat levels.
This clarification prevents the recurring confusion where a 96% media
volume triggers a GC expectation, when the GC schedule never applies
to it.
Closes the 2026-09-17 correction: 'the GC schedule is now 20:00 UTC,
protects the backup datastore, NOT the nearly-full media volume.'
HOST-RED on media volumes says 'capacity decision — owner to decide' not
'immediate owner attention'. Media volumes are report-only at all levels; the
urgency language was misleading. PBS datastore and host-root get the immediate
attention wording.
Closes the welcome-back proposal: 'how the disk-gc check should classify a
media volume so HOST-RED stops meaning nothing.'
A1: Test -k assertion now checks use_k:+-k syntax (actual bash pattern)
A2: PVE_NODES assertions now count expected nodes and verify exact array size
A3: Test now asserts liveness behavior (PVE_API_LIVENESS=1) not source text
B1: disk-gc GC schedule corrected: cron runs pbs-gc.sh (not proxmox-backup-manager),
schedule is 20:00 LOCAL (00:00 UTC, not 20:00 UTC), host timezone America/New_York
B2: PROBE SHAPE now documents actual output shape including TLS flag notes
C1: TLS kind is now printed in PVE API failure output
C2: SSH retry logic clarified - retry is in probe_http function (not unreachable)
Rewrote test to run the monitor with stubbed curl/ssh on PATH that
capture the exact argv of each probe call. The test now asserts the
URL+port of every leg actually requested, not source text or config
constants.
A2: PVE node assertions now check the exact URL in the curl log
(https://192.168.68.9:8006/... must appear), so a wrong IP
(e.g. .99) fails the test.
A3: Grafana/Prometheus/LiteLLM assertions check the URL the call
actually builds, so a hardcoded wrong port in the CALL (while the
config variable stays correct) fails the test.
Mutation evidence:
A2: sed s/192.168.68.9/192.168.68.99/ in PVE_NODES -> suite FAILS
A3: sed s/"$GRAFANA_PORT"/"9999"/ in probe call -> suite FAILS
Results: 18 passed, 0 failed (baseline); 17/18 on each mutation
When ZULIP_API_KEY is unset or contains 'placeholder'/'REDACTED', skip
the global Zulip server leg with a ⏭ marker instead of failing the whole
script. The pi/Tanko/kagentz legs do not need the Zulip API key and keep
their verdicts.
Tracked as: zulip-health-credential-placeholder-20260913 (captain-held)
This removes the repeated 'Action required' noise every cycle while
keeping the credential enforcement loud and visible.
F1: Make kind classification real — append (<kind>) to every failure
line, assign kind=tls on curl exit 35/60, fix :112 where
kind=refused was set on successful retry. Update prose shape.
F3: Move credential placeholder skip from server probe to notify()
only — server is always probed (200 without auth verified live).
F4: notify() logs ALERT SUPPRESSED when credential unusable so
alerts from other legs are not silently dropped.
F7: Restore trailing newline in infra-monitoring.sh.
F5 (DO NOT CHANGE): Verified directly — ssh root@192.168.68.6
'grep -n keep-daily /etc/pve/jobs.cfg' returns five
prune-backups keep-daily=35 lines. Prose is CORRECT.
Test: 18 passed, 0 failed (bash scripts/test_infra_monitoring.sh)
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Problem
Infrastructure-monitoring produced false service verdicts (probed non-existent ports, probed wrong host for PVE API, mislabelled TLS failures) while declaring the run OK.
Solution
Move all probes into a versioned script (scripts/infra-monitoring.sh) with targets/ports/paths in code, -k for PVE API, SSH probes for localhost bindings, non-zero exit with named failures.
Tests
Fixes: infra-monitoring-probe-targets-drift-20260917
5ab5de4704to93f15709d1A1: Test -k assertion now checks use_k:+-k syntax (actual bash pattern) A2: PVE_NODES assertions now count expected nodes and verify exact array size A3: Test now asserts liveness behavior (PVE_API_LIVENESS=1) not source text B1: disk-gc GC schedule corrected: cron runs pbs-gc.sh (not proxmox-backup-manager), schedule is 20:00 LOCAL (00:00 UTC, not 20:00 UTC), host timezone America/New_York B2: PROBE SHAPE now documents actual output shape including TLS flag notes C1: TLS kind is now printed in PVE API failure output C2: SSH retry logic clarified - retry is in probe_http function (not unreachable)Rewrote test to run the monitor with stubbed curl/ssh on PATH that capture the exact argv of each probe call. The test now asserts the URL+port of every leg actually requested, not source text or config constants. A2: PVE node assertions now check the exact URL in the curl log (https://192.168.68.9:8006/... must appear), so a wrong IP (e.g. .99) fails the test. A3: Grafana/Prometheus/LiteLLM assertions check the URL the call actually builds, so a hardcoded wrong port in the CALL (while the config variable stays correct) fails the test. Mutation evidence: A2: sed s/192.168.68.9/192.168.68.99/ in PVE_NODES -> suite FAILS A3: sed s/"$GRAFANA_PORT"/"9999"/ in probe call -> suite FAILS Results: 18 passed, 0 failed (baseline); 17/18 on each mutationF1: Make kind classification real — append (<kind>) to every failure line, assign kind=tls on curl exit 35/60, fix :112 where kind=refused was set on successful retry. Update prose shape. F3: Move credential placeholder skip from server probe to notify() only — server is always probed (200 without auth verified live). F4: notify() logs ALERT SUPPRESSED when credential unusable so alerts from other legs are not silently dropped. F7: Restore trailing newline in infra-monitoring.sh. F5 (DO NOT CHANGE): Verified directly — ssh root@192.168.68.6 'grep -n keep-daily /etc/pve/jobs.cfg' returns five prune-backups keep-daily=35 lines. Prose is CORRECT. Test: 18 passed, 0 failed (bash scripts/test_infra_monitoring.sh)F1: Set LAST_KIND on unexpected-status path (was unset, causing empty placeholder in 7 failure lines). F2: Implement TLS detection — capture curl exit code and map TLS error codes (35|51|58|59|60|77|83) to kind=tls. Previously TLS failures were misdiagnosed as timeout. F3: Header comment now lists all producible kinds: timeout | refused | tls | unexpected:<code>. F4: Add 2 test assertions: Grafana failure line exists + kind is non-empty (proves the gap that shipped in round 2). Test: 20 passed / 0 failed (bash scripts/test_infra_monitoring.sh)