fix: move infra-monitoring probes into versioned script #115
Merged
abiba-bot
merged 9 commits from 2026-09-18 18:58:28 +00:00
fix/infra-monitoring-probe-targets-drift-20260917 into master
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
933cfd223b |
fix(infra): PR #115 round 4 — fix TLS detection + remove duplicate probe
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
Round 3 had: 1. rc captured from wrong command (tr always exits 0) 2. Every probe issued TWICE (26 invocations instead of 13) 3. TLS branch unreachable Round 4 fixes: - Restructure probe_http to ONE invocation that captures both output and status: out=$(...); rc=$? - Delete the duplicated block - Fix retry classification: don't overwrite kind if already set (e.g., tls) - Add test 7b: TLS error (000 + exit 60) → kind is tls - Update header output shape to include (<kind>) suffix Test: 22 passed / 0 failed (bash scripts/test_infra_monitoring.sh) |
||
|
|
7e257ce512 |
fix(infra): PR #115 final round — complete F1/C1 + implement TLS detection
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Failing after 12m11s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
F1: Set LAST_KIND on unexpected-status path (was unset, causing empty
placeholder in 7 failure lines).
F2: Implement TLS detection — capture curl exit code and map TLS
error codes (35|51|58|59|60|77|83) to kind=tls. Previously
TLS failures were misdiagnosed as timeout.
F3: Header comment now lists all producible kinds:
timeout | refused | tls | unexpected:<code>.
F4: Add 2 test assertions: Grafana failure line exists + kind
is non-empty (proves the gap that shipped in round 2).
Test: 20 passed / 0 failed (bash scripts/test_infra_monitoring.sh)
|
||
|
|
3c7f5d7d65 |
fix(infra+zulip): PR #115 round-2 findings F1-F4+F7
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
F1: Make kind classification real — append (<kind>) to every failure
line, assign kind=tls on curl exit 35/60, fix :112 where
kind=refused was set on successful retry. Update prose shape.
F3: Move credential placeholder skip from server probe to notify()
only — server is always probed (200 without auth verified live).
F4: notify() logs ALERT SUPPRESSED when credential unusable so
alerts from other legs are not silently dropped.
F7: Restore trailing newline in infra-monitoring.sh.
F5 (DO NOT CHANGE): Verified directly — ssh root@192.168.68.6
'grep -n keep-daily /etc/pve/jobs.cfg' returns five
prune-backups keep-daily=35 lines. Prose is CORRECT.
Test: 18 passed, 0 failed (bash scripts/test_infra_monitoring.sh)
|
||
|
|
03be9b13d0 |
fix(zulip-health): skip server leg when credential is placeholder
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 12s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
When ZULIP_API_KEY is unset or contains 'placeholder'/'REDACTED', skip the global Zulip server leg with a ⏭ marker instead of failing the whole script. The pi/Tanko/kagentz legs do not need the Zulip API key and keep their verdicts. Tracked as: zulip-health-credential-placeholder-20260913 (captain-held) This removes the repeated 'Action required' noise every cycle while keeping the credential enforcement loud and visible. |
||
|
|
c295322c85 |
fix(test): stub curl/ssh to assert actual call targets (A2+A3)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 17s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
Rewrote test to run the monitor with stubbed curl/ssh on PATH that
capture the exact argv of each probe call. The test now asserts the
URL+port of every leg actually requested, not source text or config
constants.
A2: PVE node assertions now check the exact URL in the curl log
(https://192.168.68.9:8006/... must appear), so a wrong IP
(e.g. .99) fails the test.
A3: Grafana/Prometheus/LiteLLM assertions check the URL the call
actually builds, so a hardcoded wrong port in the CALL (while the
config variable stays correct) fails the test.
Mutation evidence:
A2: sed s/192.168.68.9/192.168.68.99/ in PVE_NODES -> suite FAILS
A3: sed s/"$GRAFANA_PORT"/"9999"/ in probe call -> suite FAILS
Results: 18 passed, 0 failed (baseline); 17/18 on each mutation
|
||
|
|
385f7e0623 |
fix(infra-monitoring): resolve PR #115 review findings (A1-A3, B1-B2, C1-C2)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
A1: Test -k assertion now checks use_k:+-k syntax (actual bash pattern)
A2: PVE_NODES assertions now count expected nodes and verify exact array size
A3: Test now asserts liveness behavior (PVE_API_LIVENESS=1) not source text
B1: disk-gc GC schedule corrected: cron runs pbs-gc.sh (not proxmox-backup-manager),
schedule is 20:00 LOCAL (00:00 UTC, not 20:00 UTC), host timezone America/New_York
B2: PROBE SHAPE now documents actual output shape including TLS flag notes
C1: TLS kind is now printed in PVE API failure output
C2: SSH retry logic clarified - retry is in probe_http function (not unreachable)
|
||
|
|
7efbfffe44 |
docs(disk-gc): clarify media vs pbs-datastore HOST-RED escalation
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
HOST-RED on media volumes says 'capacity decision — owner to decide' not 'immediate owner attention'. Media volumes are report-only at all levels; the urgency language was misleading. PBS datastore and host-root get the immediate attention wording. Closes the welcome-back proposal: 'how the disk-gc check should classify a media volume so HOST-RED stops meaning nothing.' |
||
|
|
315fcbae23 |
docs(disk-gc): clarify GC schedule applies to PBS datastore only
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
The 20:00 UTC PBS GC cron (proxmox-backup-manager datastore prune) applies only to /tank/pbs-backup (pbs-datastore). It does NOT touch media volumes (/media/*) which are report-only at all threat levels. This clarification prevents the recurring confusion where a 96% media volume triggers a GC expectation, when the GC schedule never applies to it. Closes the 2026-09-17 correction: 'the GC schedule is now 20:00 UTC, protects the backup datastore, NOT the nearly-full media volume.' |
||
|
|
93f15709d1 |
fix(infra-monitoring): move probes to versioned script with port-drift test
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The 2026-09-17 false-verdict incident (third recurrence) showed that prose policy is not a control: the agent probed :9325/:9405 (nonexistent ports), CT 116 for PVE API (should be real PVE nodes), and rendered TLS failures as connection-refused. This moves the canonical probe set into scripts/infra-monitoring.sh (executed verbatim by the contract) and adds scripts/test_infra_monitoring.sh which asserts every probed port matches the documented value. - scripts/infra-monitoring.sh: one script per contract pattern; all targets, ports, paths, and expected-status rules in code; -k for PVE self-signed certs; non-zero exit naming every failed target; no OK summary on failure - scripts/test_infra_monitoring.sh: 20 assertions covering port drift, monitoring-host-as-PVE-node, and missing -k flag - infrastructure-monitoring.prose.md: check-health section now references the script as executable owner; paste its raw output verbatim Proof: all 13 legs pass (exit 0); deliberately broken Grafana port (9325) produces 'probe-failed: 192.168.68.116:9325 (expected 200)' and exit 1. |