Adds a liveness leg to scripts/proxmox-monitor.sh that reads the local PBS datastore's garbage-collection state from inside CT 107 (where proxmox-backup-manager lives) and FAILS when storepve-datastore's last-run-endtime is older than 48h, reporting the age and pending bytes.
Why
The nightly GC job (/etc/cron.d/pbs-gc -> /usr/local/bin/pbs-gc.sh) had never worked: it called proxmox-backup-manager on the storepve HOST, where only proxmox-backup-client exists, so every run died 'command not found'. The host script is fixed and verified (a real run completed OK), and this leg is the signal that catches a silently dead GC.
To follow on this branch
A proxmox-monitor.prose.md section documenting the schedule and this check (superseding the separate docs PR #116, whose text states the schedule in the wrong timezone - it is 20:00 local = 00:00 UTC, not 20:00 UTC).
Firstmate runs the pre-delivery review and the merge.
## What
Adds a liveness leg to scripts/proxmox-monitor.sh that reads the local PBS datastore's garbage-collection state from inside CT 107 (where proxmox-backup-manager lives) and FAILS when storepve-datastore's last-run-endtime is older than 48h, reporting the age and pending bytes.
## Why
The nightly GC job (/etc/cron.d/pbs-gc -> /usr/local/bin/pbs-gc.sh) had never worked: it called proxmox-backup-manager on the storepve HOST, where only proxmox-backup-client exists, so every run died 'command not found'. The host script is fixed and verified (a real run completed OK), and this leg is the signal that catches a silently dead GC.
## To follow on this branch
A proxmox-monitor.prose.md section documenting the schedule and this check (superseding the separate docs PR #116, whose text states the schedule in the wrong timezone - it is 20:00 local = 00:00 UTC, not 20:00 UTC).
Firstmate runs the pre-delivery review and the merge.
Add monitoring leg that checks storepve-datastore GC health:
- Reads GC state from CT 107 via pct exec
- FAILS if last-run-endtime is older than 48h
- Reports age in hours and pending-bytes status
- Uses JSON parsing for reliable data extraction
Test: All 5 legs OK, Exit 0.
Branch: fix/pbs-gc-liveness-signal-20260919
Document the PBS GC schedule (00:00 UTC, not 20:00 UTC as PR #116 said),
what actually runs (pbs-gc.sh -> pct exec 107 -- proxmox-backup-manager),
datastore location (CT 107's /mnt/pbs-backup on /tank/pbs-backup, NOT
/media/easystore2), and the new liveness check (48h threshold, reports
age in hours, explicit healthy line).
Branch: fix/pbs-gc-liveness-signal-20260919
(a) Probe-failure detection: now treats empty OR unparseable JSON as
probe-failed, not never-run. This prevents 'command not found'
outputs from being rendered as service verdicts.
(b) pending-bytes: now extracted from JSON and reported in stale verdict.
(c) Null endtime: use .get() with explicit None check, not 0 fallback.
null values now correctly trigger never-run verdict instead of
arithmetic crash (set -u).
(d) Tests: Added 14-assertion stub-driven suite covering: healthy,
stale (>48h), probe-failed (empty and unparseable), null endtime,
datastore absent. Each test stubs ssh/curl to verify exact
behavior against the pre-fix head.
Branch: fix/pbs-gc-liveness-signal-20260919
The Liveness Check section now documents ALL SIX verdict shapes exactly
as emitted by proxmox-monitor.sh:
1. ✅ PBS GC: healthy (last run Nh ago, pending-bytes: N B)
2. 🔴 PBS GC: stale (last run Nh ago, pending-bytes: N B)
3. 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)
4. 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)
5. 🔴 PBS GC: never-run (storepve-datastore not found in GC list)
6. 🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)
Fixed the quoted healthy example (line ~114) to include the pending-bytes
suffix the code now appends. Previously the prose only documented the
probe-failed shape, missing the PR's own headline cases (never-run).
Proof: grep -n 'never-run' proxmox-monitor.prose.md now returns two lines
(lines 106 and 108), documenting both never-run variants.
Branch: fix/pbs-gc-liveness-signal-20260919
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
What
Adds a liveness leg to scripts/proxmox-monitor.sh that reads the local PBS datastore's garbage-collection state from inside CT 107 (where proxmox-backup-manager lives) and FAILS when storepve-datastore's last-run-endtime is older than 48h, reporting the age and pending bytes.
Why
The nightly GC job (/etc/cron.d/pbs-gc -> /usr/local/bin/pbs-gc.sh) had never worked: it called proxmox-backup-manager on the storepve HOST, where only proxmox-backup-client exists, so every run died 'command not found'. The host script is fixed and verified (a real run completed OK), and this leg is the signal that catches a silently dead GC.
To follow on this branch
A proxmox-monitor.prose.md section documenting the schedule and this check (superseding the separate docs PR #116, whose text states the schedule in the wrong timezone - it is 20:00 local = 00:00 UTC, not 20:00 UTC).
Firstmate runs the pre-delivery review and the merge.
(a) Probe-failure detection: now treats empty OR unparseable JSON as probe-failed, not never-run. This prevents 'command not found' outputs from being rendered as service verdicts. (b) pending-bytes: now extracted from JSON and reported in stale verdict. (c) Null endtime: use .get() with explicit None check, not 0 fallback. null values now correctly trigger never-run verdict instead of arithmetic crash (set -u). (d) Tests: Added 14-assertion stub-driven suite covering: healthy, stale (>48h), probe-failed (empty and unparseable), null endtime, datastore absent. Each test stubs ssh/curl to verify exact behavior against the pre-fix head. Branch: fix/pbs-gc-liveness-signal-20260919