docs(proxmox): Add PBS GC documentation to proxmox-monitor.prose.md
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 4s

Document the PBS GC schedule (00:00 UTC, not 20:00 UTC as PR #116 said),
what actually runs (pbs-gc.sh -> pct exec 107 -- proxmox-backup-manager),
datastore location (CT 107's /mnt/pbs-backup on /tank/pbs-backup, NOT
/media/easystore2), and the new liveness check (48h threshold, reports
age in hours, explicit healthy line).

Branch: fix/pbs-gc-liveness-signal-20260919
This commit is contained in:
root
2026-09-19 02:12:17 +00:00
parent 8ae4b59150
commit da8f5f43c9
+24
View File
@@ -89,6 +89,30 @@ agent: abiba
| minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
## PBS GC (Proxmox Backup Server)
### Schedule
Cron `0 20 * * *` on the **storepve HOST** (192.168.68.6) = 20:00 America/New_York local = **00:00 UTC**.
**NOTE**: The closed PR #116 said "20:00 UTC" — this is WRONG by four hours. Do not copy it.
### What Actually Runs
Host script `/usr/local/bin/pbs-gc.sh` runs `pct exec 107 -- proxmox-backup-manager garbage-collection start storepve-datastore`.
**IMPORTANT**: The tool `proxmox-backup-manager` exists only inside CT 107 (where the PBS server runs). The storepve host has only `proxmox-backup-client`. This is why the job had never worked before 2026-09-19 00:00 UTC.
### Datastore Location
- **Datastore**: CT 107's `/mnt/pbs-backup` on the storepve ZFS dataset `/tank/pbs-backup` (pool `tank`, ~11T free)
- **NOT** `/media/easystore2` (media library, 3.7T, 96% used — separate volume)
### Liveness Check
The new `proxmox-monitor.sh` leg checks storepve-datastore GC health:
- Reads GC state from CT 107: `pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json`
- **FAILS** if `last-run-endtime` is older than 48 hours
- Reports age in hours and pending-bytes
- A failed probe (000/timeout) is reported as `probe-failed: storepve:192.168.68.6 (expected JSON, got 000)` — **never** rendered as "GC stale"
- The healthy branch prints an explicit healthy line: `✅ PBS GC: healthy (last run Xh ago)`
## Operations
### view-dashboards