diff --git a/proxmox-monitor.prose.md b/proxmox-monitor.prose.md index a7904e1..5fbc6f8 100644 --- a/proxmox-monitor.prose.md +++ b/proxmox-monitor.prose.md @@ -89,6 +89,30 @@ agent: abiba | minipve | 192.168.68.12 | PVE | | amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) | +## PBS GC (Proxmox Backup Server) + +### Schedule +Cron `0 20 * * *` on the **storepve HOST** (192.168.68.6) = 20:00 America/New_York local = **00:00 UTC**. + +**NOTE**: The closed PR #116 said "20:00 UTC" — this is WRONG by four hours. Do not copy it. + +### What Actually Runs +Host script `/usr/local/bin/pbs-gc.sh` runs `pct exec 107 -- proxmox-backup-manager garbage-collection start storepve-datastore`. + +**IMPORTANT**: The tool `proxmox-backup-manager` exists only inside CT 107 (where the PBS server runs). The storepve host has only `proxmox-backup-client`. This is why the job had never worked before 2026-09-19 00:00 UTC. + +### Datastore Location +- **Datastore**: CT 107's `/mnt/pbs-backup` on the storepve ZFS dataset `/tank/pbs-backup` (pool `tank`, ~11T free) +- **NOT** `/media/easystore2` (media library, 3.7T, 96% used — separate volume) + +### Liveness Check +The new `proxmox-monitor.sh` leg checks storepve-datastore GC health: +- Reads GC state from CT 107: `pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json` +- **FAILS** if `last-run-endtime` is older than 48 hours +- Reports age in hours and pending-bytes +- A failed probe (000/timeout) is reported as `probe-failed: storepve:192.168.68.6 (expected JSON, got 000)` — **never** rendered as "GC stale" +- The healthy branch prints an explicit healthy line: `✅ PBS GC: healthy (last run Xh ago)` + ## Operations ### view-dashboards