The disk-gc contract banded GUEST filesystems and merely printed HOST filesystems. A host volume at 96% produced no band, no escalation and "no GC action needed" - while this weekend's two incidents were both host-side (amdpve's host root filled during container-backup staging; acerpve's root remounted read-only when its thin pool errored).
Change (disk-gc-threat-response.prose.md, +44/-1)
Separate host bands, distinct from the guest bands: HOST-WARN 85%, HOST-AMBER 90%, HOST-RED 95%.
Every host line names the volume AND what lives on it, with the percentage AND the absolute free space: storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED. 96% of 3.7T is not the same risk as 96% of 20G.
Action classes by volume type: host-root near full is a real risk (backup staging, thin-pool metadata); media near full is a capacity decision for the owner; pbs-datastore near full breaks backups.
ABSOLUTE report-only restriction: the executor must never delete media or datastore content - a host root near full requires owner investigation, never automated deletion.
State-change alerts, not per-run: a volume DMs ONCE when it enters a higher band and ONCE on recovery; while it stays in the same band it is scan-output-only. State lives in state/host-disk-bands.json (host/volume -> last band), written after every scan. Rationale: easystore2 at 96% would otherwise re-DM the owner on every 6-hour scan until someone acted, which is how a real warning becomes unread noise.
Verification performed
Measured on storepve (192.168.68.6) by firstmate: /media/easystore2 3.5T/3.7T (96%, 177G free), /media/reanim 797G/932G (86%), /dev/mapper/pve-root 73G/94G (81%), tank 1% (the PBS datastore). The quoted example output is those real figures.
The 96% volume is MEDIA, not the backup store; the datastore sits at 1% on a 12T pool - stated plainly so the band is not misread as a backup emergency.
Reviewers: (1) confirm each threshold and each number in the example is the measured value; (2) confirm the report-only restriction is absolute and that no path in the contract can delete host/media/datastore content; (3) judge whether baselining on the first run can hide a pre-existing critical volume - as written, the first scan records every band silently, so a volume already at RED raises no alert until something changes; say whether that is acceptable or whether the first run should emit a single one-line summary of anything already at AMBER/RED. Do NOT merge.
The disk-gc contract banded GUEST filesystems and merely printed HOST filesystems. A host volume at 96% produced no band, no escalation and "no GC action needed" - while this weekend's two incidents were both host-side (amdpve's host root filled during container-backup staging; acerpve's root remounted read-only when its thin pool errored).
## Change (`disk-gc-threat-response.prose.md`, +44/-1)
- **Separate host bands**, distinct from the guest bands: HOST-WARN 85%, HOST-AMBER 90%, HOST-RED 95%.
- **Every host line names the volume AND what lives on it**, with the percentage AND the absolute free space: `storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED`. 96% of 3.7T is not the same risk as 96% of 20G.
- **Action classes by volume type**: `host-root` near full is a real risk (backup staging, thin-pool metadata); `media` near full is a capacity decision for the owner; `pbs-datastore` near full breaks backups.
- **ABSOLUTE report-only restriction**: the executor must never delete media or datastore content - a host root near full requires owner investigation, never automated deletion.
- **State-change alerts, not per-run**: a volume DMs ONCE when it enters a higher band and ONCE on recovery; while it stays in the same band it is scan-output-only. State lives in `state/host-disk-bands.json` (host/volume -> last band), written after every scan. Rationale: easystore2 at 96% would otherwise re-DM the owner on every 6-hour scan until someone acted, which is how a real warning becomes unread noise.
## Verification performed
- Measured on storepve (192.168.68.6) by firstmate: `/media/easystore2` 3.5T/3.7T (96%, 177G free), `/media/reanim` 797G/932G (86%), `/dev/mapper/pve-root` 73G/94G (81%), `tank` 1% (the PBS datastore). The quoted example output is those real figures.
- The 96% volume is MEDIA, not the backup store; the datastore sits at 1% on a 12T pool - stated plainly so the band is not misread as a backup emergency.
Reviewers: (1) confirm each threshold and each number in the example is the measured value; (2) confirm the report-only restriction is absolute and that no path in the contract can delete host/media/datastore content; (3) **judge whether baselining on the first run can hide a pre-existing critical volume** - as written, the first scan records every band silently, so a volume already at RED raises no alert until something changes; say whether that is acceptable or whether the first run should emit a single one-line summary of anything already at AMBER/RED. Do NOT merge.
Add separate threat bands for HOST filesystems (distinct from guest bands):
- HOST-WARN at 85%: name volume + % + absolute free space in scan output
- HOST-AMBER at 90%: flag for owner attention, Zulip DM
- HOST-RED at 95%: flag for immediate owner attention, Zulip DM + channel alert
Volume naming rule: every host line MUST name the volume and what lives on it.
Action classes by volume type:
- host-root: near full = real risk (backup staging, thin-pool metadata)
- media (/media/*): near full = capacity decision for owner, never auto-delete
- pbs-datastore (tank): near full = breaks Proxmox Backup Server
Report-only restriction: no automatic deletion of media or datastore content ever.
Justification (measured 2026-09-15): storepve /media/easystore2 at 96% was
reported but never banded or acted on. Two incidents this weekend showed the
host filesystem is the thing that breaks, not the guest's.
Added HOST-WARN/AMBER/RED alert templates.
Added report-only execution rule for host filesystems.
A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER,
AMBER->RED) and ONCE when it drops back down (recovery notice). While it stays
in the same band, it is reported in the scan output only — no DM, no channel
alert. This stops the same 96% easystore2 from re-DMing the owner on every 6h
scan.
State lives in a small JSON state file (state/host-disk-bands.json), keyed by
host/volume -> last-seen band. The scanner reads the prior band, compares to the
current band, and DMs only on a transition; the state file is written after
every scan. Chosen over a periodic digest because the scan already runs every
6h and a transition is genuinely new, actionable state.
First-run behavior: when the state file does not yet exist, the current band of
every volume is recorded as baseline WITHOUT alerting — a first run would
otherwise DM every already-elevated volume at once.
Report-only restriction and volume-naming output kept exactly as-is.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
The disk-gc contract banded GUEST filesystems and merely printed HOST filesystems. A host volume at 96% produced no band, no escalation and "no GC action needed" - while this weekend's two incidents were both host-side (amdpve's host root filled during container-backup staging; acerpve's root remounted read-only when its thin pool errored).
Change (
disk-gc-threat-response.prose.md, +44/-1)storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED. 96% of 3.7T is not the same risk as 96% of 20G.host-rootnear full is a real risk (backup staging, thin-pool metadata);medianear full is a capacity decision for the owner;pbs-datastorenear full breaks backups.state/host-disk-bands.json(host/volume -> last band), written after every scan. Rationale: easystore2 at 96% would otherwise re-DM the owner on every 6-hour scan until someone acted, which is how a real warning becomes unread noise.Verification performed
/media/easystore23.5T/3.7T (96%, 177G free),/media/reanim797G/932G (86%),/dev/mapper/pve-root73G/94G (81%),tank1% (the PBS datastore). The quoted example output is those real figures.Reviewers: (1) confirm each threshold and each number in the example is the measured value; (2) confirm the report-only restriction is absolute and that no path in the contract can delete host/media/datastore content; (3) judge whether baselining on the first run can hide a pre-existing critical volume - as written, the first scan records every band silently, so a volume already at RED raises no alert until something changes; say whether that is acceptable or whether the first run should emit a single one-line summary of anything already at AMBER/RED. Do NOT merge.