fix(disk-gc): host filesystem bands, named volumes, report-only, and state-change alerts #105

Merged
abiba-bot merged 2 commits from fix/host-filesystem-thresholds-20260915 into master 2026-09-15 14:40:22 +00:00
Owner

The disk-gc contract banded GUEST filesystems and merely printed HOST filesystems. A host volume at 96% produced no band, no escalation and "no GC action needed" - while this weekend's two incidents were both host-side (amdpve's host root filled during container-backup staging; acerpve's root remounted read-only when its thin pool errored).

Change (disk-gc-threat-response.prose.md, +44/-1)

  • Separate host bands, distinct from the guest bands: HOST-WARN 85%, HOST-AMBER 90%, HOST-RED 95%.
  • Every host line names the volume AND what lives on it, with the percentage AND the absolute free space: storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED. 96% of 3.7T is not the same risk as 96% of 20G.
  • Action classes by volume type: host-root near full is a real risk (backup staging, thin-pool metadata); media near full is a capacity decision for the owner; pbs-datastore near full breaks backups.
  • ABSOLUTE report-only restriction: the executor must never delete media or datastore content - a host root near full requires owner investigation, never automated deletion.
  • State-change alerts, not per-run: a volume DMs ONCE when it enters a higher band and ONCE on recovery; while it stays in the same band it is scan-output-only. State lives in state/host-disk-bands.json (host/volume -> last band), written after every scan. Rationale: easystore2 at 96% would otherwise re-DM the owner on every 6-hour scan until someone acted, which is how a real warning becomes unread noise.

Verification performed

  • Measured on storepve (192.168.68.6) by firstmate: /media/easystore2 3.5T/3.7T (96%, 177G free), /media/reanim 797G/932G (86%), /dev/mapper/pve-root 73G/94G (81%), tank 1% (the PBS datastore). The quoted example output is those real figures.
  • The 96% volume is MEDIA, not the backup store; the datastore sits at 1% on a 12T pool - stated plainly so the band is not misread as a backup emergency.

Reviewers: (1) confirm each threshold and each number in the example is the measured value; (2) confirm the report-only restriction is absolute and that no path in the contract can delete host/media/datastore content; (3) judge whether baselining on the first run can hide a pre-existing critical volume - as written, the first scan records every band silently, so a volume already at RED raises no alert until something changes; say whether that is acceptable or whether the first run should emit a single one-line summary of anything already at AMBER/RED. Do NOT merge.

The disk-gc contract banded GUEST filesystems and merely printed HOST filesystems. A host volume at 96% produced no band, no escalation and "no GC action needed" - while this weekend's two incidents were both host-side (amdpve's host root filled during container-backup staging; acerpve's root remounted read-only when its thin pool errored). ## Change (`disk-gc-threat-response.prose.md`, +44/-1) - **Separate host bands**, distinct from the guest bands: HOST-WARN 85%, HOST-AMBER 90%, HOST-RED 95%. - **Every host line names the volume AND what lives on it**, with the percentage AND the absolute free space: `storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED`. 96% of 3.7T is not the same risk as 96% of 20G. - **Action classes by volume type**: `host-root` near full is a real risk (backup staging, thin-pool metadata); `media` near full is a capacity decision for the owner; `pbs-datastore` near full breaks backups. - **ABSOLUTE report-only restriction**: the executor must never delete media or datastore content - a host root near full requires owner investigation, never automated deletion. - **State-change alerts, not per-run**: a volume DMs ONCE when it enters a higher band and ONCE on recovery; while it stays in the same band it is scan-output-only. State lives in `state/host-disk-bands.json` (host/volume -> last band), written after every scan. Rationale: easystore2 at 96% would otherwise re-DM the owner on every 6-hour scan until someone acted, which is how a real warning becomes unread noise. ## Verification performed - Measured on storepve (192.168.68.6) by firstmate: `/media/easystore2` 3.5T/3.7T (96%, 177G free), `/media/reanim` 797G/932G (86%), `/dev/mapper/pve-root` 73G/94G (81%), `tank` 1% (the PBS datastore). The quoted example output is those real figures. - The 96% volume is MEDIA, not the backup store; the datastore sits at 1% on a 12T pool - stated plainly so the band is not misread as a backup emergency. Reviewers: (1) confirm each threshold and each number in the example is the measured value; (2) confirm the report-only restriction is absolute and that no path in the contract can delete host/media/datastore content; (3) **judge whether baselining on the first run can hide a pre-existing critical volume** - as written, the first scan records every band silently, so a volume already at RED raises no alert until something changes; say whether that is acceptable or whether the first run should emit a single one-line summary of anything already at AMBER/RED. Do NOT merge.
abiba-bot added 2 commits 2026-09-15 14:28:53 +00:00
Add separate threat bands for HOST filesystems (distinct from guest bands):
- HOST-WARN at 85%: name volume + % + absolute free space in scan output
- HOST-AMBER at 90%: flag for owner attention, Zulip DM
- HOST-RED at 95%: flag for immediate owner attention, Zulip DM + channel alert

Volume naming rule: every host line MUST name the volume and what lives on it.
Action classes by volume type:
- host-root: near full = real risk (backup staging, thin-pool metadata)
- media (/media/*): near full = capacity decision for owner, never auto-delete
- pbs-datastore (tank): near full = breaks Proxmox Backup Server

Report-only restriction: no automatic deletion of media or datastore content ever.

Justification (measured 2026-09-15): storepve /media/easystore2 at 96% was
reported but never banded or acted on. Two incidents this weekend showed the
host filesystem is the thing that breaks, not the guest's.

Added HOST-WARN/AMBER/RED alert templates.
Added report-only execution rule for host filesystems.
fix: make host escalations state-change driven
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
b9b1712ac6
A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER,
AMBER->RED) and ONCE when it drops back down (recovery notice). While it stays
in the same band, it is reported in the scan output only — no DM, no channel
alert. This stops the same 96% easystore2 from re-DMing the owner on every 6h
scan.

State lives in a small JSON state file (state/host-disk-bands.json), keyed by
host/volume -> last-seen band. The scanner reads the prior band, compares to the
current band, and DMs only on a transition; the state file is written after
every scan. Chosen over a periodic digest because the scan already runs every
6h and a transition is genuinely new, actionable state.

First-run behavior: when the state file does not yet exist, the current band of
every volume is recorded as baseline WITHOUT alerting — a first run would
otherwise DM every already-elevated volume at once.

Report-only restriction and volume-naming output kept exactly as-is.
abiba-bot merged commit 6c616a9e58 into master 2026-09-15 14:40:22 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SyslogSolution/prose-contracts#105