diff --git a/disk-gc-threat-response.prose.md b/disk-gc-threat-response.prose.md index 584422c..f5d72cf 100644 --- a/disk-gc-threat-response.prose.md +++ b/disk-gc-threat-response.prose.md @@ -60,8 +60,12 @@ Host filesystems have their own risk profile and their own bands. A host root ne | Level | Threshold | Response | Escalation | |-------|-----------|----------|------------| | **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None | -| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner | -| **HOST-RED** | 95% | Name the volume + % + absolute free space; flag for immediate owner attention | Zulip DM + channel alert | +| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) | +| **HOST-RED** | 95% | Name the volume + % + absolute free space; flag for immediate owner attention | Zulip DM + channel alert (state-change only) | + +**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise. + +**State lives in a small JSON state file:** `state/host-disk-bands.json`, keyed by `host/volume` → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.) **Volume naming rule:** Every host line MUST name the volume and what lives on it. Example output: @@ -73,6 +77,8 @@ storepve /dev/mapper/pve-root (host-root): 81% (73G/94G, 17G free) -> GREEN storepve tank (pbs-datastore): 1% (128K/12T, 12T free) -> GREEN ``` +**First-run behavior:** When the state file does not yet exist (first scan), the current band of every volume is recorded as the baseline WITHOUT alerting — a first run would otherwise DM every already-elevated volume at once. From the second run onward, transitions alert. + **Action classes by volume type:** - **host-root**: near full = real risk (backup staging, thin-pool metadata, journald). HOST-AMBER or above → owner must investigate. - **media** (/media/*): near full = capacity decision for the owner. HOST-AMBER or above → report only, never auto-delete.