From b9b1712ac628c2cdcc0b6372cb2841ee8a4e55fd Mon Sep 17 00:00:00 2001 From: root Date: Tue, 15 Sep 2026 14:25:49 +0000 Subject: [PATCH] fix: make host escalations state-change driven MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (recovery notice). While it stays in the same band, it is reported in the scan output only — no DM, no channel alert. This stops the same 96% easystore2 from re-DMing the owner on every 6h scan. State lives in a small JSON state file (state/host-disk-bands.json), keyed by host/volume -> last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition; the state file is written after every scan. Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state. First-run behavior: when the state file does not yet exist, the current band of every volume is recorded as baseline WITHOUT alerting — a first run would otherwise DM every already-elevated volume at once. Report-only restriction and volume-naming output kept exactly as-is. --- disk-gc-threat-response.prose.md | 10 ++++++++-- 1 file changed, 8 insertions(+), 2 deletions(-) diff --git a/disk-gc-threat-response.prose.md b/disk-gc-threat-response.prose.md index 584422c..f5d72cf 100644 --- a/disk-gc-threat-response.prose.md +++ b/disk-gc-threat-response.prose.md @@ -60,8 +60,12 @@ Host filesystems have their own risk profile and their own bands. A host root ne | Level | Threshold | Response | Escalation | |-------|-----------|----------|------------| | **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None | -| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner | -| **HOST-RED** | 95% | Name the volume + % + absolute free space; flag for immediate owner attention | Zulip DM + channel alert | +| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) | +| **HOST-RED** | 95% | Name the volume + % + absolute free space; flag for immediate owner attention | Zulip DM + channel alert (state-change only) | + +**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise. + +**State lives in a small JSON state file:** `state/host-disk-bands.json`, keyed by `host/volume` → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.) **Volume naming rule:** Every host line MUST name the volume and what lives on it. Example output: @@ -73,6 +77,8 @@ storepve /dev/mapper/pve-root (host-root): 81% (73G/94G, 17G free) -> GREEN storepve tank (pbs-datastore): 1% (128K/12T, 12T free) -> GREEN ``` +**First-run behavior:** When the state file does not yet exist (first scan), the current band of every volume is recorded as the baseline WITHOUT alerting — a first run would otherwise DM every already-elevated volume at once. From the second run onward, transitions alert. + **Action classes by volume type:** - **host-root**: near full = real risk (backup staging, thin-pool metadata, journald). HOST-AMBER or above → owner must investigate. - **media** (/media/*): near full = capacity decision for the owner. HOST-AMBER or above → report only, never auto-delete.