fix(disk-gc): host filesystem bands, named volumes, report-only, and state-change alerts #105
@@ -43,7 +43,7 @@ Docker hosts get special attention:
|
||||
|
||||
> **Note:** CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
|
||||
|
||||
## Threat Levels
|
||||
## Threat Levels (GUEST filesystems)
|
||||
|
||||
| Level | Threshold | Response | Escalation |
|
||||
|-------|-----------|----------|------------|
|
||||
@@ -53,6 +53,39 @@ Docker hosts get special attention:
|
||||
| **CRITICAL** | ≥ 95% | Aggressive GC + emergency cleanup | Zulip + relay to Kwame |
|
||||
| **FULL** | 100% (df shows 100%) | Stop writes, manual intervention | Call/Telegram Kwame |
|
||||
|
||||
## Host Filesystem Thresholds (HOST nodes — separate bands from guest bands)
|
||||
|
||||
Host filesystems have their own risk profile and their own bands. A host root near full is a real risk (backup staging writes to it, thin-pool metadata pressure), a media volume near full is a capacity decision for the owner, and the backup datastore near full breaks backups. These are **report-only** — no automatic deletion of media or datastore content ever.
|
||||
|
||||
| Level | Threshold | Response | Escalation |
|
||||
|-------|-----------|----------|------------|
|
||||
| **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None |
|
||||
| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) |
|
||||
| **HOST-RED** | 95% | Name the volume + % + absolute free space; flag for immediate owner attention | Zulip DM + channel alert (state-change only) |
|
||||
|
||||
**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.
|
||||
|
||||
**State lives in a small JSON state file:** `state/host-disk-bands.json`, keyed by `host/volume` → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.)
|
||||
|
||||
**Volume naming rule:** Every host line MUST name the volume and what lives on it. Example output:
|
||||
|
||||
```
|
||||
storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED
|
||||
storepve /media/reanim (media): 86% (797G/932G, 135G free) -> HOST-WARN
|
||||
storepve /media/mediastore (media): 77% (5.3T/7.3T, 1.6T free) -> GREEN
|
||||
storepve /dev/mapper/pve-root (host-root): 81% (73G/94G, 17G free) -> GREEN
|
||||
storepve tank (pbs-datastore): 1% (128K/12T, 12T free) -> GREEN
|
||||
```
|
||||
|
||||
**First-run behavior:** When the state file does not yet exist (first scan), the current band of every volume is recorded as the baseline WITHOUT alerting — a first run would otherwise DM every already-elevated volume at once. From the second run onward, transitions alert.
|
||||
|
||||
**Action classes by volume type:**
|
||||
- **host-root**: near full = real risk (backup staging, thin-pool metadata, journald). HOST-AMBER or above → owner must investigate.
|
||||
- **media** (/media/*): near full = capacity decision for the owner. HOST-AMBER or above → report only, never auto-delete.
|
||||
- **pbs-datastore** (tank, ZFS): near full = breaks Proxmox Backup Server. HOST-RED → escalate immediately.
|
||||
|
||||
**Justification (measured 2026-09-15, firstmate):** storepve (192.168.68.6) shows /media/easystore2 at 96%, /media/reanim at 86%, /dev/mapper/pve-root at 81%, and tank at 1%. The previous scan reported "0/19 guests >80%, no GC action needed" while easystore2 sat at 96% — the host filesystems were printed but never banded and never acted on. Two incidents this weekend (amdpve host root filled during container backup staging, acerpve root went read-only when its thin pool errored) showed the host filesystem is the thing that breaks, not the guest's.
|
||||
|
||||
## Requires
|
||||
|
||||
- SSH access to all Docker hosts (matrix from infrastructure-control pattern)
|
||||
@@ -128,6 +161,10 @@ from the `report_only_guests` YAML block above.
|
||||
|
||||
## Execution
|
||||
|
||||
### Host filesystems: report-only, NEVER auto-delete
|
||||
|
||||
Host filesystems are always report-only — no automatic deletion of media or datastore content. A host root near full requires investigation by the owner, but the executor must never delete content on a host filesystem. This restriction is absolute.
|
||||
|
||||
### Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)
|
||||
|
||||
**CT 111 / hostname `tdunna` / 192.168.68.129 is DETECT-AND-REPORT-ONLY.** It belongs to Theo.
|
||||
@@ -243,6 +280,12 @@ done
|
||||
|
||||
## Alert Templates
|
||||
|
||||
### HOST-WARN / HOST-AMBER / HOST-RED (host filesystems, report-only)
|
||||
```
|
||||
⚠️ Host Disk — {hostname} {volume} ({volume_type}): {pct}% ({used}/{total}, {free} free) -> {level}
|
||||
Action: {volume_type-specific action}
|
||||
```
|
||||
|
||||
### AMBER (75-84%)
|
||||
```
|
||||
⚠️ Disk GC — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
|
||||
|
||||
Reference in New Issue
Block a user