fix(disk-gc): actually write the host-band state file so escalations and recoveries can fire #106

Merged
abiba-bot merged 2 commits from fix/host-disk-band-state-file-20260916 into master 2026-09-16 01:00:29 +00:00
Owner

PR #105 added host filesystem bands and promised "alert once when a volume gets worse, once when it recovers". The bands displayed correctly, but the piece that distinguishes worse from same — a small state file remembering the previous band — was described in the contract and never written. Every run therefore reported the first-run case and stayed silent, so a volume moving from 96% to 99% would have passed without a word. This fixes it and, importantly, proves the behaviour rather than describing it.

Changes (3 files)

  • scripts/disk-gc-scan.py: real state I/O — reads the prior band per host/volume, compares against the current band, and reports transitions in both directions: escalation (GREEN -> HOST-WARN -> HOST-AMBER -> HOST-RED) and recovery. The state path resolves absolutely from the script's location, not from the working directory, so two execution contexts cannot write to two different places. A state file that cannot be read or written now says so loudly on stderr instead of silently degrading to "always first run".
  • .gitignore: state/host-disk-bands.json is runtime state. It is now untracked (and the mistakenly committed copy removed), because every scan rewrites it — a tracked copy would leave every executor's clone permanently dirty, produce conflicts between clones, and could present a stale baseline as current.
  • disk-gc-threat-response.prose.md: documents the absolute path, the gitignored status, and the create-on-first-run behaviour.

Evidence (produced on the branch, in this order)

  1. First run — (first run — recording baseline, no alerts), state file written with 16 volumes (including storepve//media/easystore2=HOST-RED, storepve//media/reanim=HOST-WARN).
  2. Second run, unchanged — (no band changes since last scan), with no repeating "first run" line.
  3. Forced escalation — state hand-edited to GREEN for easystore2, scan reported ⚠️ storepve//media/easystore2: GREEN → HOST-RED (ESCALATION).
  4. Run again unchanged — silent. Alerts once, not every run.
  5. Forced recovery — state hand-edited to HOST-AMBER for reanim, scan reported ✅ storepve//media/reanim: HOST-AMBER → HOST-WARN (RECOVERY).

Reviewers: confirm the state path cannot depend on CWD; confirm both transition directions are exercised and that an unchanged run is silent (that is the whole point — a control that alerts every run is noise, and one that never alerts is decoration); confirm the file is untracked and that no runtime state is committed; and confirm the failure path for an unwritable state directory is loud rather than silent.

PR #105 added host filesystem bands and promised "alert once when a volume gets worse, once when it recovers". The bands displayed correctly, but the piece that distinguishes *worse* from *same* — a small state file remembering the previous band — was described in the contract and never written. Every run therefore reported the first-run case and stayed silent, so a volume moving from 96% to 99% would have passed without a word. This fixes it and, importantly, proves the behaviour rather than describing it. ## Changes (3 files) - **`scripts/disk-gc-scan.py`**: real state I/O — reads the prior band per `host/volume`, compares against the current band, and reports transitions in both directions: escalation (`GREEN -> HOST-WARN -> HOST-AMBER -> HOST-RED`) and **recovery**. The state path resolves **absolutely from the script's location**, not from the working directory, so two execution contexts cannot write to two different places. A state file that cannot be read or written now says so loudly on stderr instead of silently degrading to "always first run". - **`.gitignore`**: `state/host-disk-bands.json` is runtime state. It is now untracked (and the mistakenly committed copy removed), because every scan rewrites it — a tracked copy would leave every executor's clone permanently dirty, produce conflicts between clones, and could present a stale baseline as current. - **`disk-gc-threat-response.prose.md`**: documents the absolute path, the gitignored status, and the create-on-first-run behaviour. ## Evidence (produced on the branch, in this order) 1. **First run** — `(first run — recording baseline, no alerts)`, state file written with 16 volumes (including `storepve//media/easystore2=HOST-RED`, `storepve//media/reanim=HOST-WARN`). 2. **Second run, unchanged** — `(no band changes since last scan)`, with no repeating "first run" line. 3. **Forced escalation** — state hand-edited to `GREEN` for easystore2, scan reported `⚠️ storepve//media/easystore2: GREEN → HOST-RED (ESCALATION)`. 4. **Run again unchanged** — silent. Alerts once, not every run. 5. **Forced recovery** — state hand-edited to `HOST-AMBER` for reanim, scan reported `✅ storepve//media/reanim: HOST-AMBER → HOST-WARN (RECOVERY)`. Reviewers: confirm the state path cannot depend on CWD; confirm both transition directions are exercised and that an unchanged run is silent (that is the whole point — a control that alerts every run is noise, and one that never alerts is decoration); confirm the file is untracked and that no runtime state is committed; and confirm the failure path for an unwritable state directory is loud rather than silent.
abiba-bot added 2 commits 2026-09-16 00:48:59 +00:00
The contract said the state file was written after every scan, but the script
had no state-file logic at all. This PR adds:

1. Host filesystem scanning (probe_host_filesystems) - probes df on all PVE nodes
2. State file I/O (read_state_file/write_state_file) - absolute path from script location
3. Band classification (classify_band) - HOST-WARN/AMBER/RED thresholds
4. Transition detection (detect_transitions) - alerts on escalation/recovery
5. CLI flags (--hosts-only, --guests-only) to control which parts run

The contract now specifies the state file path resolves from the script's own
location (not CWD-relative), so two different execution contexts cannot write
to two different places.
fix: untrack host-disk-bands.json and document gitignored status
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
39209c7ac9
The state file is runtime state (rewritten every scan), so tracking it in git
means:
- every executor's clone becomes permanently dirty after one run
- a scan in one clone produces a merge conflict with a scan in another
- the committed baseline can be stale in a way nobody notices

Added to .gitignore and removed from the index. Contract updated to say
'the state file lives at <abs path> and is gitignored runtime state - the
scanner creates it on first run'.
abiba-bot merged commit dae8d14880 into master 2026-09-16 01:00:29 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SyslogSolution/prose-contracts#106