diff --git a/disk-gc-threat-response.prose.md b/disk-gc-threat-response.prose.md index b374e59..77dbf11 100644 --- a/disk-gc-threat-response.prose.md +++ b/disk-gc-threat-response.prose.md @@ -84,6 +84,32 @@ and escalation trail. - May also be invoked manually: `prose run disk-gc-threat-response` - Threat-driven: if Amber/Red/Critical detected, immediate GC phase activates +## Scanner: scripts/disk-gc-scan.py + +The fleet scan is executed by `scripts/disk-gc-scan.py`, which makes reachability +verdicts deterministic: + +1. **Retry on failure:** Each probe retries once before declaring a guest unreachable. +2. **Named probe target:** Every rendered line names the guest, CT id, node, and + access method actually used. +3. **Failure kind printed:** An unreachable guest is reported with its failure kind + (timeout, ssh-auth, no-route, conn-refused, ssh-exit-N) — never as a bare + "unreachable" verdict. +4. **Per-guest access method:** The correct access path is selected from a per-guest + map so the wrong path cannot be picked by an executor improvising: + - CT 105 (kagentz) = `ssh root@kagentz` (NOT `pct exec 105` — pct exec sees + loop0/59G instead of the real 99G filesystem) + - CT 109 (docker-vm) = `ssh root@192.168.68.7` (NOT `pct exec` — it's a KVM VM) + - All other CTs = `pct-run ` (which uses `pct exec` via SSH to the node) +5. **Every figure traces to a named probe:** The scan output prints the exact command + that produced each disk figure, so two different guests can never render + identical numbers without the probe commands proving it. + +Run: `python3 scripts/disk-gc-scan.py` (or `--json` for machine-readable output). + +The scan feeds into `scripts/disk-gc-plan.py`, which applies the report-only gate +from the `report_only_guests` YAML block above. + ## Shape - `self`: scan all CTs via Proxmox API + SSH exec, trigger GC, alert