fix(disk-gc): deterministic reachability verdict and correctly labelled disk figures #91

Merged
abiba-bot merged 4 commits from fix/disk-gc-probe-deterministic-20260914 into master 2026-09-14 14:43:44 +00:00
Owner

Makes the disk-gc scan's verdict trustworthy. It previously reported running guests as unreachable and rendered unusable numbers; both are fixed and independently verified.

The reported defect

On 2026-09-13 the scan reported kagentz (CT 105): unreachable and docker-vm (CT 109): unreachable while both were up and reachable, and in a second run rendered byte-identical figures for two different guests (abiba (CT 100): 34% (13G/40G) and syslog-api (CT 116): 34% (13G/40G)).

The fix

  • scripts/disk-gc-scan.py (new, 304 lines): the fleet disk probe with an explicit per-guest access map - CT 105 via ssh root@kagentz (not pct exec 105, which sees loop0/59G instead of the real 99G filesystem), VM 109 via ssh root@192.168.68.7, other guests via pct-run. Retry once before declaring anything unreachable, and every line names the guest, the host, the access method and the exact probe command.
  • disk-gc-threat-response.prose.md: documents the scanner and its access map (+26 lines).

Verification performed by firstmate before this PR (extracted from the branch, run unmodified)

  1. CWD independence - the first revision resolved its helper relative to the CALLER's directory, so run from anywhere but the repo root it declared 14 false UNREACHABLE (ssh-exit-127). Now the helper resolves from the script's own location: running it from / prints bash /tmp/dgtest/scripts/pct-run.sh 104 … and produces the same verdicts as running it from the containing directory.
  2. No scanner failure rendered as a guest verdict - with the helper genuinely absent it prints a single SCANNER ERROR: helper not found: <abs path> line and zero UNREACHABLE verdicts.
  3. Figure mapping is now column-for-column correct against the authority. Raw df -P / on kagentz gives 102996088 9301244 89459408 10% (total / used / avail in 1K blocks); the scan renders 10% (8.9G used of 98.2G total, 85.3G free). The previous revision printed the pair as (102996088/9301220) - total before used, unlabelled, so the big number read as usage, and the used value did not even match the authority byte-for-byte.
  4. Full run: every covered guest resolves with a real figure and the probe that produced it (docker-vm 11% / 15.9G used of 157.3G total).

Reviewers: re-run it from two different working directories and confirm identical output; confirm a misconfigured helper yields SCANNER ERROR rather than a guest verdict; and spot-check two guests' figures against a raw df -P / on the guest itself.

Makes the disk-gc scan's verdict trustworthy. It previously reported running guests as unreachable and rendered unusable numbers; both are fixed and independently verified. ## The reported defect On 2026-09-13 the scan reported `kagentz (CT 105): unreachable` and `docker-vm (CT 109): unreachable` while both were up and reachable, and in a second run rendered byte-identical figures for two different guests (`abiba (CT 100): 34% (13G/40G)` and `syslog-api (CT 116): 34% (13G/40G)`). ## The fix - **`scripts/disk-gc-scan.py`** (new, 304 lines): the fleet disk probe with an explicit per-guest access map - CT 105 via `ssh root@kagentz` (not `pct exec 105`, which sees loop0/59G instead of the real 99G filesystem), VM 109 via `ssh root@192.168.68.7`, other guests via `pct-run`. Retry once before declaring anything unreachable, and every line names the guest, the host, the access method and the exact probe command. - **`disk-gc-threat-response.prose.md`**: documents the scanner and its access map (+26 lines). ## Verification performed by firstmate before this PR (extracted from the branch, run unmodified) 1. **CWD independence** - the first revision resolved its helper relative to the CALLER's directory, so run from anywhere but the repo root it declared 14 false `UNREACHABLE (ssh-exit-127)`. Now the helper resolves from the script's own location: running it from `/` prints `bash /tmp/dgtest/scripts/pct-run.sh 104 …` and produces the same verdicts as running it from the containing directory. 2. **No scanner failure rendered as a guest verdict** - with the helper genuinely absent it prints a single `SCANNER ERROR: helper not found: <abs path>` line and zero UNREACHABLE verdicts. 3. **Figure mapping is now column-for-column correct against the authority.** Raw `df -P /` on kagentz gives `102996088 9301244 89459408 10%` (total / used / avail in 1K blocks); the scan renders `10% (8.9G used of 98.2G total, 85.3G free)`. The previous revision printed the pair as `(102996088/9301220)` - total before used, unlabelled, so the big number read as usage, and the used value did not even match the authority byte-for-byte. 4. Full run: every covered guest resolves with a real figure and the probe that produced it (`docker-vm` 11% / 15.9G used of 157.3G total). Reviewers: re-run it from two different working directories and confirm identical output; confirm a misconfigured helper yields SCANNER ERROR rather than a guest verdict; and spot-check two guests' figures against a raw `df -P /` on the guest itself.
abiba-bot added 4 commits 2026-09-14 14:20:35 +00:00
fix: add os import for HELPER_PCT_RUN.readable() check
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 3s
5ed6f8179c
abiba-bot merged commit 6a1c4967db into master 2026-09-14 14:43:44 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SyslogSolution/prose-contracts#91