fix(disk-gc): hard guest-level report-only gate for CT 111/.129 + correct stale fleet map #81

Merged
abiba-bot merged 9 commits from fix/disk-gc-report-only-129 into master 2026-09-12 19:13:45 +00:00
Owner

Why (URGENT boundary hold)

CT 111 (tdunna, 192.168.68.129, on storepve) is Theo's box and is detect-and-report-only per the captain (2026-08-17, re-confirmed 2026-09-10: Theo handles it himself).

disk-gc-threat-response.prose.md defined AMBER as "GC scheduled for next run" and its Execution loop called gc-executor for every threat with no guest-level exclusion — so a single AMBER reading on CT 111 would have scheduled GC commands (apt clean/autoremove, journalctl --vacuum-*, /var/log + /tmp deletion, snap removal, docker prune) against someone else's box. The only marker was frontmatter report_only_agents, which names an agent while the scan unit is a guest — it can silently miss the guest it lives on.

The gate

  • Guest/host-keyed exclusion (report_only_guests), matched on id/hostname/IP. An excluded guest is alerted and structurally cannot reach gc-executor (if/else — no reliance on non-grammar keywords).
  • Implemented in scripts/disk-gc-plan.py, which the Execution loop calls; the loop must not reimplement the gate.
  • Fail-closed everywhere: an unidentified target never gets GC; an exclusion entry without an identity key is a fatal error; an empty/missing exclusion list is a fatal error (the gate can never be silently disabled).
  • Identity normalisation so encodings cannot slip past (111.0, "lxc/111", "0111" all match).

Verification

Dry-run trace (scripts/disk-gc-plan.py, CT 111 at each level and encoding):

{"id":111,"usage_pct":84}        -> 111 AMBER 84.0% -> REPORT-ONLY (no GC)
{"id":111,"usage_pct":90}        -> 111 RED 90.0% -> REPORT-ONLY (no GC)
{"id":111,"usage_pct":97}        -> 111 CRITICAL 97.0% -> REPORT-ONLY (no GC)
{"id":111.0,"usage_pct":97}      -> REPORT-ONLY (no GC)
{"ct":"lxc/111","usage_pct":97}  -> REPORT-ONLY (no GC)
{"id":"0111","usage_pct":97}     -> REPORT-ONLY (no GC)
{"usage_pct":97}                 -> REPORT-ONLY — "unidentified target - refusing to schedule GC"

Real fleet run — runs/disk-gc-ct111-report-only-verification-20260912.log (25-entry live scan):

 111 tdunna AMBER 84.0% -> REPORT-ONLY (no GC) — Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10
 acerpve AMBER 77.0% -> gc-executor
 amdpve  AMBER 76.0% -> gc-executor

No GC command was executed against 192.168.68.129. tests/test_disk_gc_report_only.py (10 passed) executes the real planner for every case above.

Same-file drift corrected (found in the same review)

  • Guest map corrected against pvesh get /cluster/resources: 105 kagentz → amdpve (was minipve), 111 tdunna → storepve (was amdpve), added 118/119/120, counts fixed to 20 guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts.
  • scripts/pct-run.sh had the same stale map (105/111 wrong node, 120 missing) — that is why pct-run 111 and pct-run 105 failed. Now verified: pct-run 111 → tdunna, 105 → kagentz, 120 → adguard2.
  • Folds in disk-gc-ct100-probe-gap-20260911: CT 100 shared the stale access layer; verified scripts/pct-run.sh 100 "df -P /" → 23%. Documented that the scanner must probe CT 100 like any other guest (it runs inside CT 100, which has no pct binary).
  • QEMU VMs are reached via direct SSH, not pct-run (scope wording fixed).
  • infrastructure-control.prose.md + scripts/prose-ai-review.sh ground truth corrected to the live map; hermes-agent-baseline.prose.md Agent Map row (Koby 111) corrected.

Validation

no-mistakes run 01M2BEA33G5105M9QQE9NW3WJE — passed after six review rounds that hardened the gate (it caught real fail-open holes: a ct-keyed scan, encoded ids, an empty exclusion list, and a non-grammar continue).

## Why (URGENT boundary hold) CT 111 (`tdunna`, `192.168.68.129`, on **storepve**) is Theo's box and is **detect-and-report-only** per the captain (2026-08-17, re-confirmed 2026-09-10: Theo handles it himself). `disk-gc-threat-response.prose.md` defined AMBER as *"GC scheduled for next run"* and its Execution loop called `gc-executor` for **every** threat with **no guest-level exclusion** — so a single AMBER reading on CT 111 would have scheduled GC commands (apt clean/autoremove, `journalctl --vacuum-*`, `/var/log` + `/tmp` deletion, snap removal, docker prune) against someone else's box. The only marker was frontmatter `report_only_agents`, which names an **agent** while the scan unit is a **guest** — it can silently miss the guest it lives on. ## The gate - **Guest/host-keyed** exclusion (`report_only_guests`), matched on id/hostname/IP. An excluded guest is alerted and **structurally** cannot reach `gc-executor` (if/else — no reliance on non-grammar keywords). - Implemented in `scripts/disk-gc-plan.py`, which the Execution loop calls; the loop must not reimplement the gate. - **Fail-closed everywhere**: an unidentified target never gets GC; an exclusion entry without an identity key is a fatal error; an empty/missing exclusion list is a fatal error (the gate can never be silently disabled). - Identity normalisation so encodings cannot slip past (`111.0`, `"lxc/111"`, `"0111"` all match). ## Verification **Dry-run trace** (`scripts/disk-gc-plan.py`, CT 111 at each level and encoding): ``` {"id":111,"usage_pct":84} -> 111 AMBER 84.0% -> REPORT-ONLY (no GC) {"id":111,"usage_pct":90} -> 111 RED 90.0% -> REPORT-ONLY (no GC) {"id":111,"usage_pct":97} -> 111 CRITICAL 97.0% -> REPORT-ONLY (no GC) {"id":111.0,"usage_pct":97} -> REPORT-ONLY (no GC) {"ct":"lxc/111","usage_pct":97} -> REPORT-ONLY (no GC) {"id":"0111","usage_pct":97} -> REPORT-ONLY (no GC) {"usage_pct":97} -> REPORT-ONLY — "unidentified target - refusing to schedule GC" ``` **Real fleet run** — `runs/disk-gc-ct111-report-only-verification-20260912.log` (25-entry live scan): ``` 111 tdunna AMBER 84.0% -> REPORT-ONLY (no GC) — Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10 acerpve AMBER 77.0% -> gc-executor amdpve AMBER 76.0% -> gc-executor ``` **No GC command was executed against 192.168.68.129.** `tests/test_disk_gc_report_only.py` (10 passed) executes the real planner for every case above. ## Same-file drift corrected (found in the same review) - Guest map corrected against `pvesh get /cluster/resources`: **105 kagentz → amdpve** (was minipve), **111 tdunna → storepve** (was amdpve), added **118/119/120**, counts fixed to **20 guests (17 LXC + 3 QEMU VMs)** + 3 GPU bare-metal hosts. - `scripts/pct-run.sh` had the **same stale map** (105/111 wrong node, 120 missing) — that is why `pct-run 111` and `pct-run 105` failed. Now verified: `pct-run 111` → tdunna, `105` → kagentz, `120` → adguard2. - **Folds in `disk-gc-ct100-probe-gap-20260911`**: CT 100 shared the stale access layer; verified `scripts/pct-run.sh 100 "df -P /"` → 23%. Documented that the scanner must probe CT 100 like any other guest (it runs *inside* CT 100, which has no `pct` binary). - QEMU VMs are reached via direct SSH, not `pct-run` (scope wording fixed). - `infrastructure-control.prose.md` + `scripts/prose-ai-review.sh` ground truth corrected to the live map; `hermes-agent-baseline.prose.md` Agent Map row (Koby 111) corrected. ## Validation no-mistakes run `01M2BEA33G5105M9QQE9NW3WJE` — **passed** after six review rounds that hardened the gate (it caught real fail-open holes: a `ct`-keyed scan, encoded ids, an empty exclusion list, and a non-grammar `continue`).
abiba-bot added 9 commits 2026-09-12 19:11:53 +00:00
CT 111 (tdunna, 192.168.68.129) is Theo's box and is report-only per the captain
(2026-08-17, re-confirmed 2026-09-10). The contract defined AMBER as 'GC scheduled
for next run' and its Execution loop called gc-executor for EVERY threat with no
guest-level exclusion - so a single AMBER reading there would have scheduled apt
clean / journal vacuum / log+tmp deletion against someone else's box. The only
marker was frontmatter report_only_agents, which names an AGENT while the scan unit
is a GUEST.

- Gate the Execution loop on a guest/host-keyed report_only_guests block (guest id,
  hostname and IP all match); an excluded guest is alerted and skipped, so no
  gc-executor call is constructed for it at any level.
- Carry the ruling in the contract body next to the loop, not only in frontmatter.
- scripts/disk-gc-plan.py: executable planner that reads the contract's
  authoritative exclusion block and emits the action plan; tests/ covers it.
- Correct the stale fleet map against pvesh /cluster/resources: CT 105 -> amdpve
  (was minipve), CT 111 -> storepve (was amdpve), add guests 118/119/120, and fix
  the '15 CTs' counts (20 guests: 17 LXC + 3 QEMU VMs).
- scripts/pct-run.sh: same stale map (105/111 wrong node, 120 missing) - this is
  why pct-run 111/105 failed.
- Fold in disk-gc-ct100-probe-gap-20260911: CT 100 verifiably works through
  pct-run now that the map is correct; documented that the scanner must probe it
  like any other guest, never via a local-only path (the scanner runs inside CT 100).
docs(disk-gc): attach real fleet-run verification for the CT 111 report-only gate
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
672bf8a912
Intent requires a production run log, not only the unit-level dry run. This is a real
scan of the live 25-entry fleet through scripts/disk-gc-plan.py: CT 111 (tdunna) at 84%
AMBER is alerted as report-only and no gc-executor row is emitted for it; the owned hosts
acerpve .9 (77%) and amdpve .15 (76%) still receive gc-executor. No GC command was run
against 192.168.68.129.
abiba-bot merged commit b3644f0292 into master 2026-09-12 19:13:45 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SyslogSolution/prose-contracts#81