Files
prose-contracts/disk-gc-threat-response.prose.md
root 39209c7ac9
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
fix: untrack host-disk-bands.json and document gitignored status
The state file is runtime state (rewritten every scan), so tracking it in git
means:
- every executor's clone becomes permanently dirty after one run
- a scan in one clone produces a merge conflict with a scan in another
- the committed baseline can be stale in a way nobody notices

Added to .gitignore and removed from the index. Contract updated to say
'the state file lives at <abs path> and is gitignored runtime state - the
scanner creates it on first run'.
2026-09-16 00:45:46 +00:00

19 KiB

report_only_agents, kind, name, description, id, version
report_only_agents kind name description id version
koby
responsibility disk-gc-threat-response Recurring disk health scan, garbage collection, and threat response across 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts (fleet verified against `pvesh get /cluster/resources` 2026-09-12). Triggered by incident 2026-07-04 where CT 105 (kagentz) hit 87% disk (49G/59G) from Docker image bloat — 5 dangling images, 15 build cache layers. Recovered 35.67GB. Second incident 2026-07-09: amdpve (.15) Docker bloat (15.5GB abandoned HIP image), reclaimed 11.56GB. 067NV8KJ03ZG71S44N41F31022 1.0.0

Disk GC & Threat Response

Goal

Every container in the fleet stays below 80% disk usage through regular inspection and automated garbage collection. Critical breaches are detected, remediated, and logged within 5 minutes of discovery.

Scope

All 20 Proxmox guests (17 LXC via pct-run + 3 QEMU VMs via direct SSH) + 3 GPU bare-metal hosts via direct SSH. Docker hosts get special attention:

Host CT Disk Risk GC Strategy
kagentz 105 HIGH — Agent Zero builds images docker system prune -a
syslog-api 116 HIGH — Prometheus data, 10 containers docker system prune, log rotate
docker-vm 109 HIGH — 16 containers across 4 stacks, NFS mounts docker system prune, check mounts
amdpve — MED — GPU bare metal, Docker for one-off builds docker system prune -a
abiba 100 LOW — local docker, go cache apt/docker/log prune
All other CTs — LOW — no Docker apt clean, log rotate
GPU bare metal (.8, .110) — LOW — no Docker on GPU hosts log rotate

Note: CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.

Threat Levels (GUEST filesystems)

Level Threshold Response Escalation
GREEN < 75% Log only None
AMBER 75-84% Warn via Zulip DM GC scheduled for next run
RED 85-94% Immediate GC attempt Zulip DM + channel alert
CRITICAL ≥ 95% Aggressive GC + emergency cleanup Zulip + relay to Kwame
FULL 100% (df shows 100%) Stop writes, manual intervention Call/Telegram Kwame

Host Filesystem Thresholds (HOST nodes — separate bands from guest bands)

Host filesystems have their own risk profile and their own bands. A host root near full is a real risk (backup staging writes to it, thin-pool metadata pressure), a media volume near full is a capacity decision for the owner, and the backup datastore near full breaks backups. These are report-only — no automatic deletion of media or datastore content ever.

Level Threshold Response Escalation
HOST-WARN 85% Name the volume + % + absolute free space in the scan output None
HOST-AMBER 90% Name the volume + % + absolute free space; flag for owner attention Zulip DM to owner (state-change only)
HOST-RED 95% Name the volume + % + absolute free space; flag for immediate owner attention Zulip DM + channel alert (state-change only)

Escalations are STATE-CHANGE driven, not per-run. A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.

State lives in a small JSON state file: the scanner resolves it to an absolute path from the script's location — $(dirname "$0")/state/host-disk-bands.json (i.e., the state/ directory next to the scripts/ directory in the prose-contracts repo). The state file is gitignored runtime state — the scanner creates it on first run. Keyed by host/volume → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.)

Volume naming rule: Every host line MUST name the volume and what lives on it. Example output:

storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED
storepve /media/reanim (media): 86% (797G/932G, 135G free) -> HOST-WARN
storepve /media/mediastore (media): 77% (5.3T/7.3T, 1.6T free) -> GREEN
storepve /dev/mapper/pve-root (host-root): 81% (73G/94G, 17G free) -> GREEN
storepve tank (pbs-datastore): 1% (128K/12T, 12T free) -> GREEN

First-run behavior: When the state file does not yet exist (first scan), the current band of every volume is recorded as the baseline WITHOUT alerting — a first run would otherwise DM every already-elevated volume at once. From the second run onward, transitions alert.

Action classes by volume type:

  • host-root: near full = real risk (backup staging, thin-pool metadata, journald). HOST-AMBER or above → owner must investigate.
  • media (/media/*): near full = capacity decision for the owner. HOST-AMBER or above → report only, never auto-delete.
  • pbs-datastore (tank, ZFS): near full = breaks Proxmox Backup Server. HOST-RED → escalate immediately.

Justification (measured 2026-09-15, firstmate): storepve (192.168.68.6) shows /media/easystore2 at 96%, /media/reanim at 86%, /dev/mapper/pve-root at 81%, and tank at 1%. The previous scan reported "0/19 guests >80%, no GC action needed" while easystore2 sat at 96% — the host filesystems were printed but never banded and never acted on. Two incidents this weekend (amdpve host root filled during container backup staging, acerpve root went read-only when its thin pool errored) showed the host filesystem is the thing that breaks, not the guest's.

Requires

  • SSH access to all Docker hosts (matrix from infrastructure-control pattern)
  • Proxmox API access from abiba (token already provisioned)
  • Zulip bot credentials for alerting

Maintains

Per-CT disk state, GC history, and threat resolution ledger. Each entry carries the CT ID, hostname, node, disk usage snapshot, GC action taken, and reclaimed bytes. Postcondition: no CT runs above 85% for more than one scan cycle without documented reason.

disk-health

Current disk state for every CT: pct_used, pct_avail, rootfs_size, last GC timestamp, and active threat level.

gc-history

Append-only log of every GC action: timestamp, CT, action taken, bytes reclaimed, and whether threat was resolved.

threat-log

Active and resolved threat entries with severity, timestamps, remediation applied, and escalation trail.

Continuity

  • self-driven: full fleet scan every 6 hours
  • May also be invoked manually: prose run disk-gc-threat-response
  • Threat-driven: if Amber/Red/Critical detected, immediate GC phase activates

Scanner: scripts/disk-gc-scan.py

The fleet scan is executed by scripts/disk-gc-scan.py, which makes reachability verdicts deterministic:

  1. Retry on failure: Each probe retries once before declaring a guest unreachable.
  2. Named probe target: Every rendered line names the guest, CT id, node, and access method actually used.
  3. Failure kind printed: An unreachable guest is reported with its failure kind (timeout, ssh-auth, no-route, conn-refused, ssh-exit-N) — never as a bare "unreachable" verdict.
  4. Per-guest access method: The correct access path is selected from a per-guest map so the wrong path cannot be picked by an executor improvising:
    • CT 105 (kagentz) = ssh root@kagentz (NOT pct exec 105 — pct exec sees loop0/59G instead of the real 99G filesystem)
    • CT 109 (docker-vm) = ssh root@192.168.68.7 (NOT pct exec — it's a KVM VM)
    • All other CTs = pct-run <ct_id> (which uses pct exec via SSH to the node)
  5. Every figure traces to a named probe: The scan output prints the exact command that produced each disk figure, so two different guests can never render identical numbers without the probe commands proving it.

Run: python3 scripts/disk-gc-scan.py (or --json for machine-readable output).

The scan feeds into scripts/disk-gc-plan.py, which applies the report-only gate from the report_only_guests YAML block above.

Shape

  • self: scan all CTs via Proxmox API + SSH exec, trigger GC, alert
  • delegates:
    • disk-scanner: per-CT disk check (pct exec df or SSH)
    • gc-docker: docker system prune execution
    • gc-system: apt clean, log rotate, tmp cleanup
    • alerter: Zulip notification dispatch
  • prohibited: deleting user data, removing running containers, force-killing production services

Runtime

  • timeout: 300 seconds per CT (GC may take time on large docker hosts)
  • retry: 2 attempts for SSH failures before marking a CT unreachable

Execution

Host filesystems: report-only, NEVER auto-delete

Host filesystems are always report-only — no automatic deletion of media or datastore content. A host root near full requires investigation by the owner, but the executor must never delete content on a host filesystem. This restriction is absolute.

Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)

CT 111 / hostname tdunna / 192.168.68.129 is DETECT-AND-REPORT-ONLY. It belongs to Theo. The captain ruled 2026-08-17 and re-confirmed 2026-09-10 that Theo handles CT 111 himself. At every threat level — AMBER, RED, or CRITICAL — the executor must:

  • push the threat row and alert the owner, and
  • never call gc-executor, and never run any GC command against that guest: no apt-get clean/autoremove, no journalctl --vacuum-*, no find /var/log -delete, no /tmp//var/tmp deletion, no snap removal, no docker system prune.

This gate is keyed on guest id / hostname / IP, not on an agent name. The frontmatter report_only_agents marker (e.g. koby) names an AGENT while the scan unit is a GUEST, so an agent-name marker can silently miss the guest it lives on — it must never be the only gate.

The authoritative machine-readable exclusion list is the YAML block below. The executor reads it at run time; scripts/disk-gc-plan.py turns a fleet scan into the action plan using it. Extend the list here, never by hand-maintaining a second copy. The Execution loop below MUST call that planner and MUST NOT reimplement the gate.

# disk-gc report-only guests — authoritative. Keyed on guest/host, not agent.
report_only_guests:
  - guest: 111
    hostname: tdunna
    ip: 192.168.68.129
    node: storepve
    reason: "Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10"

Loop

let fleet = call disk-scanner
  scope: all

-- The report-only gate is IMPLEMENTED IN scripts/disk-gc-plan.py and MUST NOT be
-- reimplemented here. That planner reads the contract's `report_only_guests` YAML block
-- and matches on guest id OR hostname OR IP, so the tested gate is the executed gate.
let plan = call disk-gc-plan
  fleet: fleet

for row in plan:
  if row.action == "report-only":
    -- Excluded guest: alert only. No gc-executor call is constructed for it, at any level.
    call alerter
      threat: row
      result: { action: "report-only", reason: row.reason }
  else:
    let result = call gc-executor
      ct: row.target
      level: row.level
      strategy: lookup-gc-strategy(row.target)

    call alerter
      threat: row
      result: result

call summary-reporter
  fleet: fleet
  plan: plan

GC Strategies by Host Type

Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109, amdpve .15)

# Phase 1: Safe prune (won't touch running containers' images)
docker system prune -f

# Phase 2: Aggressive (if still > 85% after Phase 1)
docker system prune -a --force

# Phase 3: Emergency (if still > 95%)
docker system prune -a --force --volumes
docker builder prune --all --force

# Verification after each phase
df -h /
docker system df

Non-Docker CTs

# Package cache
apt-get clean
apt-get autoremove --yes

# Log rotation
journalctl --vacuum-size=100M
find /var/log -type f -name "*.log" -mtime +30 -delete

# Temp files
find /tmp -type f -mtime +7 -delete
find /var/tmp -type f -mtime +30 -delete

# Snap (if installed)
snap list --all | awk '/disabled/ {print $1, $3}' | while read snap rev; do
  snap remove "$snap" --revision="$rev"
done

Special Cases

CT Special GC
116 (syslog-api) Prometheus retention: check --storage.tsdb.retention.time
100 (abiba) Go module cache: go clean -cache -modcache if > 500MB
106 (ra-h-os) Check relay DB size, enforce TTL
109 (docker-vm) Check /media/storage and /media/mediastore mounts first

Alert Templates

HOST-WARN / HOST-AMBER / HOST-RED (host filesystems, report-only)

⚠️ Host Disk — {hostname} {volume} ({volume_type}): {pct}% ({used}/{total}, {free} free) -> {level}
Action: {volume_type-specific action}

AMBER (75-84%)

⚠️ Disk GC — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
Next scheduled GC will attempt cleanup. No immediate action needed.

RED (85-94%)

🚨 Disk Threat — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
GC executed: reclaimed {reclaimed}G. New usage: {new_pct}%.
Status: {resolved|still elevated — {reason}}

CRITICAL (≥95%)

🔥 CRITICAL Disk — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
Emergency GC: reclaimed {reclaimed}G. New usage: {new_pct}%.
{service_status} — {manual_action if needed}
@Kwame — container nearly full, resolved: {yes_no}

Incident Log: 2026-07-04 — kagentz docker bloat

Discovery

Kwame noticed kagentz was full, asked Abiba to investigate.

Diagnosis

  • CT 105 (kagentz, amdpve): 49G used / 59G total (87%)
  • /var/lib = 55G (93% of disk)
  • Docker images: 7 total, 5 dangling, 47.83GB disk, 34.98GB reclaimable
  • 1 stopped container, 1 unused network, 15 build cache layers

Resolution

docker system prune -a --force
→ Reclaimed 35.67GB
→ Post: 16G / 59G (27%), 41G free
→ 5 dangling images removed (kagentz-bridge + old builds)
→ 1 stopped container removed
→ 1 unused network removed (kagentz-bridge_default)
→ 15 build cache layers removed

Root Cause

Agent Zero's iterative development pattern (rebuild kagentz-bridge) creates dangling images and orphaned build cache. No automated GC was in place.

Preventive Measures

  • This contract now runs disk GC fleet-wide every 6 hours
  • Docker hosts get docker system prune on amber, -a --force on red
  • kagentz is flagged as HIGH risk due to development activity
  • amdpve is flagged for Docker bloat monitoring — abandoned build images accumulate

## Incident Log: 2026-07-09 — amdpve docker bloat

### Discovery
Scheduled fleet disk scan across all 20 Proxmox guests (17 LXC via `pct-run`, 3 QEMU VMs via direct SSH) + 3 GPU bare-metal hosts.
> **Report-only gate applies to this scan:** CT 111 (`tdunna`, 192.168.68.129) is alerted but never
> garbage-collected at any level.
amdpve (.15) flagged at 78% (AMBER threshold: 75%).

### Diagnosis
- amdpve (.15, Strix Halo host): 70G used / 94G total (78%)
- Docker images: 1 image, 0 containers running, 15.47GB (100% reclaimable)
- Image: `llama-strix-hip:latest` — abandoned ROCm/HIP Docker build from 7 days ago
- Root cause: Strix Halo migrated from Docker-based HIP path to bare-metal Vulkan
  (`/root/llama.cpp/build-vk/`) but the old Docker image was never cleaned up
- Not a running service — zero containers, zero active volumes

### Resolution

docker system prune -a --force → Reclaimed 11.56GB → Post: 55G / 94G (62%), 35G free → 1 image removed (llama-strix-hip:latest, 15.5GB) → 8 build cache layers removed → amdpve now GREEN


### Root Cause
Technology migration (Docker HIP → bare-metal Vulkan) left orphaned build
artifacts. Docker on amdpve serves no running purpose — it's only used for
one-off GPU builds. No automated post-migration cleanup was in place.

### Preventive Measures
- amdpve added to Docker GC scan list
- Post-migration cleanup step added: after any GPU backend migration, prune
  the old backend's Docker images within 24 hours
- Contract now scans GPU bare-metal hosts alongside CTs
- Access via `pct-run` script for all CTs (no hardcoded IPs)

## Access Matrix (verified against `pvesh get /cluster/resources` 2026-09-12)

### Guest Access (via `pct-run` — CT id only, node resolved by `scripts/pct-run.sh`)
| Guest | Name | Node | Type | Status |
|------|------|------|------|--------|
| 100 | abiba | minipve | lxc | ✅ reachable (probed via pct-run like any other guest; no local shortcut) |
| 102 | adguard | minipve | lxc | ✅ reachable |
| 104 | authentik | minipve | lxc | ✅ reachable |
| 105 | kagentz | **amdpve** | lxc | ✅ reachable (was documented as minipve — corrected) |
| 106 | ra-h-os | storepve | lxc | ✅ reachable |
| 107 | pbs | storepve | lxc | ✅ reachable |
| 108 | media | storepve | lxc | ✅ reachable |
| 110 | gitea | minipve | lxc | ✅ reachable |
| 111 | tdunna | **storepve** | lxc | ⛔ **REPORT-ONLY** (192.168.68.129, Theo's box — no GC at any level) |
| 112 | tanko | amdpve | lxc | ✅ reachable |
| 113 | baggy | amdpve | lxc | ✅ reachable |
| 115 | scottdenya | amdpve | lxc | ✅ reachable |
| 116 | syslog-api | minipve | lxc | ✅ reachable |
| 117 | zulip | storepve | lxc | ✅ reachable |
| 118 | jdownloader | storepve | lxc | ✅ reachable |
| 119 | infisical-vault | minipve | lxc | ✅ reachable |
| 120 | adguard2 | amdpve | lxc | ✅ reachable |

### QEMU VMs (via direct SSH)
| VM | Name | Node | IP | Status |
|----|------|------|-----|--------|
| 101 | llm-gpu (workload now bare metal .8) | acerpve | — | ✅ reachable |
| 103 | ocu-llm (workload now bare metal .110) | ocupve | — | ✅ reachable |
| 109 | docker-vm | storepve | 192.168.68.7 | ✅ reachable |

### GPU Bare Metal (via direct SSH)
| Host | IP | GPU | Status |
|------|-----|-----|--------|
| llm-gpu | 192.168.68.8 | RTX 3090 | ✅ reachable |
| ocu-llm | 192.168.68.110 | RTX 5070 | ✅ reachable |
| amdpve | 192.168.68.15 | Strix Halo | ✅ reachable |

> **Fleet count:** 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts. Corrected
> 2026-09-12: CT 105 → amdpve, CT 111 → storepve, and guests 118/119/120 were missing.
>
> **CT 100 probe gap (folded in):** CT 100 previously reported "unreachable (not reported)" every
> run. Root cause is the same stale access layer: `pct-run` resolves the guest's node from its map,
> and the map/contract must reflect `pvesh /cluster/resources`. Verified working from inside CT 100:
> `scripts/pct-run.sh 100 "df -P / | tail -1"` → `23% /`. Probe CT 100 through `pct-run` like any
> other guest — never through a local-only path, since the scanner itself runs inside CT 100 and a
> container has no `pct` binary.
>
> **KVM VM:** CT 109 (docker-vm) is a QEMU VM, not LXC — access via SSH .7.
> **NOTE:** For kagentz (CT 105), use `ssh root@kagentz` (hostname), NOT `pct exec 105` — `pct exec 105` shows loop0 (59G) while `ssh root@kagentz` shows the real filesystem (99G). For docker-vm (CT 109), use `ssh root@192.168.68.7`, not `pct exec`.