CT 111 (tdunna, 192.168.68.129) is Theo's box and is report-only per the captain (2026-08-17, re-confirmed 2026-09-10). The contract defined AMBER as 'GC scheduled for next run' and its Execution loop called gc-executor for EVERY threat with no guest-level exclusion - so a single AMBER reading there would have scheduled apt clean / journal vacuum / log+tmp deletion against someone else's box. The only marker was frontmatter report_only_agents, which names an AGENT while the scan unit is a GUEST. - Gate the Execution loop on a guest/host-keyed report_only_guests block (guest id, hostname and IP all match); an excluded guest is alerted and skipped, so no gc-executor call is constructed for it at any level. - Carry the ruling in the contract body next to the loop, not only in frontmatter. - scripts/disk-gc-plan.py: executable planner that reads the contract's authoritative exclusion block and emits the action plan; tests/ covers it. - Correct the stale fleet map against pvesh /cluster/resources: CT 105 -> amdpve (was minipve), CT 111 -> storepve (was amdpve), add guests 118/119/120, and fix the '15 CTs' counts (20 guests: 17 LXC + 3 QEMU VMs). - scripts/pct-run.sh: same stale map (105/111 wrong node, 120 missing) - this is why pct-run 111/105 failed. - Fold in disk-gc-ct100-probe-gap-20260911: CT 100 verifiably works through pct-run now that the map is correct; documented that the scanner must probe it like any other guest, never via a local-only path (the scanner runs inside CT 100).
14 KiB
report_only_agents, report_only_guests, kind, name, description, id, version
| report_only_agents | report_only_guests | kind | name | description | id | version | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
responsibility | disk-gc-threat-response | Recurring disk health scan, garbage collection, and threat response across 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts (fleet verified against `pvesh get /cluster/resources` 2026-09-12). Triggered by incident 2026-07-04 where CT 105 (kagentz) hit 87% disk (49G/59G) from Docker image bloat — 5 dangling images, 15 build cache layers. Recovered 35.67GB. Second incident 2026-07-09: amdpve (.15) Docker bloat (15.5GB abandoned HIP image), reclaimed 11.56GB. | 067NV8KJ03ZG71S44N41F31022 | 1.0.0 |
Disk GC & Threat Response
Goal
Every container in the fleet stays below 80% disk usage through regular inspection and automated garbage collection. Critical breaches are detected, remediated, and logged within 5 minutes of discovery.
Scope
All 20 Proxmox guests (17 LXC + 3 QEMU VMs) via pct-run + 3 GPU bare-metal hosts via direct SSH.
Docker hosts get special attention:
| Host | CT | Disk Risk | GC Strategy |
|---|---|---|---|
| kagentz | 105 | HIGH — Agent Zero builds images | docker system prune -a |
| syslog-api | 116 | HIGH — Prometheus data, 10 containers | docker system prune, log rotate |
| docker-vm | 109 | HIGH — 16 containers across 4 stacks, NFS mounts | docker system prune, check mounts |
| amdpve | — | MED — GPU bare metal, Docker for one-off builds | docker system prune -a |
| abiba | 100 | LOW — local docker, go cache | apt/docker/log prune |
| All other CTs | — | LOW — no Docker | apt clean, log rotate |
| GPU bare metal (.8, .110) | — | LOW — no Docker on GPU hosts | log rotate |
Note: CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
Threat Levels
| Level | Threshold | Response | Escalation |
|---|---|---|---|
| GREEN | < 75% | Log only | None |
| AMBER | 75-84% | Warn via Zulip DM | GC scheduled for next run |
| RED | 85-94% | Immediate GC attempt | Zulip DM + channel alert |
| CRITICAL | ≥ 95% | Aggressive GC + emergency cleanup | Zulip + relay to Kwame |
| FULL | 100% (df shows 100%) | Stop writes, manual intervention | Call/Telegram Kwame |
Requires
- SSH access to all Docker hosts (matrix from infrastructure-control pattern)
- Proxmox API access from abiba (token already provisioned)
- Zulip bot credentials for alerting
Maintains
Per-CT disk state, GC history, and threat resolution ledger. Each entry carries the CT ID, hostname, node, disk usage snapshot, GC action taken, and reclaimed bytes. Postcondition: no CT runs above 85% for more than one scan cycle without documented reason.
disk-health
Current disk state for every CT: pct_used, pct_avail, rootfs_size, last GC
timestamp, and active threat level.
gc-history
Append-only log of every GC action: timestamp, CT, action taken, bytes reclaimed, and whether threat was resolved.
threat-log
Active and resolved threat entries with severity, timestamps, remediation applied, and escalation trail.
Continuity
- self-driven: full fleet scan every 6 hours
- May also be invoked manually:
prose run disk-gc-threat-response - Threat-driven: if Amber/Red/Critical detected, immediate GC phase activates
Shape
self: scan all CTs via Proxmox API + SSH exec, trigger GC, alertdelegates:disk-scanner: per-CT disk check (pct exec df or SSH)gc-docker: docker system prune executiongc-system: apt clean, log rotate, tmp cleanupalerter: Zulip notification dispatch
prohibited: deleting user data, removing running containers, force-killing production services
Runtime
timeout: 300 seconds per CT (GC may take time on large docker hosts)retry: 2 attempts for SSH failures before marking a CT unreachable
Execution
Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)
CT 111 / hostname tdunna / 192.168.68.129 is DETECT-AND-REPORT-ONLY. It belongs to Theo.
The captain ruled 2026-08-17 and re-confirmed 2026-09-10 that Theo handles CT 111 himself.
At every threat level — AMBER, RED, or CRITICAL — the executor must:
- push the threat row and alert the owner, and
- never call
gc-executor, and never run any GC command against that guest: noapt-get clean/autoremove, nojournalctl --vacuum-*, nofind /var/log -delete, no/tmp//var/tmpdeletion, no snap removal, nodocker system prune.
This gate is keyed on guest id / hostname / IP, not on an agent name. The frontmatter
report_only_agents marker (e.g. koby) names an AGENT while the scan unit is a GUEST, so an
agent-name marker can silently miss the guest it lives on — it must never be the only gate.
The authoritative machine-readable exclusion list is the YAML block below. The executor
reads it at run time; scripts/disk-gc-plan.py turns a fleet scan into the action plan using it.
Extend the list here, never by hand-maintaining a second copy.
# disk-gc report-only guests — authoritative. Keyed on guest/host, not agent.
report_only_guests:
- guest: 111
hostname: tdunna
ip: 192.168.68.129
node: storepve
reason: "Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10"
Loop
let fleet = call disk-scanner
scope: all
let report_only = load-report-only-guests() -- from the YAML block above
let threats = []
for ct in fleet:
if ct.usage_pct >= 95:
push threats { ct: ct.id, level: "CRITICAL", pct: ct.usage_pct }
else if ct.usage_pct >= 85:
push threats { ct: ct.id, level: "RED", pct: ct.usage_pct }
else if ct.usage_pct >= 75:
push threats { ct: ct.id, level: "AMBER", pct: ct.usage_pct }
-- sort by severity descending
sort threats by pct desc
for threat in threats:
-- HARD GATE: an excluded guest is alerted and skipped. No gc-executor call is
-- constructed for it at any level, so no GC command can be emitted for it.
if threat.ct in report_only:
call alerter
threat: threat
result: { action: "report-only", reason: report_only[threat.ct].reason }
continue
let result = call gc-executor
ct: threat.ct
level: threat.level
strategy: lookup-gc-strategy(threat.ct)
call alerter
threat: threat
result: result
call summary-reporter
fleet: fleet
threats: threats
GC Strategies by Host Type
Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109, amdpve .15)
# Phase 1: Safe prune (won't touch running containers' images)
docker system prune -f
# Phase 2: Aggressive (if still > 85% after Phase 1)
docker system prune -a --force
# Phase 3: Emergency (if still > 95%)
docker system prune -a --force --volumes
docker builder prune --all --force
# Verification after each phase
df -h /
docker system df
Non-Docker CTs
# Package cache
apt-get clean
apt-get autoremove --yes
# Log rotation
journalctl --vacuum-size=100M
find /var/log -type f -name "*.log" -mtime +30 -delete
# Temp files
find /tmp -type f -mtime +7 -delete
find /var/tmp -type f -mtime +30 -delete
# Snap (if installed)
snap list --all | awk '/disabled/ {print $1, $3}' | while read snap rev; do
snap remove "$snap" --revision="$rev"
done
Special Cases
| CT | Special GC |
|---|---|
| 116 (syslog-api) | Prometheus retention: check --storage.tsdb.retention.time |
| 100 (abiba) | Go module cache: go clean -cache -modcache if > 500MB |
| 106 (ra-h-os) | Check relay DB size, enforce TTL |
| 109 (docker-vm) | Check /media/storage and /media/mediastore mounts first |
Alert Templates
AMBER (75-84%)
⚠️ Disk GC — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
Next scheduled GC will attempt cleanup. No immediate action needed.
RED (85-94%)
🚨 Disk Threat — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
GC executed: reclaimed {reclaimed}G. New usage: {new_pct}%.
Status: {resolved|still elevated — {reason}}
CRITICAL (≥95%)
🔥 CRITICAL Disk — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
Emergency GC: reclaimed {reclaimed}G. New usage: {new_pct}%.
{service_status} — {manual_action if needed}
@Kwame — container nearly full, resolved: {yes_no}
Incident Log: 2026-07-04 — kagentz docker bloat
Discovery
Kwame noticed kagentz was full, asked Abiba to investigate.
Diagnosis
- CT 105 (kagentz, amdpve): 49G used / 59G total (87%)
/var/lib= 55G (93% of disk)- Docker images: 7 total, 5 dangling, 47.83GB disk, 34.98GB reclaimable
- 1 stopped container, 1 unused network, 15 build cache layers
Resolution
docker system prune -a --force
→ Reclaimed 35.67GB
→ Post: 16G / 59G (27%), 41G free
→ 5 dangling images removed (kagentz-bridge + old builds)
→ 1 stopped container removed
→ 1 unused network removed (kagentz-bridge_default)
→ 15 build cache layers removed
Root Cause
Agent Zero's iterative development pattern (rebuild kagentz-bridge) creates dangling images and orphaned build cache. No automated GC was in place.
Preventive Measures
- This contract now runs disk GC fleet-wide every 6 hours
- Docker hosts get
docker system pruneon amber,-a --forceon red - kagentz is flagged as HIGH risk due to development activity
- amdpve is flagged for Docker bloat monitoring — abandoned build images accumulate
## Incident Log: 2026-07-09 — amdpve docker bloat
### Discovery
Scheduled fleet disk scan via `pct-run` across all 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts.
> **Report-only gate applies to this scan:** CT 111 (`tdunna`, 192.168.68.129) is alerted but never
> garbage-collected at any level.
amdpve (.15) flagged at 78% (AMBER threshold: 75%).
### Diagnosis
- amdpve (.15, Strix Halo host): 70G used / 94G total (78%)
- Docker images: 1 image, 0 containers running, 15.47GB (100% reclaimable)
- Image: `llama-strix-hip:latest` — abandoned ROCm/HIP Docker build from 7 days ago
- Root cause: Strix Halo migrated from Docker-based HIP path to bare-metal Vulkan
(`/root/llama.cpp/build-vk/`) but the old Docker image was never cleaned up
- Not a running service — zero containers, zero active volumes
### Resolution
docker system prune -a --force → Reclaimed 11.56GB → Post: 55G / 94G (62%), 35G free → 1 image removed (llama-strix-hip:latest, 15.5GB) → 8 build cache layers removed → amdpve now GREEN
### Root Cause
Technology migration (Docker HIP → bare-metal Vulkan) left orphaned build
artifacts. Docker on amdpve serves no running purpose — it's only used for
one-off GPU builds. No automated post-migration cleanup was in place.
### Preventive Measures
- amdpve added to Docker GC scan list
- Post-migration cleanup step added: after any GPU backend migration, prune
the old backend's Docker images within 24 hours
- Contract now scans GPU bare-metal hosts alongside CTs
- Access via `pct-run` script for all CTs (no hardcoded IPs)
## Access Matrix (verified against `pvesh get /cluster/resources` 2026-09-12)
### Guest Access (via `pct-run` — CT id only, node resolved by `scripts/pct-run.sh`)
| Guest | Name | Node | Type | Status |
|------|------|------|------|--------|
| 100 | abiba | minipve | lxc | ✅ reachable (probed via pct-run like any other guest; no local shortcut) |
| 102 | adguard | minipve | lxc | ✅ reachable |
| 104 | authentik | minipve | lxc | ✅ reachable |
| 105 | kagentz | **amdpve** | lxc | ✅ reachable (was documented as minipve — corrected) |
| 106 | ra-h-os | storepve | lxc | ✅ reachable |
| 107 | pbs | storepve | lxc | ✅ reachable |
| 108 | media | storepve | lxc | ✅ reachable |
| 110 | gitea | minipve | lxc | ✅ reachable |
| 111 | tdunna | **storepve** | lxc | ⛔ **REPORT-ONLY** (192.168.68.129, Theo's box — no GC at any level) |
| 112 | tanko | amdpve | lxc | ✅ reachable |
| 113 | baggy | amdpve | lxc | ✅ reachable |
| 115 | scottdenya | amdpve | lxc | ✅ reachable |
| 116 | syslog-api | minipve | lxc | ✅ reachable |
| 117 | zulip | storepve | lxc | ✅ reachable |
| 118 | jdownloader | storepve | lxc | ✅ reachable |
| 119 | infisical-vault | minipve | lxc | ✅ reachable |
| 120 | adguard2 | amdpve | lxc | ✅ reachable |
### QEMU VMs (via direct SSH)
| VM | Name | Node | IP | Status |
|----|------|------|-----|--------|
| 101 | llm-gpu (workload now bare metal .8) | acerpve | — | ✅ reachable |
| 103 | ocu-llm (workload now bare metal .110) | ocupve | — | ✅ reachable |
| 109 | docker-vm | storepve | 192.168.68.7 | ✅ reachable |
### GPU Bare Metal (via direct SSH)
| Host | IP | GPU | Status |
|------|-----|-----|--------|
| llm-gpu | 192.168.68.8 | RTX 3090 | ✅ reachable |
| ocu-llm | 192.168.68.110 | RTX 5070 | ✅ reachable |
| amdpve | 192.168.68.15 | Strix Halo | ✅ reachable |
> **Fleet count:** 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts. Corrected
> 2026-09-12: CT 105 → amdpve, CT 111 → storepve, and guests 118/119/120 were missing.
>
> **CT 100 probe gap (folded in):** CT 100 previously reported "unreachable (not reported)" every
> run. Root cause is the same stale access layer: `pct-run` resolves the guest's node from its map,
> and the map/contract must reflect `pvesh /cluster/resources`. Verified working from inside CT 100:
> `scripts/pct-run.sh 100 "df -P / | tail -1"` → `23% /`. Probe CT 100 through `pct-run` like any
> other guest — never through a local-only path, since the scanner itself runs inside CT 100 and a
> container has no `pct` binary.
>
> **KVM VM:** CT 109 (docker-vm) is a QEMU VM, not LXC — access via SSH .7.