PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 2s
1. Replace '(## Execution' + ')' with '## Execution' in 6 contract files 2. Make LOG_DIR honor CONTRACT_RUN_LOG_DIR env var (default: /var/log/contract-runs) 3. Fix Zulip alert: correct URL (https://chat.sysloggh.net/api/v1), user (abiba-bot@chat.sysloggh.net), take ZULIP_API_KEY from environment, and write alert failures to run log
452 lines
21 KiB
Markdown
452 lines
21 KiB
Markdown
---
|
|
report_only_agents:
|
|
- koby # ⛔ KOBY IS NEVER REPAIRED (Rule 17, 2026-08-17) — detect + report, never fix on .129
|
|
# ⛔ The guest/host-keyed report-only gate the GC executor MUST honour lives in the body
|
|
# "Hard gate" YAML block below — that block is authoritative and is the only copy.
|
|
kind: responsibility
|
|
name: disk-gc-threat-response
|
|
description: >
|
|
Recurring disk health scan, garbage collection, and threat response
|
|
across 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts
|
|
(fleet verified against `pvesh get /cluster/resources` 2026-09-12). Triggered by incident
|
|
2026-07-04 where CT 105 (kagentz) hit 87% disk (49G/59G) from
|
|
Docker image bloat — 5 dangling images, 15 build cache layers.
|
|
Recovered 35.67GB. Second incident 2026-07-09: amdpve (.15) Docker
|
|
bloat (15.5GB abandoned HIP image), reclaimed 11.56GB.
|
|
id: 067NV8KJ03ZG71S44N41F31022
|
|
version: 1.0.0
|
|
---
|
|
---
|
|
|
|
# Disk GC & Threat Response
|
|
|
|
## Goal
|
|
|
|
Every container in the fleet stays below 80% disk usage through regular inspection
|
|
and automated garbage collection. Critical breaches are detected, remediated, and
|
|
logged within 5 minutes of discovery.
|
|
|
|
## Scope
|
|
|
|
All 20 Proxmox guests (17 LXC via `pct-run` + 3 QEMU VMs via direct SSH) + 3 GPU bare-metal hosts via direct SSH.
|
|
Docker hosts get special attention:
|
|
|
|
| Host | CT | Disk Risk | GC Strategy |
|
|
|------|----|-----------|-------------|
|
|
| kagentz | 105 | HIGH — Agent Zero builds images | `docker system prune -a` |
|
|
| syslog-api | 116 | HIGH — Prometheus data, 10 containers | `docker system prune`, log rotate |
|
|
| docker-vm | 109 | HIGH — 16 containers across 4 stacks, NFS mounts | `docker system prune`, check mounts |
|
|
| amdpve | — | MED — GPU bare metal, Docker for one-off builds | `docker system prune -a` |
|
|
| abiba | 100 | LOW — local docker, go cache | apt/docker/log prune |
|
|
| All other CTs | — | LOW — no Docker | apt clean, log rotate |
|
|
| GPU bare metal (.8, .110) | — | LOW — no Docker on GPU hosts | log rotate |
|
|
|
|
> **Note:** CT 118 is now jdownloader (active on storepve). CT 101 (llm-gpu) → bare metal .8, CT 103 (ocu-llm) → bare metal .110.
|
|
|
|
## Threat Levels (GUEST filesystems)
|
|
|
|
| Level | Threshold | Response | Escalation |
|
|
|-------|-----------|----------|------------|
|
|
| **GREEN** | < 75% | Log only | None |
|
|
| **AMBER** | 75-84% | Warn via Zulip DM | GC scheduled for next run |
|
|
| **RED** | 85-94% | Immediate GC attempt | Zulip DM + channel alert |
|
|
| **CRITICAL** | ≥ 95% | Aggressive GC + emergency cleanup | Zulip + relay to Kwame |
|
|
| **FULL** | 100% (df shows 100%) | Stop writes, manual intervention | Call/Telegram Kwame |
|
|
|
|
## Host Filesystem Thresholds (HOST nodes — separate bands from guest bands)
|
|
|
|
Host filesystems have their own risk profile and their own bands. A host root near full is a real risk (backup staging writes to it, thin-pool metadata pressure), a media volume near full is a capacity decision for the owner, and the backup datastore near full breaks backups. These are **report-only** — no automatic deletion of media or datastore content ever.
|
|
|
|
| Level | Threshold | Response | Escalation |
|
|
|-------|-----------|----------|------------|
|
|
| **HOST-WARN** | 85% | Name the volume + % + absolute free space in the scan output | None |
|
|
| **HOST-AMBER** | 90% | Name the volume + % + absolute free space; flag for owner attention | Zulip DM to owner (state-change only) |
|
|
| **HOST-RED** | 95% | Name the volume + % + absolute free space; **media volumes: "capacity decision — owner to decide"; pbs-datastore/host-root: "immediate owner attention"** | Zulip DM + channel alert (state-change only) |
|
|
|
|
**Escalations are STATE-CHANGE driven, not per-run.** A volume alerts ONCE when it enters a higher band (GREEN->WARN, WARN->AMBER, AMBER->RED) and ONCE when it drops back down (a recovery notice). While a volume stays in the same band, it is reported in the scan output only — no DM, no channel alert. This prevents the same 96% easystore2 from re-DMing the owner on every 6-hour scan and drowning a real warning in noise.
|
|
|
|
**State lives in a small JSON state file:** the scanner resolves it to an absolute path from the script's location — `$(dirname "$0")/state/host-disk-bands.json` (i.e., the `state/` directory next to the `scripts/` directory in the prose-contracts repo). The state file is gitignored runtime state — the scanner creates it on first run. Keyed by `host/volume` → last-seen band. The scanner reads the prior band, compares to the current band, and DMs only on a transition. The state file is written after every scan. (Chosen over a periodic digest because the scan already runs every 6h and a transition is genuinely new, actionable state that warrants an immediate DM — but only once.)
|
|
|
|
**Volume naming rule:** Every host line MUST name the volume and what lives on it. Example output:
|
|
|
|
```
|
|
storepve /media/easystore2 (media): 96% (3.5T/3.7T, 177G free) -> HOST-RED
|
|
storepve /media/reanim (media): 86% (797G/932G, 135G free) -> HOST-WARN
|
|
storepve /media/mediastore (media): 77% (5.3T/7.3T, 1.6T free) -> GREEN
|
|
storepve /dev/mapper/pve-root (host-root): 81% (73G/94G, 17G free) -> GREEN
|
|
storepve tank (pbs-datastore): 1% (128K/12T, 12T free) -> GREEN
|
|
```
|
|
|
|
**First-run behavior:** When the state file does not yet exist (first scan), the current band of every volume is recorded as the baseline WITHOUT alerting — a first run would otherwise DM every already-elevated volume at once. From the second run onward, transitions alert.
|
|
|
|
**Action classes by volume type:**
|
|
- **host-root**: near full = real risk (backup staging, thin-pool metadata, journald). HOST-AMBER or above → owner must investigate.
|
|
- **media** (/media/*): near full = capacity decision for the owner. HOST-AMBER or above → report only, never auto-delete.
|
|
- **pbs-datastore** (tank, ZFS): near full = breaks Proxmox Backup Server. HOST-RED → escalate immediately.
|
|
|
|
**Justification (measured 2026-09-15, firstmate):** storepve (192.168.68.6) shows /media/easystore2 at 96%, /media/reanim at 86%, /dev/mapper/pve-root at 81%, and tank at 1%. The previous scan reported "0/19 guests >80%, no GC action needed" while easystore2 sat at 96% — the host filesystems were printed but never banded and never acted on. Two incidents this weekend (amdpve host root filled during container backup staging, acerpve root went read-only when its thin pool errored) showed the host filesystem is the thing that breaks, not the guest's.
|
|
|
|
## Requires
|
|
|
|
- SSH access to all Docker hosts (matrix from infrastructure-control pattern)
|
|
- Proxmox API access from abiba (token already provisioned)
|
|
- Zulip bot credentials for alerting
|
|
|
|
## Maintains
|
|
|
|
Per-CT disk state, GC history, and threat resolution ledger. Each entry carries
|
|
the CT ID, hostname, node, disk usage snapshot, GC action taken, and reclaimed
|
|
bytes. Postcondition: no CT runs above 85% for more than one scan cycle without
|
|
documented reason.
|
|
|
|
### disk-health
|
|
Current disk state for every CT: `pct_used`, `pct_avail`, `rootfs_size`, last GC
|
|
timestamp, and active threat level.
|
|
|
|
### gc-history
|
|
Append-only log of every GC action: timestamp, CT, action taken, bytes reclaimed,
|
|
and whether threat was resolved.
|
|
|
|
### threat-log
|
|
Active and resolved threat entries with severity, timestamps, remediation applied,
|
|
and escalation trail.
|
|
|
|
## Continuity
|
|
|
|
- self-driven: full fleet scan every 6 hours
|
|
- May also be invoked manually: `prose run disk-gc-threat-response`
|
|
- Threat-driven: if Amber/Red/Critical detected, immediate GC phase activates
|
|
|
|
## Scanner: scripts/disk-gc-scan.py
|
|
|
|
The fleet scan is executed by `scripts/disk-gc-scan.py`, which makes reachability
|
|
verdicts deterministic:
|
|
|
|
1. **Retry on failure:** Each probe retries once before declaring a guest unreachable.
|
|
2. **Named probe target:** Every rendered line names the guest, CT id, node, and
|
|
access method actually used.
|
|
3. **Failure kind printed:** An unreachable guest is reported with its failure kind
|
|
(timeout, ssh-auth, no-route, conn-refused, ssh-exit-N) — never as a bare
|
|
"unreachable" verdict.
|
|
4. **Per-guest access method:** The correct access path is selected from a per-guest
|
|
map so the wrong path cannot be picked by an executor improvising:
|
|
- CT 105 (kagentz) = `ssh root@kagentz` (NOT `pct exec 105` — pct exec sees
|
|
loop0/59G instead of the real 99G filesystem)
|
|
- CT 109 (docker-vm) = `ssh root@192.168.68.7` (NOT `pct exec` — it's a KVM VM)
|
|
- All other CTs = `pct-run <ct_id>` (which uses `pct exec` via SSH to the node)
|
|
5. **Every figure traces to a named probe:** The scan output prints the exact command
|
|
that produced each disk figure, so two different guests can never render
|
|
identical numbers without the probe commands proving it.
|
|
|
|
Run: `python3 scripts/disk-gc-scan.py` (or `--json` for machine-readable output).
|
|
|
|
The scan feeds into `scripts/disk-gc-plan.py`, which applies the report-only gate
|
|
from the `report_only_guests` YAML block above.
|
|
|
|
## Shape
|
|
|
|
- `self`: scan all CTs via Proxmox API + SSH exec, trigger GC, alert
|
|
- `delegates`:
|
|
- `disk-scanner`: per-CT disk check (pct exec df or SSH)
|
|
- `gc-docker`: docker system prune execution
|
|
- `gc-system`: apt clean, log rotate, tmp cleanup
|
|
- `alerter`: Zulip notification dispatch
|
|
- `prohibited`: deleting user data, removing running containers, force-killing
|
|
production services
|
|
|
|
## Runtime
|
|
|
|
- `timeout`: 300 seconds per CT (GC may take time on large docker hosts)
|
|
- `retry`: 2 attempts for SSH failures before marking a CT unreachable
|
|
|
|
## Execution
|
|
**Execution model**: This contract is executed by a host-scheduled cron job (see
|
|
`scripts/contract-run.sh`). The cron job runs the monitoring script directly on the
|
|
target host and appends the result to `/var/log/contract-runs/<contract>.log`. An
|
|
agent-session acknowledgement (a `done:` line in the ops status log) is NOT
|
|
execution — it only proves the agent read the result and reported it. The actual
|
|
monitoring work happens in the host cron job.
|
|
|
|
|
|
### Host filesystems: report-only, NEVER auto-delete
|
|
|
|
Host filesystems are always report-only — no automatic deletion of media or datastore content. A host root near full requires investigation by the owner, but the executor must never delete content on a host filesystem. This restriction is absolute.
|
|
|
|
### Hard gate: report-only guests (READ THIS BEFORE RUNNING GC)
|
|
|
|
**CT 111 / hostname `tdunna` / 192.168.68.129 is DETECT-AND-REPORT-ONLY.** It belongs to Theo.
|
|
The captain ruled 2026-08-17 and re-confirmed 2026-09-10 that Theo handles CT 111 himself.
|
|
At **every** threat level — AMBER, RED, or CRITICAL — the executor must:
|
|
|
|
- push the threat row and alert the owner, and
|
|
- **never** call `gc-executor`, and **never** run any GC command against that guest: no
|
|
`apt-get clean/autoremove`, no `journalctl --vacuum-*`, no `find /var/log -delete`, no
|
|
`/tmp`/`/var/tmp` deletion, no snap removal, no `docker system prune`.
|
|
|
|
This gate is keyed on **guest id / hostname / IP**, not on an agent name. The frontmatter
|
|
`report_only_agents` marker (e.g. `koby`) names an AGENT while the scan unit is a GUEST, so an
|
|
agent-name marker can silently miss the guest it lives on — it must never be the only gate.
|
|
|
|
**The authoritative machine-readable exclusion list is the YAML block below.** The executor
|
|
reads it at run time; `scripts/disk-gc-plan.py` turns a fleet scan into the action plan using it.
|
|
Extend the list here, never by hand-maintaining a second copy. The Execution loop below MUST
|
|
call that planner and MUST NOT reimplement the gate.
|
|
|
|
```yaml
|
|
# disk-gc report-only guests — authoritative. Keyed on guest/host, not agent.
|
|
report_only_guests:
|
|
- guest: 111
|
|
hostname: tdunna
|
|
ip: 192.168.68.129
|
|
node: storepve
|
|
reason: "Theo's box — captain ruling 2026-08-17, re-confirmed 2026-09-10"
|
|
```
|
|
|
|
### Loop
|
|
|
|
```prose
|
|
let fleet = call disk-scanner
|
|
scope: all
|
|
|
|
-- The report-only gate is IMPLEMENTED IN scripts/disk-gc-plan.py and MUST NOT be
|
|
-- reimplemented here. That planner reads the contract's `report_only_guests` YAML block
|
|
-- and matches on guest id OR hostname OR IP, so the tested gate is the executed gate.
|
|
let plan = call disk-gc-plan
|
|
fleet: fleet
|
|
|
|
for row in plan:
|
|
if row.action == "report-only":
|
|
-- Excluded guest: alert only. No gc-executor call is constructed for it, at any level.
|
|
call alerter
|
|
threat: row
|
|
result: { action: "report-only", reason: row.reason }
|
|
else:
|
|
let result = call gc-executor
|
|
ct: row.target
|
|
level: row.level
|
|
strategy: lookup-gc-strategy(row.target)
|
|
|
|
call alerter
|
|
threat: row
|
|
result: result
|
|
|
|
call summary-reporter
|
|
fleet: fleet
|
|
plan: plan
|
|
```
|
|
|
|
## GC SCHEDULE (PBS datastore only)
|
|
|
|
The PBS GC schedule is defined in ONE authoritative place: `/etc/cron.d/pbs-gc` on storepve.
|
|
The schedule is `0 20 * * *` (20:00 LOCAL = 00:00 UTC, since host timezone is America/New_York).
|
|
This applies **only** to the PBS datastore (`/tank/pbs-backup`), NOT to media volumes.
|
|
Media volumes (/media/*) are report-only at all threat levels.
|
|
|
|
The cron runs `/usr/local/bin/pbs-gc.sh` which executes:
|
|
```bash
|
|
proxmox-backup-manager garbage-collection start storepve-datastore
|
|
```
|
|
This is NOT a `prune` operation; it is a GC pass that reclaims unreferenced chunks.
|
|
There is no `--keep-daily` flag; retention is governed by jobs.cfg (keep-daily=35).
|
|
The GC does not touch media volumes or any other filesystem.
|
|
|
|
## GC Strategies by Host Type
|
|
|
|
### Docker Hosts (kagentz 105, syslog-api 116, docker-vm 109, amdpve .15)
|
|
|
|
```bash
|
|
# Phase 1: Safe prune (won't touch running containers' images)
|
|
docker system prune -f
|
|
|
|
# Phase 2: Aggressive (if still > 85% after Phase 1)
|
|
docker system prune -a --force
|
|
|
|
# Phase 3: Emergency (if still > 95%)
|
|
docker system prune -a --force --volumes
|
|
docker builder prune --all --force
|
|
|
|
# Verification after each phase
|
|
df -h /
|
|
docker system df
|
|
```
|
|
|
|
### Non-Docker CTs
|
|
|
|
```bash
|
|
# Package cache
|
|
apt-get clean
|
|
apt-get autoremove --yes
|
|
|
|
# Log rotation
|
|
journalctl --vacuum-size=100M
|
|
find /var/log -type f -name "*.log" -mtime +30 -delete
|
|
|
|
# Temp files
|
|
find /tmp -type f -mtime +7 -delete
|
|
find /var/tmp -type f -mtime +30 -delete
|
|
|
|
# Snap (if installed)
|
|
snap list --all | awk '/disabled/ {print $1, $3}' | while read snap rev; do
|
|
snap remove "$snap" --revision="$rev"
|
|
done
|
|
```
|
|
|
|
### Special Cases
|
|
|
|
| CT | Special GC |
|
|
|----|-----------|
|
|
| 116 (syslog-api) | Prometheus retention: check `--storage.tsdb.retention.time` |
|
|
| 100 (abiba) | Go module cache: `go clean -cache -modcache` if > 500MB |
|
|
| 106 (ra-h-os) | Check relay DB size, enforce TTL |
|
|
| 109 (docker-vm) | Check `/media/storage` and `/media/mediastore` mounts first |
|
|
|
|
## Alert Templates
|
|
|
|
### HOST-WARN / HOST-AMBER / HOST-RED (host filesystems, report-only)
|
|
```
|
|
⚠️ Host Disk — {hostname} {volume} ({volume_type}): {pct}% ({used}/{total}, {free} free) -> {level}
|
|
Action: {volume_type-specific action}
|
|
```
|
|
|
|
### AMBER (75-84%)
|
|
```
|
|
⚠️ Disk GC — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
|
|
Next scheduled GC will attempt cleanup. No immediate action needed.
|
|
```
|
|
|
|
### RED (85-94%)
|
|
```
|
|
🚨 Disk Threat — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
|
|
GC executed: reclaimed {reclaimed}G. New usage: {new_pct}%.
|
|
Status: {resolved|still elevated — {reason}}
|
|
```
|
|
|
|
### CRITICAL (≥95%)
|
|
```
|
|
🔥 CRITICAL Disk — {hostname} (CT {id}) at {pct}% ({used}G/{total}G)
|
|
Emergency GC: reclaimed {reclaimed}G. New usage: {new_pct}%.
|
|
{service_status} — {manual_action if needed}
|
|
@Kwame — container nearly full, resolved: {yes_no}
|
|
```
|
|
|
|
## Incident Log: 2026-07-04 — kagentz docker bloat
|
|
|
|
### Discovery
|
|
Kwame noticed kagentz was full, asked Abiba to investigate.
|
|
|
|
### Diagnosis
|
|
- CT 105 (kagentz, amdpve): 49G used / 59G total (87%)
|
|
- `/var/lib` = 55G (93% of disk)
|
|
- Docker images: 7 total, 5 dangling, 47.83GB disk, 34.98GB reclaimable
|
|
- 1 stopped container, 1 unused network, 15 build cache layers
|
|
|
|
### Resolution
|
|
```
|
|
docker system prune -a --force
|
|
→ Reclaimed 35.67GB
|
|
→ Post: 16G / 59G (27%), 41G free
|
|
→ 5 dangling images removed (kagentz-bridge + old builds)
|
|
→ 1 stopped container removed
|
|
→ 1 unused network removed (kagentz-bridge_default)
|
|
→ 15 build cache layers removed
|
|
```
|
|
|
|
### Root Cause
|
|
Agent Zero's iterative development pattern (rebuild kagentz-bridge) creates
|
|
dangling images and orphaned build cache. No automated GC was in place.
|
|
|
|
### Preventive Measures
|
|
- This contract now runs disk GC fleet-wide every 6 hours
|
|
- Docker hosts get `docker system prune` on amber, `-a --force` on red
|
|
- kagentz is flagged as HIGH risk due to development activity
|
|
- amdpve is flagged for Docker bloat monitoring — abandoned build images accumulate
|
|
```
|
|
|
|
## Incident Log: 2026-07-09 — amdpve docker bloat
|
|
|
|
### Discovery
|
|
Scheduled fleet disk scan across all 20 Proxmox guests (17 LXC via `pct-run`, 3 QEMU VMs via direct SSH) + 3 GPU bare-metal hosts.
|
|
> **Report-only gate applies to this scan:** CT 111 (`tdunna`, 192.168.68.129) is alerted but never
|
|
> garbage-collected at any level.
|
|
amdpve (.15) flagged at 78% (AMBER threshold: 75%).
|
|
|
|
### Diagnosis
|
|
- amdpve (.15, Strix Halo host): 70G used / 94G total (78%)
|
|
- Docker images: 1 image, 0 containers running, 15.47GB (100% reclaimable)
|
|
- Image: `llama-strix-hip:latest` — abandoned ROCm/HIP Docker build from 7 days ago
|
|
- Root cause: Strix Halo migrated from Docker-based HIP path to bare-metal Vulkan
|
|
(`/root/llama.cpp/build-vk/`) but the old Docker image was never cleaned up
|
|
- Not a running service — zero containers, zero active volumes
|
|
|
|
### Resolution
|
|
```
|
|
docker system prune -a --force
|
|
→ Reclaimed 11.56GB
|
|
→ Post: 55G / 94G (62%), 35G free
|
|
→ 1 image removed (llama-strix-hip:latest, 15.5GB)
|
|
→ 8 build cache layers removed
|
|
→ amdpve now GREEN
|
|
```
|
|
|
|
### Root Cause
|
|
Technology migration (Docker HIP → bare-metal Vulkan) left orphaned build
|
|
artifacts. Docker on amdpve serves no running purpose — it's only used for
|
|
one-off GPU builds. No automated post-migration cleanup was in place.
|
|
|
|
### Preventive Measures
|
|
- amdpve added to Docker GC scan list
|
|
- Post-migration cleanup step added: after any GPU backend migration, prune
|
|
the old backend's Docker images within 24 hours
|
|
- Contract now scans GPU bare-metal hosts alongside CTs
|
|
- Access via `pct-run` script for all CTs (no hardcoded IPs)
|
|
|
|
## Access Matrix (verified against `pvesh get /cluster/resources` 2026-09-12)
|
|
|
|
### Guest Access (via `pct-run` — CT id only, node resolved by `scripts/pct-run.sh`)
|
|
| Guest | Name | Node | Type | Status |
|
|
|------|------|------|------|--------|
|
|
| 100 | abiba | minipve | lxc | ✅ reachable (probed via pct-run like any other guest; no local shortcut) |
|
|
| 102 | adguard | minipve | lxc | ✅ reachable |
|
|
| 104 | authentik | minipve | lxc | ✅ reachable |
|
|
| 105 | kagentz | **amdpve** | lxc | ✅ reachable (was documented as minipve — corrected) |
|
|
| 106 | ra-h-os | storepve | lxc | ✅ reachable |
|
|
| 107 | pbs | storepve | lxc | ✅ reachable |
|
|
| 108 | media | storepve | lxc | ✅ reachable |
|
|
| 110 | gitea | minipve | lxc | ✅ reachable |
|
|
| 111 | tdunna | **storepve** | lxc | ⛔ **REPORT-ONLY** (192.168.68.129, Theo's box — no GC at any level) |
|
|
| 112 | tanko | amdpve | lxc | ✅ reachable |
|
|
| 113 | baggy | amdpve | lxc | ✅ reachable |
|
|
| 115 | scottdenya | amdpve | lxc | ✅ reachable |
|
|
| 116 | syslog-api | minipve | lxc | ✅ reachable |
|
|
| 117 | zulip | storepve | lxc | ✅ reachable |
|
|
| 118 | jdownloader | storepve | lxc | ✅ reachable |
|
|
| 119 | infisical-vault | minipve | lxc | ✅ reachable |
|
|
| 120 | adguard2 | amdpve | lxc | ✅ reachable |
|
|
|
|
### QEMU VMs (via direct SSH)
|
|
| VM | Name | Node | IP | Status |
|
|
|----|------|------|-----|--------|
|
|
| 101 | llm-gpu (workload now bare metal .8) | acerpve | — | ✅ reachable |
|
|
| 103 | ocu-llm (workload now bare metal .110) | ocupve | — | ✅ reachable |
|
|
| 109 | docker-vm | storepve | 192.168.68.7 | ✅ reachable |
|
|
|
|
### GPU Bare Metal (via direct SSH)
|
|
| Host | IP | GPU | Status |
|
|
|------|-----|-----|--------|
|
|
| llm-gpu | 192.168.68.8 | RTX 3090 | ✅ reachable |
|
|
| ocu-llm | 192.168.68.110 | RTX 5070 | ✅ reachable |
|
|
| amdpve | 192.168.68.15 | Strix Halo | ✅ reachable |
|
|
|
|
> **Fleet count:** 20 Proxmox guests (17 LXC + 3 QEMU VMs) + 3 GPU bare-metal hosts. Corrected
|
|
> 2026-09-12: CT 105 → amdpve, CT 111 → storepve, and guests 118/119/120 were missing.
|
|
>
|
|
> **CT 100 probe gap (folded in):** CT 100 previously reported "unreachable (not reported)" every
|
|
> run. Root cause is the same stale access layer: `pct-run` resolves the guest's node from its map,
|
|
> and the map/contract must reflect `pvesh /cluster/resources`. Verified working from inside CT 100:
|
|
> `scripts/pct-run.sh 100 "df -P / | tail -1"` → `23% /`. Probe CT 100 through `pct-run` like any
|
|
> other guest — never through a local-only path, since the scanner itself runs inside CT 100 and a
|
|
> container has no `pct` binary.
|
|
>
|
|
> **KVM VM:** CT 109 (docker-vm) is a QEMU VM, not LXC — access via SSH .7.
|
|
> **NOTE:** For kagentz (CT 105), use `ssh root@kagentz` (hostname), NOT `pct exec 105` — `pct exec 105` shows loop0 (59G) while `ssh root@kagentz` shows the real filesystem (99G). For docker-vm (CT 109), use `ssh root@192.168.68.7`, not `pct exec`.
|