Files
prose-contracts/proxmox-monitor.prose.md
T
root 8ff13d38f3
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
docs(proxmox): Document all six PBS GC verdict shapes
The Liveness Check section now documents ALL SIX verdict shapes exactly
as emitted by proxmox-monitor.sh:

1. ✅ PBS GC: healthy (last run Nh ago, pending-bytes: N B)
2. 🔴 PBS GC: stale (last run Nh ago, pending-bytes: N B)
3. 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)
4. 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)
5. 🔴 PBS GC: never-run (storepve-datastore not found in GC list)
6. 🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)

Fixed the quoted healthy example (line ~114) to include the pending-bytes
suffix the code now appends. Previously the prose only documented the
probe-failed shape, missing the PR's own headline cases (never-run).

Proof: grep -n 'never-run' proxmox-monitor.prose.md now returns two lines
(lines 106 and 108), documenting both never-run variants.

Branch: fix/pbs-gc-liveness-signal-20260919
2026-09-19 02:43:49 +00:00

10 KiB

kind, name, description, agent
kind name description agent
responsibility proxmox-monitor Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single instance), node_exporter (all 5 PVE nodes), and a custom docker-stats-exporter (Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx). abiba

Architecture

┌─────────────────────────────────────────────────────────────┐
│  CT 116 (syslog-api) — monitoring compose /opt/monitoring/   │
│  Prometheus :9090  →  Grafana :3001 (direct LAN, 0.0.0.0)    │
└───────┬──────────────┬──────────────┬───────────────────────┘
        │              │              │
        ▼              ▼              ▼
   pve-exporter    docker-stats    (scrapes 5x node_exporter)
   :9221           :9324
        │              │
        ▼              ▼
   PVE API         Docker API
   (amdpve .15,    (unix socket,
    cluster-wide)  10 containers)

Exporters

Exporter Host:Port Scope Notes
prometheus-pve-exporter .116:9221 (container) All 5 nodes + 14 guests + 36 storage pools Single instance, cluster-aware via amdpve API. Config /opt/monitoring/pve.yml (token monitoring@pve!prometheus, PVEAuditor role). Metric schema is label-based (id=node/amdpve, id=lxc/100).
node_exporter .5/.6/.9/.12/.15:9100 (systemd) Per-node CPU/mem/disk/net/temp Installed via apt on all 5 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool.
docker-stats-exporter .116:9324 (container) 10 Docker containers on .116 Custom (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script /opt/monitoring/docker-stats-exporter.py.

PVE API Token

  • User: monitoring@pve (cluster-replicated)
  • Role: PVEAuditor on / (read-only, whole cluster)
  • Token: monitoring@pve!prometheus — stored in Infisical vault (PROXMOX_MONITOR_TOKEN)
  • verify_ssl: false (proxmoxer uses verify_ssl, NOT verify_tls)

Grafana Dashboards (file-provisioned, folder "Syslog Fleet")

UID Title Panels Source
proxmox-cluster Proxmox Cluster Overview 16 cluster status, 5-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries
proxmox-node Proxmox Node Detail 13 per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node)
docker-containers Docker Containers 10 per-container CPU/mem/network, restarts, memory limit ratio (variable: $container)
gpu-fleet GPU Fleet 7 (existing, preserved in DB, not provisioned)
  • Dashboards built by /opt/monitoring/grafana/dashboards/build-dashboards.py → JSON in .../dashboards/json/
  • Provider config: /opt/monitoring/grafana/dashboards/dashboards.yml
  • Datasource: Prometheus uid afqpgfay4g9hce (provisioned, /opt/monitoring/grafana/datasources/prometheus.yml)
  • Edit dashboards in build-dashboards.py + re-run; allowUiUpdates: true for ad-hoc UI tweaks

Access

  • URL: http://192.168.68.116:3001/ (LAN, direct — Grafana bound to 0.0.0.0:3001)
  • Dashboards: http://192.168.68.116:3001/d/gpu-fleet, .../d/proxmox-cluster, .../d/proxmox-node, .../d/docker-containers
  • Credentials: admin / password stored in Infisical vault (GRAFANA_ADMIN_PASSWORD)
  • Grafana is NOT behind nginx — access port 3001 directly. The harness-nginx /grafana/ sub-path route was tried and reverted (broke the existing :3001 URL and gpu-fleet path). Do not re-add GF_SERVER_SERVE_FROM_SUB_PATH or an nginx /grafana/ route.
  • grafana compose port mapping: "3001:3000" (0.0.0.0, not 127.0.0.1)

Configuration Files

File Host Purpose
/opt/monitoring/docker-compose.yml .116 monitoring stack (prometheus, grafana, pve-exporter, docker-stats)
/opt/monitoring/prometheus.yml .116 6 scrape jobs (3 GPU, pve, node x5, docker-stats)
/opt/monitoring/pve.yml .116 PVE API credentials (chmod 644, contains token)
/opt/monitoring/docker-stats-exporter.py .116 custom Docker metrics exporter
/opt/monitoring/grafana/dashboards/build-dashboards.py .116 dashboard JSON generator
/opt/monitoring/grafana/dashboards/json/*.json .116 provisioned dashboard definitions
/opt/monitoring/grafana/datasources/prometheus.yml .116 datasource provisioning
/etc/default/prometheus-node-exporter .5/.6/.9/.12/.15 node_exporter collector config

Cluster "Tabiri" — 5 Nodes

Node IP Role
ocupve 192.168.68.5 PVE
storepve 192.168.68.6 PVE
acerpve 192.168.68.9 PVE (hosts llm-gpu qemu/101)
minipve 192.168.68.12 PVE
amdpve 192.168.68.15 PVE + Strix Halo LLM (strix-moe)

PBS GC (Proxmox Backup Server)

Schedule

Cron 0 20 * * * on the storepve HOST (192.168.68.6) = 20:00 America/New_York local = 00:00 UTC.

NOTE: The closed PR #116 said "20:00 UTC" — this is WRONG by four hours. Do not copy it.

What Actually Runs

Host script /usr/local/bin/pbs-gc.sh runs pct exec 107 -- proxmox-backup-manager garbage-collection start storepve-datastore.

IMPORTANT: The tool proxmox-backup-manager exists only inside CT 107 (where the PBS server runs). The storepve host has only proxmox-backup-client. This is why the job had never worked before 2026-09-19 00:00 UTC.

Datastore Location

  • Datastore: CT 107's /mnt/pbs-backup on the storepve ZFS dataset /tank/pbs-backup (pool tank, ~11T free)
  • NOT /media/easystore2 (media library, 3.7T, 96% used — separate volume)

Liveness Check

The new proxmox-monitor.sh leg checks storepve-datastore GC health:

  • Reads GC state from CT 107: pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json
  • FAILS if last-run-endtime is older than 48 hours
  • Reports age in hours and pending-bytes

All six verdict shapes (exactly as emitted by the script):

  1. Healthy (fresh GC, 0 B pending): ✅ PBS GC: healthy (last run 1h ago, pending-bytes: 0 B)

  2. Stale (GC ran >48h ago): 🔴 PBS GC: stale (last run 49h ago, pending-bytes: 1048576 B)

  3. Probe-failed: empty read (000/timeout/unreadable): 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)

  4. Probe-failed: unparseable (non-empty but invalid JSON — the "command not found" case): 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)

  5. Never-run: datastore absent (valid JSON but storepve-datastore not in list): 🔴 PBS GC: never-run (storepve-datastore not found in GC list)

  6. Never-run: no endtime (valid JSON with datastore present but last-run-endtime is null/0): 🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)

Operations

view-dashboards

Open http://192.168.68.116:3001/ → "Syslog Fleet" folder

add-dashboard

Edit build-dashboards.py, run it, docker restart harness-grafana

check-targets

curl http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job,health}'

check-health

RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.

# Prometheus health (bound to 0.0.0.0:9090 on .116)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:9090/-/healthy
# Expected: 200 (Prometheus is up and healthy)

# Grafana health (bound to 0.0.0.0:3001 on .116)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:3001/api/health
# Expected: 200 (Grafana is up and healthy)

# Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — must probe from .116 localhost)
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9324/metrics"
# Expected: 200 (docker-stats-exporter is up and responding)

# PVE exporter (bound to 127.0.0.1:9221 on .116 — must probe from .116 localhost)
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9221/metrics"
# Expected: 200 (pve-exporter is up and responding)

Report format: Begin every report with the absolute path the probe executed from (pwd -P, or the script's absolute path) so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual results from each probe. If any probe returns non-200, flag as alert.

Note: Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only) so they must be probed from .116 via SSH. Prometheus and Grafana are bound to 0.0.0.0 so they can be probed from the LAN.

restart-exporter

cd /opt/monitoring && docker compose restart pve-exporter docker-stats

rotate-pve-token

pveum user token add monitoring@pve prometheus on any PVE node → update /opt/monitoring/pve.yml → docker compose restart pve-exporter

Known Issues & Notes

  • cAdvisor abandoned: v0.51 can't resolve Docker 29 containerd image-store layerdb (/var/lib/docker/image/ only has identity-cache.db). Returns 0 named containers. Replaced by custom docker-stats-exporter.
  • Docker native metrics (/etc/docker/daemon.json metrics-addr: 127.0.0.1:9323, experimental:true) enabled but only gives engine_daemon_* (daemon-level), not per-container. Kept for daemon health.
  • PVE exporter metric schema: NOT name-prefixed. pve_cpu_usage_ratio, pve_memory_usage_bytes, pve_disk_usage_bytes, pve_uptime_seconds are GUEST-level only (24 series, id=lxc/100 etc). Node-level host metrics come from node_exporter. Storage pool usage: pve_storage_info (info only, no usage bytes — use node_filesystem_* for actual disk usage).
  • grafana piechart plugin removed from GF_INSTALL_PLUGINS (Angular, unsupported in Grafana 13).
  • Single pve-exporter points at amdpve .15 — if amdpve API is down, cluster metrics gap (other node_exporters still report host metrics). Acceptable; amdpve is primary.