--- kind: responsibility name: proxmox-monitor description: > Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single instance), node_exporter (all 5 PVE nodes), and a custom docker-stats-exporter (Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx). agent: abiba --- ## Architecture ``` ┌─────────────────────────────────────────────────────────────┐ │ CT 116 (syslog-api) — monitoring compose /opt/monitoring/ │ │ Prometheus :9090 → Grafana :3001 (direct LAN, 0.0.0.0) │ └───────┬──────────────┬──────────────┬───────────────────────┘ │ │ │ ▼ ▼ ▼ pve-exporter docker-stats (scrapes 5x node_exporter) :9221 :9324 │ │ ▼ ▼ PVE API Docker API (amdpve .15, (unix socket, cluster-wide) 10 containers) ``` ## Exporters | Exporter | Host:Port | Scope | Notes | |----------|-----------|-------|-------| | prometheus-pve-exporter | .116:9221 (container) | All 5 nodes + 14 guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). | | node_exporter | .5/.6/.9/.12/.15:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 5 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. | | docker-stats-exporter | .116:9324 (container) | 10 Docker containers on .116 | **Custom** (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script `/opt/monitoring/docker-stats-exporter.py`. | ## PVE API Token - User: `monitoring@pve` (cluster-replicated) - Role: `PVEAuditor` on `/` (read-only, whole cluster) - Token: `monitoring@pve!prometheus` — stored in Infisical vault (`PROXMOX_MONITOR_TOKEN`) - `verify_ssl: false` (proxmoxer uses `verify_ssl`, NOT `verify_tls`) ## Grafana Dashboards (file-provisioned, folder "Syslog Fleet") | UID | Title | Panels | Source | |-----|-------|--------|--------| | proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 5-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries | | proxmox-node | Proxmox Node Detail | 13 | per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node) | | docker-containers | Docker Containers | 10 | per-container CPU/mem/network, restarts, memory limit ratio (variable: $container) | | gpu-fleet | GPU Fleet | 7 | (existing, preserved in DB, not provisioned) | - Dashboards built by `/opt/monitoring/grafana/dashboards/build-dashboards.py` → JSON in `.../dashboards/json/` - Provider config: `/opt/monitoring/grafana/dashboards/dashboards.yml` - Datasource: Prometheus uid `afqpgfay4g9hce` (provisioned, `/opt/monitoring/grafana/datasources/prometheus.yml`) - Edit dashboards in build-dashboards.py + re-run; `allowUiUpdates: true` for ad-hoc UI tweaks ## Access - **URL**: `http://192.168.68.116:3001/` (LAN, direct — Grafana bound to `0.0.0.0:3001`) - **Dashboards**: `http://192.168.68.116:3001/d/gpu-fleet`, `.../d/proxmox-cluster`, `.../d/proxmox-node`, `.../d/docker-containers` - **Credentials**: admin / password stored in Infisical vault (`GRAFANA_ADMIN_PASSWORD`) - Grafana is NOT behind nginx — access port 3001 directly. The `harness-nginx` `/grafana/` sub-path route was tried and reverted (broke the existing `:3001` URL and gpu-fleet path). Do not re-add `GF_SERVER_SERVE_FROM_SUB_PATH` or an nginx `/grafana/` route. - grafana compose port mapping: `"3001:3000"` (0.0.0.0, not 127.0.0.1) ## Configuration Files | File | Host | Purpose | |------|------|---------| | `/opt/monitoring/docker-compose.yml` | .116 | monitoring stack (prometheus, grafana, pve-exporter, docker-stats) | | `/opt/monitoring/prometheus.yml` | .116 | 6 scrape jobs (3 GPU, pve, node x5, docker-stats) | | `/opt/monitoring/pve.yml` | .116 | PVE API credentials (chmod 644, contains token) | | `/opt/monitoring/docker-stats-exporter.py` | .116 | custom Docker metrics exporter | | `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator | | `/opt/monitoring/grafana/dashboards/json/*.json` | .116 | provisioned dashboard definitions | | `/opt/monitoring/grafana/datasources/prometheus.yml` | .116 | datasource provisioning | | `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15 | node_exporter collector config | ## Cluster "Tabiri" — 5 Nodes | Node | IP | Role | |------|----|----| | ocupve | 192.168.68.5 | PVE | | storepve | 192.168.68.6 | PVE | | acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) | | minipve | 192.168.68.12 | PVE | | amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) | ## PBS GC (Proxmox Backup Server) ### Schedule Cron `0 20 * * *` on the **storepve HOST** (192.168.68.6) = 20:00 America/New_York local = **00:00 UTC**. **NOTE**: The closed PR #116 said "20:00 UTC" — this is WRONG by four hours. Do not copy it. ### What Actually Runs Host script `/usr/local/bin/pbs-gc.sh` runs `pct exec 107 -- proxmox-backup-manager garbage-collection start storepve-datastore`. **IMPORTANT**: The tool `proxmox-backup-manager` exists only inside CT 107 (where the PBS server runs). The storepve host has only `proxmox-backup-client`. This is why the job had never worked before 2026-09-19 00:00 UTC. ### Datastore Location - **Datastore**: CT 107's `/mnt/pbs-backup` on the storepve ZFS dataset `/tank/pbs-backup` (pool `tank`, ~11T free) - **NOT** `/media/easystore2` (media library, 3.7T, 96% used — separate volume) ### Liveness Check The new `proxmox-monitor.sh` leg checks storepve-datastore GC health: - Reads GC state from CT 107: `pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json` - **FAILS** if `last-run-endtime` is older than 48 hours - Reports age in hours and pending-bytes **All six verdict shapes** (exactly as emitted by the script): 1. **Healthy** (fresh GC, 0 B pending): `✅ PBS GC: healthy (last run 1h ago, pending-bytes: 0 B)` 2. **Stale** (GC ran >48h ago): `🔴 PBS GC: stale (last run 49h ago, pending-bytes: 1048576 B)` 3. **Probe-failed: empty read** (000/timeout/unreadable): `🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)` 4. **Probe-failed: unparseable** (non-empty but invalid JSON — the "command not found" case): `🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)` 5. **Never-run: datastore absent** (valid JSON but storepve-datastore not in list): `🔴 PBS GC: never-run (storepve-datastore not found in GC list)` 6. **Never-run: no endtime** (valid JSON with datastore present but last-run-endtime is null/0): `🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)` ## Operations ### view-dashboards Open `http://192.168.68.116:3001/` → "Syslog Fleet" folder ### add-dashboard Edit `build-dashboards.py`, run it, `docker restart harness-grafana` ### check-targets `curl http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job,health}'` ### check-health **RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** ```bash # Prometheus health (bound to 0.0.0.0:9090 on .116) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:9090/-/healthy # Expected: 200 (Prometheus is up and healthy) # Grafana health (bound to 0.0.0.0:3001 on .116) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:3001/api/health # Expected: 200 (Grafana is up and healthy) # Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — must probe from .116 localhost) ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9324/metrics" # Expected: 200 (docker-stats-exporter is up and responding) # PVE exporter (bound to 127.0.0.1:9221 on .116 — must probe from .116 localhost) ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9221/metrics" # Expected: 200 (pve-exporter is up and responding) ``` **Report format**: Begin every report with the **absolute path the probe executed from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual results from each probe. If any probe returns non-200, flag as alert. **Note**: Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only) so they must be probed from .116 via SSH. Prometheus and Grafana are bound to 0.0.0.0 so they can be probed from the LAN. ### restart-exporter `cd /opt/monitoring && docker compose restart pve-exporter docker-stats` ### rotate-pve-token `pveum user token add monitoring@pve prometheus` on any PVE node → update `/opt/monitoring/pve.yml` → `docker compose restart pve-exporter` ## Known Issues & Notes - **cAdvisor abandoned**: v0.51 can't resolve Docker 29 containerd image-store layerdb (`/var/lib/docker/image/` only has `identity-cache.db`). Returns 0 named containers. Replaced by custom docker-stats-exporter. - **Docker native metrics** (`/etc/docker/daemon.json` `metrics-addr: 127.0.0.1:9323`, experimental:true) enabled but only gives `engine_daemon_*` (daemon-level), not per-container. Kept for daemon health. - **PVE exporter metric schema**: NOT name-prefixed. `pve_cpu_usage_ratio`, `pve_memory_usage_bytes`, `pve_disk_usage_bytes`, `pve_uptime_seconds` are GUEST-level only (24 series, `id=lxc/100` etc). Node-level host metrics come from node_exporter. Storage pool usage: `pve_storage_info` (info only, no usage bytes — use node_filesystem_* for actual disk usage). - **grafana piechart plugin removed** from `GF_INSTALL_PLUGINS` (Angular, unsupported in Grafana 13). - **Single pve-exporter points at amdpve .15** — if amdpve API is down, cluster metrics gap (other node_exporters still report host metrics). Acceptable; amdpve is primary.