PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The Liveness Check section now documents ALL SIX verdict shapes exactly as emitted by proxmox-monitor.sh: 1. ✅ PBS GC: healthy (last run Nh ago, pending-bytes: N B) 2. 🔴 PBS GC: stale (last run Nh ago, pending-bytes: N B) 3. 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000) 4. 🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON) 5. 🔴 PBS GC: never-run (storepve-datastore not found in GC list) 6. 🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime) Fixed the quoted healthy example (line ~114) to include the pending-bytes suffix the code now appends. Previously the prose only documented the probe-failed shape, missing the PR's own headline cases (never-run). Proof: grep -n 'never-run' proxmox-monitor.prose.md now returns two lines (lines 106 and 108), documenting both never-run variants. Branch: fix/pbs-gc-liveness-signal-20260919
187 lines
10 KiB
Markdown
187 lines
10 KiB
Markdown
---
|
|
kind: responsibility
|
|
name: proxmox-monitor
|
|
description: >
|
|
Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack
|
|
on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three
|
|
exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single
|
|
instance), node_exporter (all 5 PVE nodes), and a custom docker-stats-exporter
|
|
(Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve
|
|
the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx).
|
|
agent: abiba
|
|
---
|
|
|
|
## Architecture
|
|
|
|
```
|
|
┌─────────────────────────────────────────────────────────────┐
|
|
│ CT 116 (syslog-api) — monitoring compose /opt/monitoring/ │
|
|
│ Prometheus :9090 → Grafana :3001 (direct LAN, 0.0.0.0) │
|
|
└───────┬──────────────┬──────────────┬───────────────────────┘
|
|
│ │ │
|
|
▼ ▼ ▼
|
|
pve-exporter docker-stats (scrapes 5x node_exporter)
|
|
:9221 :9324
|
|
│ │
|
|
▼ ▼
|
|
PVE API Docker API
|
|
(amdpve .15, (unix socket,
|
|
cluster-wide) 10 containers)
|
|
```
|
|
|
|
## Exporters
|
|
|
|
| Exporter | Host:Port | Scope | Notes |
|
|
|----------|-----------|-------|-------|
|
|
| prometheus-pve-exporter | .116:9221 (container) | All 5 nodes + 14 guests + 36 storage pools | Single instance, cluster-aware via amdpve API. Config `/opt/monitoring/pve.yml` (token `monitoring@pve!prometheus`, PVEAuditor role). Metric schema is label-based (`id=node/amdpve`, `id=lxc/100`). |
|
|
| node_exporter | .5/.6/.9/.12/.15:9100 (systemd) | Per-node CPU/mem/disk/net/temp | Installed via apt on all 5 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool. |
|
|
| docker-stats-exporter | .116:9324 (container) | 10 Docker containers on .116 | **Custom** (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script `/opt/monitoring/docker-stats-exporter.py`. |
|
|
|
|
## PVE API Token
|
|
|
|
- User: `monitoring@pve` (cluster-replicated)
|
|
- Role: `PVEAuditor` on `/` (read-only, whole cluster)
|
|
- Token: `monitoring@pve!prometheus` — stored in Infisical vault (`PROXMOX_MONITOR_TOKEN`)
|
|
- `verify_ssl: false` (proxmoxer uses `verify_ssl`, NOT `verify_tls`)
|
|
|
|
## Grafana Dashboards (file-provisioned, folder "Syslog Fleet")
|
|
|
|
| UID | Title | Panels | Source |
|
|
|-----|-------|--------|--------|
|
|
| proxmox-cluster | Proxmox Cluster Overview | 16 | cluster status, 5-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries |
|
|
| proxmox-node | Proxmox Node Detail | 13 | per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node) |
|
|
| docker-containers | Docker Containers | 10 | per-container CPU/mem/network, restarts, memory limit ratio (variable: $container) |
|
|
| gpu-fleet | GPU Fleet | 7 | (existing, preserved in DB, not provisioned) |
|
|
|
|
- Dashboards built by `/opt/monitoring/grafana/dashboards/build-dashboards.py` → JSON in `.../dashboards/json/`
|
|
- Provider config: `/opt/monitoring/grafana/dashboards/dashboards.yml`
|
|
- Datasource: Prometheus uid `afqpgfay4g9hce` (provisioned, `/opt/monitoring/grafana/datasources/prometheus.yml`)
|
|
- Edit dashboards in build-dashboards.py + re-run; `allowUiUpdates: true` for ad-hoc UI tweaks
|
|
|
|
## Access
|
|
|
|
- **URL**: `http://192.168.68.116:3001/` (LAN, direct — Grafana bound to `0.0.0.0:3001`)
|
|
- **Dashboards**: `http://192.168.68.116:3001/d/gpu-fleet`, `.../d/proxmox-cluster`, `.../d/proxmox-node`, `.../d/docker-containers`
|
|
- **Credentials**: admin / password stored in Infisical vault (`GRAFANA_ADMIN_PASSWORD`)
|
|
- Grafana is NOT behind nginx — access port 3001 directly. The `harness-nginx` `/grafana/` sub-path route was tried and reverted (broke the existing `:3001` URL and gpu-fleet path). Do not re-add `GF_SERVER_SERVE_FROM_SUB_PATH` or an nginx `/grafana/` route.
|
|
- grafana compose port mapping: `"3001:3000"` (0.0.0.0, not 127.0.0.1)
|
|
|
|
## Configuration Files
|
|
|
|
| File | Host | Purpose |
|
|
|------|------|---------|
|
|
| `/opt/monitoring/docker-compose.yml` | .116 | monitoring stack (prometheus, grafana, pve-exporter, docker-stats) |
|
|
| `/opt/monitoring/prometheus.yml` | .116 | 6 scrape jobs (3 GPU, pve, node x5, docker-stats) |
|
|
| `/opt/monitoring/pve.yml` | .116 | PVE API credentials (chmod 644, contains token) |
|
|
| `/opt/monitoring/docker-stats-exporter.py` | .116 | custom Docker metrics exporter |
|
|
| `/opt/monitoring/grafana/dashboards/build-dashboards.py` | .116 | dashboard JSON generator |
|
|
| `/opt/monitoring/grafana/dashboards/json/*.json` | .116 | provisioned dashboard definitions |
|
|
| `/opt/monitoring/grafana/datasources/prometheus.yml` | .116 | datasource provisioning |
|
|
| `/etc/default/prometheus-node-exporter` | .5/.6/.9/.12/.15 | node_exporter collector config |
|
|
|
|
## Cluster "Tabiri" — 5 Nodes
|
|
|
|
| Node | IP | Role |
|
|
|------|----|----|
|
|
| ocupve | 192.168.68.5 | PVE |
|
|
| storepve | 192.168.68.6 | PVE |
|
|
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
|
|
| minipve | 192.168.68.12 | PVE |
|
|
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
|
|
|
|
## PBS GC (Proxmox Backup Server)
|
|
|
|
### Schedule
|
|
Cron `0 20 * * *` on the **storepve HOST** (192.168.68.6) = 20:00 America/New_York local = **00:00 UTC**.
|
|
|
|
**NOTE**: The closed PR #116 said "20:00 UTC" — this is WRONG by four hours. Do not copy it.
|
|
|
|
### What Actually Runs
|
|
Host script `/usr/local/bin/pbs-gc.sh` runs `pct exec 107 -- proxmox-backup-manager garbage-collection start storepve-datastore`.
|
|
|
|
**IMPORTANT**: The tool `proxmox-backup-manager` exists only inside CT 107 (where the PBS server runs). The storepve host has only `proxmox-backup-client`. This is why the job had never worked before 2026-09-19 00:00 UTC.
|
|
|
|
### Datastore Location
|
|
- **Datastore**: CT 107's `/mnt/pbs-backup` on the storepve ZFS dataset `/tank/pbs-backup` (pool `tank`, ~11T free)
|
|
- **NOT** `/media/easystore2` (media library, 3.7T, 96% used — separate volume)
|
|
|
|
### Liveness Check
|
|
The new `proxmox-monitor.sh` leg checks storepve-datastore GC health:
|
|
- Reads GC state from CT 107: `pct exec 107 -- proxmox-backup-manager garbage-collection list --output-format json`
|
|
- **FAILS** if `last-run-endtime` is older than 48 hours
|
|
- Reports age in hours and pending-bytes
|
|
|
|
**All six verdict shapes** (exactly as emitted by the script):
|
|
|
|
1. **Healthy** (fresh GC, 0 B pending):
|
|
`✅ PBS GC: healthy (last run 1h ago, pending-bytes: 0 B)`
|
|
|
|
2. **Stale** (GC ran >48h ago):
|
|
`🔴 PBS GC: stale (last run 49h ago, pending-bytes: 1048576 B)`
|
|
|
|
3. **Probe-failed: empty read** (000/timeout/unreadable):
|
|
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (expected JSON, got 000)`
|
|
|
|
4. **Probe-failed: unparseable** (non-empty but invalid JSON — the "command not found" case):
|
|
`🔴 PBS GC: probe-failed: storepve:192.168.68.6 (unparseable JSON)`
|
|
|
|
5. **Never-run: datastore absent** (valid JSON but storepve-datastore not in list):
|
|
`🔴 PBS GC: never-run (storepve-datastore not found in GC list)`
|
|
|
|
6. **Never-run: no endtime** (valid JSON with datastore present but last-run-endtime is null/0):
|
|
`🔴 PBS GC: never-run (storepve-datastore has no last-run-endtime)`
|
|
|
|
## Operations
|
|
|
|
### view-dashboards
|
|
Open `http://192.168.68.116:3001/` → "Syslog Fleet" folder
|
|
|
|
### add-dashboard
|
|
Edit `build-dashboards.py`, run it, `docker restart harness-grafana`
|
|
|
|
### check-targets
|
|
`curl http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job,health}'`
|
|
|
|
### check-health
|
|
|
|
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.**
|
|
|
|
```bash
|
|
# Prometheus health (bound to 0.0.0.0:9090 on .116)
|
|
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:9090/-/healthy
|
|
# Expected: 200 (Prometheus is up and healthy)
|
|
|
|
# Grafana health (bound to 0.0.0.0:3001 on .116)
|
|
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:3001/api/health
|
|
# Expected: 200 (Grafana is up and healthy)
|
|
|
|
# Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — must probe from .116 localhost)
|
|
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9324/metrics"
|
|
# Expected: 200 (docker-stats-exporter is up and responding)
|
|
|
|
# PVE exporter (bound to 127.0.0.1:9221 on .116 — must probe from .116 localhost)
|
|
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9221/metrics"
|
|
# Expected: 200 (pve-exporter is up and responding)
|
|
```
|
|
|
|
**Report format**: Begin every report with the **absolute path the probe executed
|
|
from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is
|
|
distinguishable from a real fault at read time. Summarize actual results from
|
|
each probe. If any probe returns non-200, flag as alert.
|
|
|
|
**Note**: Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only) so they must be probed from .116 via SSH. Prometheus and Grafana are bound to 0.0.0.0 so they can be probed from the LAN.
|
|
|
|
### restart-exporter
|
|
`cd /opt/monitoring && docker compose restart pve-exporter docker-stats`
|
|
|
|
### rotate-pve-token
|
|
`pveum user token add monitoring@pve prometheus` on any PVE node → update `/opt/monitoring/pve.yml` → `docker compose restart pve-exporter`
|
|
|
|
## Known Issues & Notes
|
|
|
|
- **cAdvisor abandoned**: v0.51 can't resolve Docker 29 containerd image-store layerdb (`/var/lib/docker/image/` only has `identity-cache.db`). Returns 0 named containers. Replaced by custom docker-stats-exporter.
|
|
- **Docker native metrics** (`/etc/docker/daemon.json` `metrics-addr: 127.0.0.1:9323`, experimental:true) enabled but only gives `engine_daemon_*` (daemon-level), not per-container. Kept for daemon health.
|
|
- **PVE exporter metric schema**: NOT name-prefixed. `pve_cpu_usage_ratio`, `pve_memory_usage_bytes`, `pve_disk_usage_bytes`, `pve_uptime_seconds` are GUEST-level only (24 series, `id=lxc/100` etc). Node-level host metrics come from node_exporter. Storage pool usage: `pve_storage_info` (info only, no usage bytes — use node_filesystem_* for actual disk usage).
|
|
- **grafana piechart plugin removed** from `GF_INSTALL_PLUGINS` (Angular, unsupported in Grafana 13).
|
|
- **Single pve-exporter points at amdpve .15** — if amdpve API is down, cluster metrics gap (other node_exporters still report host metrics). Acceptable; amdpve is primary.
|