Files
prose-contracts/proxmox-monitor.prose.md
T
abiba-bot c66d9c1e20 fix: probe-drift round 2 — stale monitoring expectations (health check, PVE API, GPU port 80, report provenance)
Second probe-drift correction pass after #65/#66/#68. All four legs were stale
consumer expectations, not live faults.

1. scripts/agent-health-check.py (v4)
   - abiba declared pi-only runtime (harness purge): Hermes-era gateway, config
     and wrapper legs are skipped instead of failing.
   - koby declared report_only (captain ruling 2026-08-17, Rule 17): every koby
     leg is detected and reported, never counted as a fleet failure or repaired.
   - koby's CT 111 mapping corrected to storepve (.6); the old amdpve mapping
     made `pct status 111` fail and read as ct-unreachable.
   - wrapper check no longer FAILs .env-based wrappers that legitimately never
     invoke infisical (koonimo).
   - keys load in main() (load_agent_keys) so the module is importable/testable.
   - every run prints absolute execution provenance (script + cwd), in the header
     and in --json.
   Before: 6 FAILURE(S). After: 0 failures, koby reported read-only.

2. infrastructure-monitoring.prose.md
   - PVE API probe repointed from CT 116 (no pveproxy, 000) to the five real
     nodes on https://<node>:8006/api2/json/version, all 401 = alive.
   - any-HTTP-response liveness rule added (401/3xx alive; 000/timeout = DOWN).
   - LiteLLM health documented as 301 -> /litellm/health/liveliness, not bare 200.

3. gpu-monitor.prose.md
   - GPU health probes on :8080 (or router /health/unified); bare port 80 on a
     GPU host is forbidden (no listener -> false DEGRADED).
   - router /health/unified 301 -> /gpu/gpu-data documented as alive.
   - port-discipline + liveness rule + direct-fallback execution step.

4. Report provenance (all contracts)
   - docs/AUTHORING-GUIDE.md documents the rule; scripts/prose-lint.sh enforces
     that any **Report format** contract states an absolute path (pwd -P).
   - provenance added to gpu-monitor, infrastructure-monitoring, proxmox-monitor.

Tests: tests/test_probe_drift.py (23 passed); prose-lint.sh clean + shellcheck
clean; CI frontmatter validation passes. Evidence with per-leg before/after and
absolute paths: docs/probe-drift-round2-evidence.md.
2026-09-10 01:23:54 +00:00

8.4 KiB

kind, name, description, agent
kind name description agent
responsibility proxmox-monitor Proxmox cluster + Docker monitoring via the existing Grafana/Prometheus stack on CT 116. Replaces Pulse with file-provisioned Grafana dashboards. Three exporters feed Prometheus: prometheus-pve-exporter (cluster-aware, single instance), node_exporter (all 5 PVE nodes), and a custom docker-stats-exporter (Docker 29 / containerd image-store compatible, since cAdvisor cannot resolve the layerdb). Dashboards exposed at http://192.168.68.116:3001/ (direct LAN, not behind nginx). abiba

Architecture

┌─────────────────────────────────────────────────────────────┐
│  CT 116 (syslog-api) — monitoring compose /opt/monitoring/   │
│  Prometheus :9090  →  Grafana :3001 (direct LAN, 0.0.0.0)    │
└───────┬──────────────┬──────────────┬───────────────────────┘
        │              │              │
        ▼              ▼              ▼
   pve-exporter    docker-stats    (scrapes 5x node_exporter)
   :9221           :9324
        │              │
        ▼              ▼
   PVE API         Docker API
   (amdpve .15,    (unix socket,
    cluster-wide)  10 containers)

Exporters

Exporter Host:Port Scope Notes
prometheus-pve-exporter .116:9221 (container) All 5 nodes + 14 guests + 36 storage pools Single instance, cluster-aware via amdpve API. Config /opt/monitoring/pve.yml (token monitoring@pve!prometheus, PVEAuditor role). Metric schema is label-based (id=node/amdpve, id=lxc/100).
node_exporter .5/.6/.9/.12/.15:9100 (systemd) Per-node CPU/mem/disk/net/temp Installed via apt on all 5 PVE nodes, enabled (reboot-persistent). Collectors: textfile, systemd, tcpstat, ethtool.
docker-stats-exporter .116:9324 (container) 10 Docker containers on .116 Custom (cAdvisor v0.51 incompatible with Docker 29 containerd image store — layerdb gone). Uses Docker Engine API over unix socket. Script /opt/monitoring/docker-stats-exporter.py.

PVE API Token

  • User: monitoring@pve (cluster-replicated)
  • Role: PVEAuditor on / (read-only, whole cluster)
  • Token: monitoring@pve!prometheus — stored in Infisical vault (PROXMOX_MONITOR_TOKEN)
  • verify_ssl: false (proxmoxer uses verify_ssl, NOT verify_tls)

Grafana Dashboards (file-provisioned, folder "Syslog Fleet")

UID Title Panels Source
proxmox-cluster Proxmox Cluster Overview 16 cluster status, 5-node CPU/mem/disk/load gauges, guests table, storage pools, guest CPU/mem timeseries
proxmox-node Proxmox Node Detail 13 per-node CPU per-core, memory, network, disk IO/IOPS/latency, temperature, disk space (variable: $node)
docker-containers Docker Containers 10 per-container CPU/mem/network, restarts, memory limit ratio (variable: $container)
gpu-fleet GPU Fleet 7 (existing, preserved in DB, not provisioned)
  • Dashboards built by /opt/monitoring/grafana/dashboards/build-dashboards.py → JSON in .../dashboards/json/
  • Provider config: /opt/monitoring/grafana/dashboards/dashboards.yml
  • Datasource: Prometheus uid afqpgfay4g9hce (provisioned, /opt/monitoring/grafana/datasources/prometheus.yml)
  • Edit dashboards in build-dashboards.py + re-run; allowUiUpdates: true for ad-hoc UI tweaks

Access

  • URL: http://192.168.68.116:3001/ (LAN, direct — Grafana bound to 0.0.0.0:3001)
  • Dashboards: http://192.168.68.116:3001/d/gpu-fleet, .../d/proxmox-cluster, .../d/proxmox-node, .../d/docker-containers
  • Credentials: admin / password stored in Infisical vault (GRAFANA_ADMIN_PASSWORD)
  • Grafana is NOT behind nginx — access port 3001 directly. The harness-nginx /grafana/ sub-path route was tried and reverted (broke the existing :3001 URL and gpu-fleet path). Do not re-add GF_SERVER_SERVE_FROM_SUB_PATH or an nginx /grafana/ route.
  • grafana compose port mapping: "3001:3000" (0.0.0.0, not 127.0.0.1)

Configuration Files

File Host Purpose
/opt/monitoring/docker-compose.yml .116 monitoring stack (prometheus, grafana, pve-exporter, docker-stats)
/opt/monitoring/prometheus.yml .116 6 scrape jobs (3 GPU, pve, node x5, docker-stats)
/opt/monitoring/pve.yml .116 PVE API credentials (chmod 644, contains token)
/opt/monitoring/docker-stats-exporter.py .116 custom Docker metrics exporter
/opt/monitoring/grafana/dashboards/build-dashboards.py .116 dashboard JSON generator
/opt/monitoring/grafana/dashboards/json/*.json .116 provisioned dashboard definitions
/opt/monitoring/grafana/datasources/prometheus.yml .116 datasource provisioning
/etc/default/prometheus-node-exporter .5/.6/.9/.12/.15 node_exporter collector config

Cluster "Tabiri" — 5 Nodes

Node IP Role
ocupve 192.168.68.5 PVE
storepve 192.168.68.6 PVE
acerpve 192.168.68.9 PVE (hosts llm-gpu qemu/101)
minipve 192.168.68.12 PVE
amdpve 192.168.68.15 PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe)

Operations

view-dashboards

Open http://192.168.68.116:3001/ → "Syslog Fleet" folder

add-dashboard

Edit build-dashboards.py, run it, docker restart harness-grafana

check-targets

curl http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job,health}'

check-health

RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.

# Prometheus health (bound to 0.0.0.0:9090 on .116)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:9090/-/healthy
# Expected: 200 (Prometheus is up and healthy)

# Grafana health (bound to 0.0.0.0:3001 on .116)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116:3001/api/health
# Expected: 200 (Grafana is up and healthy)

# Docker Stats exporter (bound to 127.0.0.1:9324 on .116 — must probe from .116 localhost)
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9324/metrics"
# Expected: 200 (docker-stats-exporter is up and responding)

# PVE exporter (bound to 127.0.0.1:9221 on .116 — must probe from .116 localhost)
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:9221/metrics"
# Expected: 200 (pve-exporter is up and responding)

Report format: Begin every report with the absolute path the probe executed from (pwd -P, or the script's absolute path) so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual results from each probe. If any probe returns non-200, flag as alert.

Note: Docker Stats and PVE Exporter are bound to 127.0.0.1 (localhost-only) so they must be probed from .116 via SSH. Prometheus and Grafana are bound to 0.0.0.0 so they can be probed from the LAN.

restart-exporter

cd /opt/monitoring && docker compose restart pve-exporter docker-stats

rotate-pve-token

pveum user token add monitoring@pve prometheus on any PVE node → update /opt/monitoring/pve.yml → docker compose restart pve-exporter

Known Issues & Notes

  • cAdvisor abandoned: v0.51 can't resolve Docker 29 containerd image-store layerdb (/var/lib/docker/image/ only has identity-cache.db). Returns 0 named containers. Replaced by custom docker-stats-exporter.
  • Docker native metrics (/etc/docker/daemon.json metrics-addr: 127.0.0.1:9323, experimental:true) enabled but only gives engine_daemon_* (daemon-level), not per-container. Kept for daemon health.
  • PVE exporter metric schema: NOT name-prefixed. pve_cpu_usage_ratio, pve_memory_usage_bytes, pve_disk_usage_bytes, pve_uptime_seconds are GUEST-level only (24 series, id=lxc/100 etc). Node-level host metrics come from node_exporter. Storage pool usage: pve_storage_info (info only, no usage bytes — use node_filesystem_* for actual disk usage).
  • grafana piechart plugin removed from GF_INSTALL_PLUGINS (Angular, unsupported in Grafana 13).
  • Single pve-exporter points at amdpve .15 — if amdpve API is down, cluster metrics gap (other node_exporters still report host metrics). Acceptable; amdpve is primary.