Files
prose-contracts/infrastructure-monitoring.prose.md
T

10 KiB

kind, name, description, version
kind name description version
function infrastructure-monitoring Deploys Prometheus + GPU exporters + Grafana to monitor the entire inference fleet (3 GPU hosts + LiteLLM) from CT 116. GPU metrics from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics via existing /metrics Prometheus endpoint. DEPLOYMENT STATUS (2026-08-09): ✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters (via proxmox-monitor contract). Grafana at :3001, all scrape targets active. ✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on :9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09. ✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09. ⚠️ This contract is target-state aspirational — but GPU export + alerting are now as-built (verified 2026-08-09). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). 1.0.0

Architecture

GPU .8 (RTX 3090)    GPU .110 (RTX 5070)   GPU .15 (Strix Halo)
  nvidia-exporter       nvidia-exporter       amdgpu-exporter
  :9400                 :9400                 :9400
       │                     │                     │
       └─────────────────────┼─────────────────────┘
                             ▼
              ┌──────────────────────────┐
              │       Prometheus         │
              │       CT 116 :9090       │
              │                          │
              │  Scrape targets:         │
              │  • 192.168.68.8:9400     │
              │  • 192.168.68.110:9400   │
              │  • 192.168.68.15:9400    │
              │  • litellm:4000/metrics  │
              └──────────┬───────────────┘
                         │
              ┌──────────▼───────────────┐
              │        Grafana           │
              │       CT 116 :3001       │
              │                          │
              │  Preloaded dashboards:   │
              │  • GPU Fleet Overview    │
              │  • LiteLLM Proxy Stats   │
              └──────────────────────────┘

Components

1. NVIDIA GPU Exporter (hosts: .8, .110)

  • Tool: utkuozdemir/nvidia_gpu_exporter (Go binary, single static binary)
  • Listens on :9400, exposes /metrics in Prometheus format
  • Metrics: utilization, temp, VRAM, power, clock speeds, fan speed

2. AMD GPU Exporter (host: .15)

  • Custom exporter: Python script wrapping amdgpu_top --json
  • Listens on :9400, exposes /metrics in Prometheus format
  • Metrics: power (W), temp (°C), VRAM used/total, GFX clock, utilization
  • Runs as systemd service for persistence

3. Prometheus (CT 116)

  • Container: prom/prometheus:latest
  • Port: 9090 (internal Docker network)
  • Scrape interval: 15s
  • Config: /opt/monitoring/prometheus.yml
  • Storage: Docker volume prometheus-data

4. Grafana (CT 116)

  • Container: grafana/grafana:latest
  • Port: 3001 (mapped to host)
  • Data source: Prometheus at http://prometheus:9090
  • Provisioned dashboards for GPU fleet + LiteLLM
  • Accessible at http://192.168.68.116:3001

Parameters

  • gpu_nvidia_hosts: ["192.168.68.8", "192.168.68.110"]
  • gpu_amd_hosts: ["192.168.68.15"]
  • monitoring_host: "192.168.68.116"
  • prometheus_port: 9090
  • grafana_port: 3001
  • gpu_exporter_port: 9400

Requires

  • SSH access to all GPU hosts for exporter deployment
  • Docker on CT 116 for Prometheus + Grafana containers
  • Python 3 on AMD host for custom exporter
  • nvidia-smi on NVIDIA hosts

Maintains

  • All 3 GPU hosts export metrics at :9400/metrics in Prometheus format
  • Prometheus scrapes all targets every 15s
  • Grafana dashboards show real-time GPU utilization, temp, VRAM, power
  • LiteLLM metrics (requests, tokens, latency, errors) visible alongside GPU metrics
  • Stack persists across reboots (systemd for exporters, Docker restart policy)

Execution

Liveness rule (scoped)

The any-HTTP-response rule applies ONLY to unauthenticated/auth-gated endpoints, where any HTTP answer proves a listener is up: the PVE API (https://<node>:8006/api2/json/version) and LiteLLM health (/litellm/health, 301 → /litellm/health/liveliness). For those endpoints a probe is ALIVE on ANY HTTP status — 401/403 auth challenges and 3xx redirects included — and DOWN = connection refused (000) or timeout only. The PVE API legitimately answers 401 to an unauthenticated probe — that is the healthy signal, not a failure. Same scoped rule as zulip-health (Tanko) and gpu-monitor.

Probes whose success condition is specifically a bare 200 are NOT covered by the any-HTTP rule. On those — the authenticated Zulip POST and the router /health — an unexpected status (401/403 from a bad or missing credential, 5xx, or anything other than the expected 200) is an ALERT, not "alive".

check-health

RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.

# Provenance — run first; paste the absolute path into the report
pwd -P

# Zulip API health (POST ping)
source /etc/litellm-monitor.env
ZULIP_USER="abiba-bot@chat.sysloggh.net"
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}"
# Expected: 200 (HTTP 000 = unreachable/cache)

# PM2 process health
pm2 jlist
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)

# GPU exporters (may be down per DEPLOYMENT STATUS)
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"

# Router health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
# Expected: 200 (Router is up and responding)

# LiteLLM health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive

# PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring
# host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 —
# that was the stale-vantage bug this replaces (CT 116 is the monitoring host,
# not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE
# by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response);
# DOWN = connection refused (000) or timeout only.
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
  printf '%s:8006 -> %s\n' "$node" \
    "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
done
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
# A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault.

# Prometheus targets
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
# Expected: All targets UP (may show some down if exporters not deployed)

# Grafana health
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
# Expected: {"status":"ok","version":"..."}

# LiteLLM metrics (Prometheus endpoint)
curl -s http://192.168.68.116:4001/metrics | head -20
# Expected: Prometheus-formatted metrics output

Report format: Begin every report with the absolute path the probe executed from (pwd -P, or the script's absolute path) so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual results from each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE API and LiteLLM endpoints above: only connection-refused (000) or timeout is DOWN; empty output is a warning. For probes whose expected result is a bare 200 (the authenticated Zulip POST, router /health), flag an alert on any unexpected status (401/403/5xx) — do not summarize it as alive. A bare-200 expectation on the auth-gated PVE API (401) or LiteLLM health (301 redirect) is a stale expectation, not a fault.

Phase 1: GPU Exporters

NVIDIA (.8 and .110):

  1. Download nvidia_gpu_exporter binary
  2. Create systemd service nvidia-gpu-exporter.service
  3. Start and enable

AMD (.15):

  1. Create Python exporter script at /opt/amdgpu-exporter/exporter.py
  2. Parses amdgpu_top --json -d 1000 output
  3. Exposes key metrics at :9400/metrics via Python http.server
  4. Create systemd service
  5. Start and enable

Phase 2: Prometheus

  1. Create /opt/monitoring/ directory on CT 116
  2. Write prometheus.yml with scrape configs for all targets
  3. Add to docker-compose (or separate compose file)
  4. Start container

Phase 3: Grafana

  1. Create /opt/monitoring/grafana/ directories
  2. Provision Prometheus datasource
  3. Provision GPU fleet dashboard JSON
  4. Provision LiteLLM dashboard JSON
  5. Add to docker-compose
  6. Start container

Phase 4: Verification

  1. Verify all 3 GPU exporters return 200 at :9400/metrics
  2. Verify Prometheus targets all UP at :9090/targets
  3. Verify Grafana accessible at :3001 with dashboards
  4. Verify LiteLLM metrics flowing to Prometheus
  5. Update nginx to proxy /monitoring/ → Grafana (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard)

Verification Commands

# GPU exporters
curl -s http://192.168.68.8:9400/metrics | grep nvidia
curl -s http://192.168.68.110:9400/metrics | grep nvidia
curl -s http://192.168.68.15:9400/metrics | grep amdgpu

# Prometheus
curl -s http://192.168.68.116:9090/api/v1/targets

# Grafana
curl -s http://192.168.68.116:3001/api/health

# LiteLLM metrics (already live)
curl -s http://192.168.68.116:4001/metrics | head -20