10 KiB
kind, name, description, version
| kind | name | description | version |
|---|---|---|---|
| function | infrastructure-monitoring | Deploys Prometheus + GPU exporters + Grafana to monitor the entire inference fleet (3 GPU hosts + LiteLLM) from CT 116. GPU metrics from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics via existing /metrics Prometheus endpoint. DEPLOYMENT STATUS (2026-08-09): ✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters (via proxmox-monitor contract). Grafana at :3001, all scrape targets active. ✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on :9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09. ✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09. ⚠️ This contract is target-state aspirational — but GPU export + alerting are now as-built (verified 2026-08-09). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). | 1.0.0 |
Architecture
GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
nvidia-exporter nvidia-exporter amdgpu-exporter
:9400 :9400 :9400
│ │ │
└─────────────────────┼─────────────────────┘
▼
┌──────────────────────────┐
│ Prometheus │
│ CT 116 :9090 │
│ │
│ Scrape targets: │
│ • 192.168.68.8:9400 │
│ • 192.168.68.110:9400 │
│ • 192.168.68.15:9400 │
│ • litellm:4000/metrics │
└──────────┬───────────────┘
│
┌──────────▼───────────────┐
│ Grafana │
│ CT 116 :3001 │
│ │
│ Preloaded dashboards: │
│ • GPU Fleet Overview │
│ • LiteLLM Proxy Stats │
└──────────────────────────┘
Components
1. NVIDIA GPU Exporter (hosts: .8, .110)
- Tool:
utkuozdemir/nvidia_gpu_exporter(Go binary, single static binary) - Listens on
:9400, exposes/metricsin Prometheus format - Metrics: utilization, temp, VRAM, power, clock speeds, fan speed
2. AMD GPU Exporter (host: .15)
- Custom exporter: Python script wrapping
amdgpu_top --json - Listens on
:9400, exposes/metricsin Prometheus format - Metrics: power (W), temp (°C), VRAM used/total, GFX clock, utilization
- Runs as systemd service for persistence
3. Prometheus (CT 116)
- Container:
prom/prometheus:latest - Port:
9090(internal Docker network) - Scrape interval: 15s
- Config:
/opt/monitoring/prometheus.yml - Storage: Docker volume
prometheus-data
4. Grafana (CT 116)
- Container:
grafana/grafana:latest - Port:
3001(mapped to host) - Data source: Prometheus at
http://prometheus:9090 - Provisioned dashboards for GPU fleet + LiteLLM
- Accessible at
http://192.168.68.116:3001
Parameters
- gpu_nvidia_hosts: ["192.168.68.8", "192.168.68.110"]
- gpu_amd_hosts: ["192.168.68.15"]
- monitoring_host: "192.168.68.116"
- prometheus_port: 9090
- grafana_port: 3001
- gpu_exporter_port: 9400
Requires
- SSH access to all GPU hosts for exporter deployment
- Docker on CT 116 for Prometheus + Grafana containers
- Python 3 on AMD host for custom exporter
- nvidia-smi on NVIDIA hosts
Maintains
- All 3 GPU hosts export metrics at :9400/metrics in Prometheus format
- Prometheus scrapes all targets every 15s
- Grafana dashboards show real-time GPU utilization, temp, VRAM, power
- LiteLLM metrics (requests, tokens, latency, errors) visible alongside GPU metrics
- Stack persists across reboots (systemd for exporters, Docker restart policy)
Execution
Liveness rule (scoped)
The any-HTTP-response rule applies ONLY to unauthenticated/auth-gated endpoints,
where any HTTP answer proves a listener is up: the PVE API
(https://<node>:8006/api2/json/version) and LiteLLM health
(/litellm/health, 301 → /litellm/health/liveliness). For those endpoints a
probe is ALIVE on ANY HTTP status — 401/403 auth challenges and 3xx
redirects included — and DOWN = connection refused (000) or timeout only.
The PVE API legitimately answers 401 to an unauthenticated probe — that is the
healthy signal, not a failure. Same scoped rule as zulip-health (Tanko) and
gpu-monitor.
Probes whose success condition is specifically a bare 200 are NOT covered by
the any-HTTP rule. On those — the authenticated Zulip POST and the router
/health — an unexpected status (401/403 from a bad or missing credential,
5xx, or anything other than the expected 200) is an ALERT, not "alive".
check-health
RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.
# Provenance — run first; paste the absolute path into the report
pwd -P
# Zulip API health (POST ping)
source /etc/litellm-monitor.env
ZULIP_USER="abiba-bot@chat.sysloggh.net"
curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}"
# Expected: 200 (HTTP 000 = unreachable/cache)
# PM2 process health
pm2 jlist
# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service)
# GPU exporters (may be down per DEPLOYMENT STATUS)
curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL"
curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL"
# Router health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
# Expected: 200 (Router is up and responding)
# LiteLLM health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive
# PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring
# host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 —
# that was the stale-vantage bug this replaces (CT 116 is the monitoring host,
# not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE
# by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response);
# DOWN = connection refused (000) or timeout only.
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
printf '%s:8006 -> %s\n' "$node" \
"$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")"
done
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
# A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault.
# Prometheus targets
curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets'
# Expected: All targets UP (may show some down if exporters not deployed)
# Grafana health
curl -s http://192.168.68.116:3001/api/health | jq '{status, version}'
# Expected: {"status":"ok","version":"..."}
# LiteLLM metrics (Prometheus endpoint)
curl -s http://192.168.68.116:4001/metrics | head -20
# Expected: Prometheus-formatted metrics output
Report format: Begin every report with the absolute path the probe executed
from (pwd -P, or the script's absolute path) so a stale-consumer report is
distinguishable from a real fault at read time. Summarize actual results from
each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE
API and LiteLLM endpoints above: only connection-refused (000) or timeout is
DOWN; empty output is a warning. For probes whose expected result is a bare 200
(the authenticated Zulip POST, router /health), flag an alert on any unexpected
status (401/403/5xx) — do not summarize it as alive. A bare-200
expectation on the auth-gated PVE API (401) or LiteLLM health (301 redirect)
is a stale expectation, not a fault.
Phase 1: GPU Exporters
NVIDIA (.8 and .110):
- Download
nvidia_gpu_exporterbinary - Create systemd service
nvidia-gpu-exporter.service - Start and enable
AMD (.15):
- Create Python exporter script at
/opt/amdgpu-exporter/exporter.py - Parses
amdgpu_top --json -d 1000output - Exposes key metrics at
:9400/metricsvia Python http.server - Create systemd service
- Start and enable
Phase 2: Prometheus
- Create
/opt/monitoring/directory on CT 116 - Write
prometheus.ymlwith scrape configs for all targets - Add to docker-compose (or separate compose file)
- Start container
Phase 3: Grafana
- Create
/opt/monitoring/grafana/directories - Provision Prometheus datasource
- Provision GPU fleet dashboard JSON
- Provision LiteLLM dashboard JSON
- Add to docker-compose
- Start container
Phase 4: Verification
- Verify all 3 GPU exporters return 200 at :9400/metrics
- Verify Prometheus targets all UP at :9090/targets
- Verify Grafana accessible at :3001 with dashboards
- Verify LiteLLM metrics flowing to Prometheus
Update nginx to proxy(NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard)/monitoring/→ Grafana
Verification Commands
# GPU exporters
curl -s http://192.168.68.8:9400/metrics | grep nvidia
curl -s http://192.168.68.110:9400/metrics | grep nvidia
curl -s http://192.168.68.15:9400/metrics | grep amdgpu
# Prometheus
curl -s http://192.168.68.116:9090/api/v1/targets
# Grafana
curl -s http://192.168.68.116:3001/api/health
# LiteLLM metrics (already live)
curl -s http://192.168.68.116:4001/metrics | head -20