--- kind: function name: infrastructure-monitoring description: > Deploys Prometheus + GPU exporters + Grafana to monitor the entire inference fleet (3 GPU hosts + LiteLLM) from CT 116. GPU metrics from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics via existing /metrics Prometheus endpoint. DEPLOYMENT STATUS (2026-08-09): ✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters (via proxmox-monitor contract). Grafana at :3001, all scrape targets active. ✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on :9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09. ✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09. ⚠️ This contract is target-state aspirational — but GPU export + alerting are now as-built (verified 2026-08-09). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). version: 1.0.0 --- ## Architecture ``` GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo) nvidia-exporter nvidia-exporter amdgpu-exporter :9400 :9400 :9400 │ │ │ └─────────────────────┼─────────────────────┘ ▼ ┌──────────────────────────┐ │ Prometheus │ │ CT 116 :9090 │ │ │ │ Scrape targets: │ │ • 192.168.68.8:9400 │ │ • 192.168.68.110:9400 │ │ • 192.168.68.15:9400 │ │ • litellm:4000/metrics │ └──────────┬───────────────┘ │ ┌──────────▼───────────────┐ │ Grafana │ │ CT 116 :3001 │ │ │ │ Preloaded dashboards: │ │ • GPU Fleet Overview │ │ • LiteLLM Proxy Stats │ └──────────────────────────┘ ``` ## Components ### 1. NVIDIA GPU Exporter (hosts: .8, .110) - Tool: `utkuozdemir/nvidia_gpu_exporter` (Go binary, single static binary) - Listens on `:9400`, exposes `/metrics` in Prometheus format - Metrics: utilization, temp, VRAM, power, clock speeds, fan speed ### 2. AMD GPU Exporter (host: .15) - Custom exporter: Python script wrapping `amdgpu_top --json` - Listens on `:9400`, exposes `/metrics` in Prometheus format - Metrics: power (W), temp (°C), VRAM used/total, GFX clock, utilization - Runs as systemd service for persistence ### 3. Prometheus (CT 116) - Container: `prom/prometheus:latest` - Port: `9090` (internal Docker network) - Scrape interval: 15s - Config: `/opt/monitoring/prometheus.yml` - Storage: Docker volume `prometheus-data` ### 4. Grafana (CT 116) - Container: `grafana/grafana:latest` - Port: `3001` (mapped to host) - Data source: Prometheus at `http://prometheus:9090` - Provisioned dashboards for GPU fleet + LiteLLM - Accessible at `http://192.168.68.116:3001` ## Parameters - gpu_nvidia_hosts: ["192.168.68.8", "192.168.68.110"] - gpu_amd_hosts: ["192.168.68.15"] - monitoring_host: "192.168.68.116" - prometheus_port: 9090 - grafana_port: 3001 - gpu_exporter_port: 9400 ## Requires - SSH access to all GPU hosts for exporter deployment - Docker on CT 116 for Prometheus + Grafana containers - Python 3 on AMD host for custom exporter - nvidia-smi on NVIDIA hosts ## Maintains - All 3 GPU hosts export metrics at :9400/metrics in Prometheus format - Prometheus scrapes all targets every 15s - Grafana dashboards show real-time GPU utilization, temp, VRAM, power - LiteLLM metrics (requests, tokens, latency, errors) visible alongside GPU metrics - Stack persists across reboots (systemd for exporters, Docker restart policy) ## Execution ### Liveness rule (scoped) The any-HTTP-response rule applies ONLY to unauthenticated/auth-gated endpoints, where any HTTP answer proves a listener is up: the PVE API (`https://:8006/api2/json/version`) and LiteLLM health (`/litellm/health`, `301` → `/litellm/health/liveliness`). For those endpoints a probe is **ALIVE** on **ANY** HTTP status — `401`/`403` auth challenges and `3xx` redirects included — and **DOWN = connection refused (`000`) or timeout only**. The PVE API legitimately answers `401` to an unauthenticated probe — that is the healthy signal, not a failure. Same scoped rule as zulip-health (Tanko) and gpu-monitor. Probes whose success condition is specifically a bare `200` are NOT covered by the any-HTTP rule. On those — the authenticated Zulip POST and the router `/health` — an unexpected status (`401`/`403` from a bad or missing credential, `5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive". ### check-health **RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** ```bash # Provenance — run first; paste the absolute path into the report pwd -P # Zulip API health (POST ping) source /etc/litellm-monitor.env ZULIP_USER="abiba-bot@chat.sysloggh.net" curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${ZULIP_BOT_KEY}" # Expected: 200 (HTTP 000 = unreachable/cache) # PM2 process health pm2 jlist # Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service) # GPU exporters (may be down per DEPLOYMENT STATUS) curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL" curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL" curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL" # Router health (via nginx on port 80) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health # Expected: 200 (Router is up and responding) # LiteLLM health (via nginx on port 80) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health # Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive # PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring # host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 — # that was the stale-vantage bug this replaces (CT 116 is the monitoring host, # not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE # by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response); # DOWN = connection refused (000) or timeout only. for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do printf '%s:8006 -> %s\n' "$node" \ "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")" done # Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12) # A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault. # Prometheus targets curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets' # Expected: All targets UP (may show some down if exporters not deployed) # Grafana health curl -s http://192.168.68.116:3001/api/health | jq '{status, version}' # Expected: {"status":"ok","version":"..."} # LiteLLM metrics (Prometheus endpoint) curl -s http://192.168.68.116:4001/metrics | head -20 # Expected: Prometheus-formatted metrics output ``` **Report format**: Begin every report with the **absolute path the probe executed from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual results from each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE API and LiteLLM endpoints above: only connection-refused (`000`) or timeout is DOWN; empty output is a warning. For probes whose expected result is a bare `200` (the authenticated Zulip POST, router `/health`), flag an alert on any unexpected status (`401`/`403`/`5xx`) — do not summarize it as alive. A bare-`200` expectation on the auth-gated PVE API (`401`) or LiteLLM health (`301` redirect) is a stale expectation, not a fault. ### Phase 1: GPU Exporters **NVIDIA (.8 and .110)**: 1. Download `nvidia_gpu_exporter` binary 2. Create systemd service `nvidia-gpu-exporter.service` 3. Start and enable **AMD (.15)**: 1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py` 2. Parses `amdgpu_top --json -d 1000` output 3. Exposes key metrics at `:9400/metrics` via Python http.server 4. Create systemd service 5. Start and enable ### Phase 2: Prometheus 1. Create `/opt/monitoring/` directory on CT 116 2. Write `prometheus.yml` with scrape configs for all targets 3. Add to docker-compose (or separate compose file) 4. Start container ### Phase 3: Grafana 1. Create `/opt/monitoring/grafana/` directories 2. Provision Prometheus datasource 3. Provision GPU fleet dashboard JSON 4. Provision LiteLLM dashboard JSON 5. Add to docker-compose 6. Start container ### Phase 4: Verification 1. Verify all 3 GPU exporters return 200 at :9400/metrics 2. Verify Prometheus targets all UP at :9090/targets 3. Verify Grafana accessible at :3001 with dashboards 4. Verify LiteLLM metrics flowing to Prometheus 5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard) ## Verification Commands ```bash # GPU exporters curl -s http://192.168.68.8:9400/metrics | grep nvidia curl -s http://192.168.68.110:9400/metrics | grep nvidia curl -s http://192.168.68.15:9400/metrics | grep amdgpu # Prometheus curl -s http://192.168.68.116:9090/api/v1/targets # Grafana curl -s http://192.168.68.116:3001/api/health # LiteLLM metrics (already live) curl -s http://192.168.68.116:4001/metrics | head -20 ```