--- kind: function name: infrastructure-monitoring description: > Deploys Prometheus + GPU exporters + Grafana to monitor the entire inference fleet (3 GPU hosts + LiteLLM) from CT 116. GPU metrics from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics via existing /metrics Prometheus endpoint. DEPLOYMENT STATUS (2026-08-09): ✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters (via proxmox-monitor contract). Grafana at :3001, all scrape targets active. ✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on :9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09. ✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09. ⚠️ This contract is target-state aspirational — but GPU export + alerting are now as-built (verified 2026-08-09). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). version: 1.0.1 --- ## Architecture ``` GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo) nvidia-exporter nvidia-exporter amdgpu-exporter :9400 :9400 :9400 │ │ │ └─────────────────────┼─────────────────────┘ ▼ ┌──────────────────────────┐ │ Prometheus │ │ CT 116 :9090 │ │ │ │ Scrape targets: │ │ • 192.168.68.8:9400 │ │ • 192.168.68.110:9400 │ │ • 192.168.68.15:9400 │ │ • litellm:4000/metrics │ └──────────┬───────────────┘ │ ┌──────────▼───────────────┐ │ Grafana │ │ CT 116 :3001 │ │ │ │ Preloaded dashboards: │ │ • GPU Fleet Overview │ │ • LiteLLM Proxy Stats │ └──────────────────────────┘ ``` ## Components ### 1. NVIDIA GPU Exporter (hosts: .8, .110) - Tool: `utkuozdemir/nvidia_gpu_exporter` (Go binary, single static binary) - Listens on `:9400`, exposes `/metrics` in Prometheus format - Metrics: utilization, temp, VRAM, power, clock speeds, fan speed ### 2. AMD GPU Exporter (host: .15) - Custom exporter: Python script wrapping `amdgpu_top --json` - Listens on `:9400`, exposes `/metrics` in Prometheus format - Metrics: power (W), temp (°C), VRAM used/total, GFX clock, utilization - Runs as systemd service for persistence ### 3. Prometheus (CT 116) - Container: `prom/prometheus:latest` - Port: `9090` (internal Docker network) - Scrape interval: 15s - Config: `/opt/monitoring/prometheus.yml` - Storage: Docker volume `prometheus-data` ### 4. Grafana (CT 116) - Container: `grafana/grafana:latest` - Port: `3001` (mapped to host) - Data source: Prometheus at `http://prometheus:9090` - Provisioned dashboards for GPU fleet + LiteLLM - Accessible at `http://192.168.68.116:3001` ## Parameters - gpu_nvidia_hosts: ["192.168.68.8", "192.168.68.110"] - gpu_amd_hosts: ["192.168.68.15"] - monitoring_host: "192.168.68.116" - prometheus_port: 9090 - grafana_port: 3001 - gpu_exporter_port: 9400 ## Requires - SSH access to all GPU hosts for exporter deployment - Docker on CT 116 for Prometheus + Grafana containers - Python 3 on AMD host for custom exporter - nvidia-smi on NVIDIA hosts ## Maintains - All 3 GPU hosts export metrics at :9400/metrics in Prometheus format - Prometheus scrapes all targets every 15s - Grafana dashboards show real-time GPU utilization, temp, VRAM, power - LiteLLM metrics (requests, tokens, latency, errors) visible alongside GPU metrics - Stack persists across reboots (systemd for exporters, Docker restart policy) ## Execution ### Liveness rule (scoped) The any-HTTP-response rule applies ONLY to unauthenticated/auth-gated endpoints, where any HTTP answer proves a listener is up: the PVE API (`https://:8006/api2/json/version`) and LiteLLM health (`/litellm/health`, `301` → `/litellm/health/liveliness`). For those endpoints a probe is **ALIVE** on **ANY** HTTP status — `401`/`403` auth challenges and `3xx` redirects included — and **DOWN = connection refused (`000`) or timeout only**. The PVE API legitimately answers `401` to an unauthenticated probe — that is the healthy signal, not a failure. Same scoped rule as zulip-health (Tanko) and gpu-monitor. Probes whose success condition is specifically a bare `200` are NOT covered by the any-HTTP rule. On those — the authenticated Zulip POST and the router `/health` — an unexpected status (`401`/`403` from a bad or missing credential, `5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive". **STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):** 1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe, and that is a statement about YOUR PROBE, not about the service. 2. **A failed probe is never a service verdict.** Print `probe-failed: ` naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report. Apply the same shape as scripts/disk-gc-scan.py. 3. **Say which probe produced each number.** "Grafana: 000" is unusable; "Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s (retried at 25s: also timeout)" is actionable. ### check-health **RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** **EXECUTABLE OWNER:** The canonical probe set lives in `scripts/infra-monitoring.sh`. A check run is a single command: `bash scripts/infra-monitoring.sh` (from the repository root). Paste its raw output verbatim into the report. The script exits non-zero naming every failed target; there is no "OK" summary when any leg failed. Port drift is caught by `scripts/test_infra_monitoring.sh` which asserts every probed port matches the documented value. **PROBE SHAPE (per standing rules above):** - Every probe prints: `✅ : alive` on success, or `🔴 : probe-failed: : (expected ) ()` on failure - PVE API failures include `(any-HTTP liveness, -k for self-signed)` to distinguish TLS vs connection - Retry once on connection failure at longer timeout (25s connect, 30s max) - Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed - Report the actual probe output, not a summary verdict ```bash # Provenance — run first; paste the absolute path into the report pwd -P # ============================================================ # 1. ZULIP API HEALTH (POST ping) — bare-200 probe # ============================================================ # NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via: zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2") if [ -z "$zulip_key" ]; then echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env" else ZULIP_USER="abiba-bot@chat.sysloggh.net" code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}") echo "Zulip API https://chat.sysloggh.net/api/v1/messages -> $code" # Expected: 200 (bare-200 probe; any other status is an ALERT) fi # ============================================================ # 2. PM2 PROCESS HEALTH # ============================================================ pm2 jlist # Expected: 4/4 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner) # spoton-service removed 2026-09-14 (not in live set) # ============================================================ # 3. GPU EXPORTERS — any-HTTP probe (metrics endpoint) # ============================================================ # Probe /metrics (the Prometheus scrape target), not bare / for host in 192.168.68.8 192.168.68.110 192.168.68.15; do code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 "http://$host:9400/metrics") if [ "$code" == "000" ]; then # Retry with longer timeout code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 "http://$host:9400/metrics") echo "GPU exporter http://$host:9400/metrics -> probe-failed: timeout (retried at 25s: still $code)" else echo "GPU exporter http://$host:9400/metrics -> $code" fi done # Expected: 200 on all 3 hosts (RTX 3090, RTX 5070, Strix Halo) # ============================================================ # 4. ROUTER HEALTH (via nginx on port 80) — bare-200 probe # ============================================================ code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/health) if [ "$code" == "000" ]; then code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/health) echo "Router http://192.168.68.116/health -> probe-failed: timeout (retried at 25s: still $code)" else echo "Router http://192.168.68.116/health -> $code" fi # Expected: 200 (bare-200 probe; any other status is an ALERT) # ============================================================ # 5. LITELLM HEALTH (via nginx on port 80) — any-HTTP probe # ============================================================ code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/litellm/health) if [ "$code" == "000" ]; then code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/litellm/health) echo "LiteLLM http://192.168.68.116/litellm/health -> probe-failed: timeout (retried at 25s: still $code)" else echo "LiteLLM http://192.168.68.116/litellm/health -> $code" fi # Expected: 301 → /litellm/health/liveliness (any HTTP status = ALIVE) # ============================================================ # 6. PVE API LIVENESS — any-HTTP probe (auth-gated) # ============================================================ # Probe the REAL PVE nodes on :8006, never the monitoring host CT 116. for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 10 "https://$node:8006/api2/json/version") if [ "$code" == "000" ]; then code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 25 "https://$node:8006/api2/json/version") echo "PVE API https://$node:8006/api2/json/version -> probe-failed: timeout (retried at 25s: still $code)" else echo "PVE API https://$node:8006/api2/json/version -> $code" fi done # Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12) # 401 is the EXPECTED healthy response (auth-gated); 000/timeout = DOWN # ============================================================ # 7. PROMETHEUS TARGETS — bare-200 probe # ============================================================ code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:9090/api/v1/targets) if [ "$code" == "000" ]; then code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:9090/api/v1/targets) echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> probe-failed: timeout (retried at 25s: still $code)" else echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> $code" fi # Expected: 200 (bare-200 probe; any other status is an ALERT) # ============================================================ # 8. GRAFANA HEALTH — any-HTTP probe # ============================================================ code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:3001/api/health) if [ "$code" == "000" ]; then code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:3001/api/health) echo "Grafana http://192.168.68.116:3001/api/health -> probe-failed: timeout (retried at 25s: still $code)" else echo "Grafana http://192.168.68.116:3001/api/health -> $code" fi # Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN) # ============================================================ # 9. LITELLM METRICS (Prometheus endpoint) — any-HTTP probe # ============================================================ code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:4000/metrics) if [ "$code" == "000" ]; then code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:4000/metrics) echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> probe-failed: timeout (retried at 25s: still $code)" else echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> $code" fi # Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN) ``` **Report format**: Begin every report with the **absolute path the probe executed from** (`pwd -P`) so a stale-consumer report is distinguishable from a real fault at read time. For each probe, print the target name, the full URL, and the HTTP code (or failure kind with retry details). Apply the standing probe rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not a failure. ### Docker Stats and PVE Exporter Ports These two exporters bind to 127.0.0.1 on CT 116 (localhost-only) and must be probed via SSH: | Exporter | Port | Container | Metrics | |----------|------|-----------|---------| | **Docker Stats** | **9324** | harness-docker-stats | `docker_container_*` (per-container CPU/mem/network) | | **PVE Exporter** | **9221** | harness-pve-exporter | `pve_*` (5 cluster-level metrics) | **IMPORTANT**: Do not confuse with port 9323, which is owned by dockerd and serves the Docker Engine's own metrics (`builder_builds_*`, `containerd_build_info_*`). ```bash # Docker Stats (harness-docker-stats) ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics" # Expected: 200 or 404 (any HTTP status = ALIVE) # PVE Exporter (harness-pve-exporter) ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics" # Expected: 200 or 404 (any HTTP status = ALIVE) ``` ### Phase 1: GPU Exporters **NVIDIA (.8 and .110)**: 1. Download `nvidia_gpu_exporter` binary