PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Previously probed wrong ports: - Docker Stats was at 9323 (dockerd metrics) but should be 9324 (harness-docker-stats, docker_container_* metrics) - PVE Exporter was at 9324 (harness-docker-stats) but should be 9221 (harness-pve-exporter, 5 pve_* metrics) Both exporters bind to 127.0.0.1 on CT 116 and must be probed via SSH. Updated infrastructure-monitoring.prose.md to document the correct ports. Added test assertions verifying the exact ports are probed. Branch: fix/infra-monitoring-probe-ports-20260919
305 lines
15 KiB
Markdown
305 lines
15 KiB
Markdown
---
|
|
kind: function
|
|
name: infrastructure-monitoring
|
|
description: >
|
|
Deploys Prometheus + GPU exporters + Grafana to monitor the entire
|
|
inference fleet (3 GPU hosts + LiteLLM) from CT 116. GPU metrics
|
|
from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics
|
|
via existing /metrics Prometheus endpoint.
|
|
|
|
DEPLOYMENT STATUS (2026-08-09):
|
|
✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters
|
|
(via proxmox-monitor contract). Grafana at :3001, all scrape targets active.
|
|
✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on
|
|
:9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09.
|
|
✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via
|
|
master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09.
|
|
⚠️ This contract is target-state aspirational — but GPU export + alerting
|
|
are now as-built (verified 2026-08-09).
|
|
As-built GPU monitoring is via gpu-monitor contract (port 9100 poll).
|
|
version: 1.0.1
|
|
---
|
|
|
|
## Architecture
|
|
|
|
```
|
|
GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
|
|
nvidia-exporter nvidia-exporter amdgpu-exporter
|
|
:9400 :9400 :9400
|
|
│ │ │
|
|
└─────────────────────┼─────────────────────┘
|
|
▼
|
|
┌──────────────────────────┐
|
|
│ Prometheus │
|
|
│ CT 116 :9090 │
|
|
│ │
|
|
│ Scrape targets: │
|
|
│ • 192.168.68.8:9400 │
|
|
│ • 192.168.68.110:9400 │
|
|
│ • 192.168.68.15:9400 │
|
|
│ • litellm:4000/metrics │
|
|
└──────────┬───────────────┘
|
|
│
|
|
┌──────────▼───────────────┐
|
|
│ Grafana │
|
|
│ CT 116 :3001 │
|
|
│ │
|
|
│ Preloaded dashboards: │
|
|
│ • GPU Fleet Overview │
|
|
│ • LiteLLM Proxy Stats │
|
|
└──────────────────────────┘
|
|
```
|
|
|
|
## Components
|
|
|
|
### 1. NVIDIA GPU Exporter (hosts: .8, .110)
|
|
- Tool: `utkuozdemir/nvidia_gpu_exporter` (Go binary, single static binary)
|
|
- Listens on `:9400`, exposes `/metrics` in Prometheus format
|
|
- Metrics: utilization, temp, VRAM, power, clock speeds, fan speed
|
|
|
|
### 2. AMD GPU Exporter (host: .15)
|
|
- Custom exporter: Python script wrapping `amdgpu_top --json`
|
|
- Listens on `:9400`, exposes `/metrics` in Prometheus format
|
|
- Metrics: power (W), temp (°C), VRAM used/total, GFX clock, utilization
|
|
- Runs as systemd service for persistence
|
|
|
|
### 3. Prometheus (CT 116)
|
|
- Container: `prom/prometheus:latest`
|
|
- Port: `9090` (internal Docker network)
|
|
- Scrape interval: 15s
|
|
- Config: `/opt/monitoring/prometheus.yml`
|
|
- Storage: Docker volume `prometheus-data`
|
|
|
|
### 4. Grafana (CT 116)
|
|
- Container: `grafana/grafana:latest`
|
|
- Port: `3001` (mapped to host)
|
|
- Data source: Prometheus at `http://prometheus:9090`
|
|
- Provisioned dashboards for GPU fleet + LiteLLM
|
|
- Accessible at `http://192.168.68.116:3001`
|
|
|
|
## Parameters
|
|
|
|
- gpu_nvidia_hosts: ["192.168.68.8", "192.168.68.110"]
|
|
- gpu_amd_hosts: ["192.168.68.15"]
|
|
- monitoring_host: "192.168.68.116"
|
|
- prometheus_port: 9090
|
|
- grafana_port: 3001
|
|
- gpu_exporter_port: 9400
|
|
|
|
## Requires
|
|
|
|
- SSH access to all GPU hosts for exporter deployment
|
|
- Docker on CT 116 for Prometheus + Grafana containers
|
|
- Python 3 on AMD host for custom exporter
|
|
- nvidia-smi on NVIDIA hosts
|
|
|
|
## Maintains
|
|
|
|
- All 3 GPU hosts export metrics at :9400/metrics in Prometheus format
|
|
- Prometheus scrapes all targets every 15s
|
|
- Grafana dashboards show real-time GPU utilization, temp, VRAM, power
|
|
- LiteLLM metrics (requests, tokens, latency, errors) visible alongside GPU metrics
|
|
- Stack persists across reboots (systemd for exporters, Docker restart policy)
|
|
|
|
## Execution
|
|
|
|
### Liveness rule (scoped)
|
|
|
|
The any-HTTP-response rule applies ONLY to unauthenticated/auth-gated endpoints,
|
|
where any HTTP answer proves a listener is up: the PVE API
|
|
(`https://<node>:8006/api2/json/version`) and LiteLLM health
|
|
(`/litellm/health`, `301` → `/litellm/health/liveliness`). For those endpoints a
|
|
probe is **ALIVE** on **ANY** HTTP status — `401`/`403` auth challenges and `3xx`
|
|
redirects included — and **DOWN = connection refused (`000`) or timeout only**.
|
|
The PVE API legitimately answers `401` to an unauthenticated probe — that is the
|
|
healthy signal, not a failure. Same scoped rule as zulip-health (Tanko) and
|
|
gpu-monitor.
|
|
|
|
Probes whose success condition is specifically a bare `200` are NOT covered by
|
|
the any-HTTP rule. On those — the authenticated Zulip POST and the router
|
|
`/health` — an unexpected status (`401`/`403` from a bad or missing credential,
|
|
`5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive".
|
|
|
|
**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):**
|
|
1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the
|
|
service answered — report the code, never "down". A redirect is not a failure.
|
|
Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe,
|
|
and that is a statement about YOUR PROBE, not about the service.
|
|
2. **A failed probe is never a service verdict.** Print
|
|
`probe-failed: <target> <kind>` naming the exact URL/host/port and the failure
|
|
kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only
|
|
then report. Apply the same shape as scripts/disk-gc-scan.py.
|
|
3. **Say which probe produced each number.** "Grafana: 000" is unusable;
|
|
"Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s
|
|
(retried at 25s: also timeout)" is actionable.
|
|
|
|
### check-health
|
|
|
|
**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real
|
|
tool calls; never repeat a prior report unless a live probe fails.**
|
|
|
|
**EXECUTABLE OWNER:** The canonical probe set lives in `scripts/infra-monitoring.sh`.
|
|
A check run is a single command: `bash scripts/infra-monitoring.sh` (from the
|
|
repository root). Paste its raw output verbatim into the report. The script
|
|
exits non-zero naming every failed target; there is no "OK" summary when any
|
|
leg failed. Port drift is caught by `scripts/test_infra_monitoring.sh` which
|
|
asserts every probed port matches the documented value.
|
|
|
|
**PROBE SHAPE (per standing rules above):**
|
|
- Every probe prints: `✅ <name>: alive` on success, or `🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)` on failure
|
|
- PVE API failures include `(any-HTTP liveness, -k for self-signed)` to distinguish TLS vs connection
|
|
- Retry once on connection failure at longer timeout (25s connect, 30s max)
|
|
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
|
|
- Report the actual probe output, not a summary verdict
|
|
|
|
```bash
|
|
# Provenance — run first; paste the absolute path into the report
|
|
pwd -P
|
|
|
|
# ============================================================
|
|
# 1. ZULIP API HEALTH (POST ping) — bare-200 probe
|
|
# ============================================================
|
|
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
|
|
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
|
|
if [ -z "$zulip_key" ]; then
|
|
echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
|
|
else
|
|
ZULIP_USER="abiba-bot@chat.sysloggh.net"
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}")
|
|
echo "Zulip API https://chat.sysloggh.net/api/v1/messages -> $code"
|
|
# Expected: 200 (bare-200 probe; any other status is an ALERT)
|
|
fi
|
|
|
|
# ============================================================
|
|
# 2. PM2 PROCESS HEALTH
|
|
# ============================================================
|
|
pm2 jlist
|
|
# Expected: 4/4 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner)
|
|
# spoton-service removed 2026-09-14 (not in live set)
|
|
|
|
# ============================================================
|
|
# 3. GPU EXPORTERS — any-HTTP probe (metrics endpoint)
|
|
# ============================================================
|
|
# Probe /metrics (the Prometheus scrape target), not bare /
|
|
for host in 192.168.68.8 192.168.68.110 192.168.68.15; do
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 "http://$host:9400/metrics")
|
|
if [ "$code" == "000" ]; then
|
|
# Retry with longer timeout
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 "http://$host:9400/metrics")
|
|
echo "GPU exporter http://$host:9400/metrics -> probe-failed: timeout (retried at 25s: still $code)"
|
|
else
|
|
echo "GPU exporter http://$host:9400/metrics -> $code"
|
|
fi
|
|
done
|
|
# Expected: 200 on all 3 hosts (RTX 3090, RTX 5070, Strix Halo)
|
|
|
|
# ============================================================
|
|
# 4. ROUTER HEALTH (via nginx on port 80) — bare-200 probe
|
|
# ============================================================
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/health)
|
|
if [ "$code" == "000" ]; then
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/health)
|
|
echo "Router http://192.168.68.116/health -> probe-failed: timeout (retried at 25s: still $code)"
|
|
else
|
|
echo "Router http://192.168.68.116/health -> $code"
|
|
fi
|
|
# Expected: 200 (bare-200 probe; any other status is an ALERT)
|
|
|
|
# ============================================================
|
|
# 5. LITELLM HEALTH (via nginx on port 80) — any-HTTP probe
|
|
# ============================================================
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/litellm/health)
|
|
if [ "$code" == "000" ]; then
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/litellm/health)
|
|
echo "LiteLLM http://192.168.68.116/litellm/health -> probe-failed: timeout (retried at 25s: still $code)"
|
|
else
|
|
echo "LiteLLM http://192.168.68.116/litellm/health -> $code"
|
|
fi
|
|
# Expected: 301 → /litellm/health/liveliness (any HTTP status = ALIVE)
|
|
|
|
# ============================================================
|
|
# 6. PVE API LIVENESS — any-HTTP probe (auth-gated)
|
|
# ============================================================
|
|
# Probe the REAL PVE nodes on :8006, never the monitoring host CT 116.
|
|
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
|
|
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 10 "https://$node:8006/api2/json/version")
|
|
if [ "$code" == "000" ]; then
|
|
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 25 "https://$node:8006/api2/json/version")
|
|
echo "PVE API https://$node:8006/api2/json/version -> probe-failed: timeout (retried at 25s: still $code)"
|
|
else
|
|
echo "PVE API https://$node:8006/api2/json/version -> $code"
|
|
fi
|
|
done
|
|
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
|
|
# 401 is the EXPECTED healthy response (auth-gated); 000/timeout = DOWN
|
|
|
|
# ============================================================
|
|
# 7. PROMETHEUS TARGETS — bare-200 probe
|
|
# ============================================================
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:9090/api/v1/targets)
|
|
if [ "$code" == "000" ]; then
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:9090/api/v1/targets)
|
|
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> probe-failed: timeout (retried at 25s: still $code)"
|
|
else
|
|
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> $code"
|
|
fi
|
|
# Expected: 200 (bare-200 probe; any other status is an ALERT)
|
|
|
|
# ============================================================
|
|
# 8. GRAFANA HEALTH — any-HTTP probe
|
|
# ============================================================
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:3001/api/health)
|
|
if [ "$code" == "000" ]; then
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:3001/api/health)
|
|
echo "Grafana http://192.168.68.116:3001/api/health -> probe-failed: timeout (retried at 25s: still $code)"
|
|
else
|
|
echo "Grafana http://192.168.68.116:3001/api/health -> $code"
|
|
fi
|
|
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
|
|
|
|
# ============================================================
|
|
# 9. LITELLM METRICS (Prometheus endpoint) — any-HTTP probe
|
|
# ============================================================
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:4000/metrics)
|
|
if [ "$code" == "000" ]; then
|
|
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:4000/metrics)
|
|
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> probe-failed: timeout (retried at 25s: still $code)"
|
|
else
|
|
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> $code"
|
|
fi
|
|
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
|
|
```
|
|
|
|
**Report format**: Begin every report with the **absolute path the probe executed
|
|
from** (`pwd -P`) so a stale-consumer report is distinguishable from a real fault
|
|
at read time. For each probe, print the target name, the full URL, and the HTTP
|
|
code (or failure kind with retry details). Apply the standing probe rules: any
|
|
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
|
|
a failure.
|
|
|
|
### Docker Stats and PVE Exporter Ports
|
|
|
|
These two exporters bind to 127.0.0.1 on CT 116 (localhost-only) and must be probed via SSH:
|
|
|
|
| Exporter | Port | Container | Metrics |
|
|
|----------|------|-----------|---------|
|
|
| **Docker Stats** | **9324** | harness-docker-stats | `docker_container_*` (per-container CPU/mem/network) |
|
|
| **PVE Exporter** | **9221** | harness-pve-exporter | `pve_*` (5 cluster-level metrics) |
|
|
|
|
**IMPORTANT**: Do not confuse with port 9323, which is owned by dockerd and serves the Docker Engine's own metrics (`builder_builds_*`, `containerd_build_info_*`).
|
|
|
|
```bash
|
|
# Docker Stats (harness-docker-stats)
|
|
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics"
|
|
# Expected: 200 or 404 (any HTTP status = ALIVE)
|
|
|
|
# PVE Exporter (harness-pve-exporter)
|
|
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics"
|
|
# Expected: 200 or 404 (any HTTP status = ALIVE)
|
|
```
|
|
|
|
### Phase 1: GPU Exporters
|
|
|
|
**NVIDIA (.8 and .110)**:
|
|
1. Download `nvidia_gpu_exporter` binary
|