Files
prose-contracts/infrastructure-monitoring.prose.md
root a13457bcd6
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
fix(infra): Fix Docker Stats (9324) and PVE Exporter (9221) probe ports
Previously probed wrong ports:
- Docker Stats was at 9323 (dockerd metrics) but should be 9324
  (harness-docker-stats, docker_container_* metrics)
- PVE Exporter was at 9324 (harness-docker-stats) but should be 9221
  (harness-pve-exporter, 5 pve_* metrics)

Both exporters bind to 127.0.0.1 on CT 116 and must be probed via SSH.

Updated infrastructure-monitoring.prose.md to document the correct ports.
Added test assertions verifying the exact ports are probed.

Branch: fix/infra-monitoring-probe-ports-20260919
2026-09-19 02:32:00 +00:00

15 KiB

kind, name, description, version
kind name description version
function infrastructure-monitoring Deploys Prometheus + GPU exporters + Grafana to monitor the entire inference fleet (3 GPU hosts + LiteLLM) from CT 116. GPU metrics from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics via existing /metrics Prometheus endpoint. DEPLOYMENT STATUS (2026-08-09): ✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters (via proxmox-monitor contract). Grafana at :3001, all scrape targets active. ✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on :9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09. ✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09. ⚠️ This contract is target-state aspirational — but GPU export + alerting are now as-built (verified 2026-08-09). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). 1.0.1

Architecture

GPU .8 (RTX 3090)    GPU .110 (RTX 5070)   GPU .15 (Strix Halo)
  nvidia-exporter       nvidia-exporter       amdgpu-exporter
  :9400                 :9400                 :9400
       │                     │                     │
       └─────────────────────┼─────────────────────┘
                             ▼
              ┌──────────────────────────┐
              │       Prometheus         │
              │       CT 116 :9090       │
              │                          │
              │  Scrape targets:         │
              │  • 192.168.68.8:9400     │
              │  • 192.168.68.110:9400   │
              │  • 192.168.68.15:9400    │
              │  • litellm:4000/metrics  │
              └──────────┬───────────────┘
                         │
              ┌──────────▼───────────────┐
              │        Grafana           │
              │       CT 116 :3001       │
              │                          │
              │  Preloaded dashboards:   │
              │  • GPU Fleet Overview    │
              │  • LiteLLM Proxy Stats   │
              └──────────────────────────┘

Components

1. NVIDIA GPU Exporter (hosts: .8, .110)

  • Tool: utkuozdemir/nvidia_gpu_exporter (Go binary, single static binary)
  • Listens on :9400, exposes /metrics in Prometheus format
  • Metrics: utilization, temp, VRAM, power, clock speeds, fan speed

2. AMD GPU Exporter (host: .15)

  • Custom exporter: Python script wrapping amdgpu_top --json
  • Listens on :9400, exposes /metrics in Prometheus format
  • Metrics: power (W), temp (°C), VRAM used/total, GFX clock, utilization
  • Runs as systemd service for persistence

3. Prometheus (CT 116)

  • Container: prom/prometheus:latest
  • Port: 9090 (internal Docker network)
  • Scrape interval: 15s
  • Config: /opt/monitoring/prometheus.yml
  • Storage: Docker volume prometheus-data

4. Grafana (CT 116)

  • Container: grafana/grafana:latest
  • Port: 3001 (mapped to host)
  • Data source: Prometheus at http://prometheus:9090
  • Provisioned dashboards for GPU fleet + LiteLLM
  • Accessible at http://192.168.68.116:3001

Parameters

  • gpu_nvidia_hosts: ["192.168.68.8", "192.168.68.110"]
  • gpu_amd_hosts: ["192.168.68.15"]
  • monitoring_host: "192.168.68.116"
  • prometheus_port: 9090
  • grafana_port: 3001
  • gpu_exporter_port: 9400

Requires

  • SSH access to all GPU hosts for exporter deployment
  • Docker on CT 116 for Prometheus + Grafana containers
  • Python 3 on AMD host for custom exporter
  • nvidia-smi on NVIDIA hosts

Maintains

  • All 3 GPU hosts export metrics at :9400/metrics in Prometheus format
  • Prometheus scrapes all targets every 15s
  • Grafana dashboards show real-time GPU utilization, temp, VRAM, power
  • LiteLLM metrics (requests, tokens, latency, errors) visible alongside GPU metrics
  • Stack persists across reboots (systemd for exporters, Docker restart policy)

Execution

Liveness rule (scoped)

The any-HTTP-response rule applies ONLY to unauthenticated/auth-gated endpoints, where any HTTP answer proves a listener is up: the PVE API (https://<node>:8006/api2/json/version) and LiteLLM health (/litellm/health, 301 → /litellm/health/liveliness). For those endpoints a probe is ALIVE on ANY HTTP status — 401/403 auth challenges and 3xx redirects included — and DOWN = connection refused (000) or timeout only. The PVE API legitimately answers 401 to an unauthenticated probe — that is the healthy signal, not a failure. Same scoped rule as zulip-health (Tanko) and gpu-monitor.

Probes whose success condition is specifically a bare 200 are NOT covered by the any-HTTP rule. On those — the authenticated Zulip POST and the router /health — an unexpected status (401/403 from a bad or missing credential, 5xx, or anything other than the expected 200) is an ALERT, not "alive".

STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):

  1. Any HTTP status means ALIVE. 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe, and that is a statement about YOUR PROBE, not about the service.
  2. A failed probe is never a service verdict. Print probe-failed: <target> <kind> naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report. Apply the same shape as scripts/disk-gc-scan.py.
  3. Say which probe produced each number. "Grafana: 000" is unusable; "Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.

check-health

RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.

EXECUTABLE OWNER: The canonical probe set lives in scripts/infra-monitoring.sh. A check run is a single command: bash scripts/infra-monitoring.sh (from the repository root). Paste its raw output verbatim into the report. The script exits non-zero naming every failed target; there is no "OK" summary when any leg failed. Port drift is caught by scripts/test_infra_monitoring.sh which asserts every probed port matches the documented value.

PROBE SHAPE (per standing rules above):

  • Every probe prints: ✅ <name>: alive on success, or 🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>) on failure
  • PVE API failures include (any-HTTP liveness, -k for self-signed) to distinguish TLS vs connection
  • Retry once on connection failure at longer timeout (25s connect, 30s max)
  • Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
  • Report the actual probe output, not a summary verdict
# Provenance — run first; paste the absolute path into the report
pwd -P

# ============================================================
# 1. ZULIP API HEALTH (POST ping) — bare-200 probe
# ============================================================
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
if [ -z "$zulip_key" ]; then
  echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
else
  ZULIP_USER="abiba-bot@chat.sysloggh.net"
  code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}")
  echo "Zulip API https://chat.sysloggh.net/api/v1/messages -> $code"
  # Expected: 200 (bare-200 probe; any other status is an ALERT)
fi

# ============================================================
# 2. PM2 PROCESS HEALTH
# ============================================================
pm2 jlist
# Expected: 4/4 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner)
# spoton-service removed 2026-09-14 (not in live set)

# ============================================================
# 3. GPU EXPORTERS — any-HTTP probe (metrics endpoint)
# ============================================================
# Probe /metrics (the Prometheus scrape target), not bare /
for host in 192.168.68.8 192.168.68.110 192.168.68.15; do
  code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 "http://$host:9400/metrics")
  if [ "$code" == "000" ]; then
    # Retry with longer timeout
    code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 "http://$host:9400/metrics")
    echo "GPU exporter http://$host:9400/metrics -> probe-failed: timeout (retried at 25s: still $code)"
  else
    echo "GPU exporter http://$host:9400/metrics -> $code"
  fi
done
# Expected: 200 on all 3 hosts (RTX 3090, RTX 5070, Strix Halo)

# ============================================================
# 4. ROUTER HEALTH (via nginx on port 80) — bare-200 probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/health)
if [ "$code" == "000" ]; then
  code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/health)
  echo "Router http://192.168.68.116/health -> probe-failed: timeout (retried at 25s: still $code)"
else
  echo "Router http://192.168.68.116/health -> $code"
fi
# Expected: 200 (bare-200 probe; any other status is an ALERT)

# ============================================================
# 5. LITELLM HEALTH (via nginx on port 80) — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/litellm/health)
if [ "$code" == "000" ]; then
  code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/litellm/health)
  echo "LiteLLM http://192.168.68.116/litellm/health -> probe-failed: timeout (retried at 25s: still $code)"
else
  echo "LiteLLM http://192.168.68.116/litellm/health -> $code"
fi
# Expected: 301 → /litellm/health/liveliness (any HTTP status = ALIVE)

# ============================================================
# 6. PVE API LIVENESS — any-HTTP probe (auth-gated)
# ============================================================
# Probe the REAL PVE nodes on :8006, never the monitoring host CT 116.
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
  code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 10 "https://$node:8006/api2/json/version")
  if [ "$code" == "000" ]; then
    code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 25 "https://$node:8006/api2/json/version")
    echo "PVE API https://$node:8006/api2/json/version -> probe-failed: timeout (retried at 25s: still $code)"
  else
    echo "PVE API https://$node:8006/api2/json/version -> $code"
  fi
done
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
# 401 is the EXPECTED healthy response (auth-gated); 000/timeout = DOWN

# ============================================================
# 7. PROMETHEUS TARGETS — bare-200 probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:9090/api/v1/targets)
if [ "$code" == "000" ]; then
  code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:9090/api/v1/targets)
  echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> probe-failed: timeout (retried at 25s: still $code)"
else
  echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> $code"
fi
# Expected: 200 (bare-200 probe; any other status is an ALERT)

# ============================================================
# 8. GRAFANA HEALTH — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:3001/api/health)
if [ "$code" == "000" ]; then
  code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:3001/api/health)
  echo "Grafana http://192.168.68.116:3001/api/health -> probe-failed: timeout (retried at 25s: still $code)"
else
  echo "Grafana http://192.168.68.116:3001/api/health -> $code"
fi
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)

# ============================================================
# 9. LITELLM METRICS (Prometheus endpoint) — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:4000/metrics)
if [ "$code" == "000" ]; then
  code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:4000/metrics)
  echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> probe-failed: timeout (retried at 25s: still $code)"
else
  echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> $code"
fi
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)

Report format: Begin every report with the absolute path the probe executed from (pwd -P) so a stale-consumer report is distinguishable from a real fault at read time. For each probe, print the target name, the full URL, and the HTTP code (or failure kind with retry details). Apply the standing probe rules: any HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not a failure.

Docker Stats and PVE Exporter Ports

These two exporters bind to 127.0.0.1 on CT 116 (localhost-only) and must be probed via SSH:

Exporter Port Container Metrics
Docker Stats 9324 harness-docker-stats docker_container_* (per-container CPU/mem/network)
PVE Exporter 9221 harness-pve-exporter pve_* (5 cluster-level metrics)

IMPORTANT: Do not confuse with port 9323, which is owned by dockerd and serves the Docker Engine's own metrics (builder_builds_*, containerd_build_info_*).

# Docker Stats (harness-docker-stats)
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics"
# Expected: 200 or 404 (any HTTP status = ALIVE)

# PVE Exporter (harness-pve-exporter)
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics"
# Expected: 200 or 404 (any HTTP status = ALIVE)

Phase 1: GPU Exporters

NVIDIA (.8 and .110):

  1. Download nvidia_gpu_exporter binary