1. Replace '(## Execution' + ')' with '## Execution' in 6 contract files 2. Make LOG_DIR honor CONTRACT_RUN_LOG_DIR env var (default: /var/log/contract-runs) 3. Fix Zulip alert: correct URL (https://chat.sysloggh.net/api/v1), user (abiba-bot@chat.sysloggh.net), take ZULIP_API_KEY from environment, and write alert failures to run log
15 KiB
kind, name, description, version
| kind | name | description | version |
|---|---|---|---|
| function | infrastructure-monitoring | Deploys Prometheus + GPU exporters + Grafana to monitor the entire inference fleet (3 GPU hosts + LiteLLM) from CT 116. GPU metrics from nvidia-smi (.8, .110) and amdgpu_top (.15). LiteLLM metrics via existing /metrics Prometheus endpoint. DEPLOYMENT STATUS (2026-08-09): ✅ Core stack deployed: Prometheus + Grafana + pve/node/docker exporters (via proxmox-monitor contract). Grafana at :3001, all scrape targets active. ✅ GPU exporters DEPLOYED: all 3 GPU hosts (.8/.110/.15) run exporters on :9400 (nvidia_gpu_exporter / amdgpu exporter) — verified 200 on 2026-08-09. ✅ LiteLLM /metrics scraping live (success_callback: prometheus; auth via master key) + Alertmanager + Zulip bridge (alerts-infra) added 2026-08-09. ⚠️ This contract is target-state aspirational — but GPU export + alerting are now as-built (verified 2026-08-09). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). | 1.0.1 |
Architecture
GPU .8 (RTX 3090) GPU .110 (RTX 5070) GPU .15 (Strix Halo)
nvidia-exporter nvidia-exporter amdgpu-exporter
:9400 :9400 :9400
│ │ │
└─────────────────────┼─────────────────────┘
▼
┌──────────────────────────┐
│ Prometheus │
│ CT 116 :9090 │
│ │
│ Scrape targets: │
│ • 192.168.68.8:9400 │
│ • 192.168.68.110:9400 │
│ • 192.168.68.15:9400 │
│ • litellm:4000/metrics │
└──────────┬───────────────┘
│
┌──────────▼───────────────┐
│ Grafana │
│ CT 116 :3001 │
│ │
│ Preloaded dashboards: │
│ • GPU Fleet Overview │
│ • LiteLLM Proxy Stats │
└──────────────────────────┘
Components
1. NVIDIA GPU Exporter (hosts: .8, .110)
- Tool:
utkuozdemir/nvidia_gpu_exporter(Go binary, single static binary) - Listens on
:9400, exposes/metricsin Prometheus format - Metrics: utilization, temp, VRAM, power, clock speeds, fan speed
2. AMD GPU Exporter (host: .15)
- Custom exporter: Python script wrapping
amdgpu_top --json - Listens on
:9400, exposes/metricsin Prometheus format - Metrics: power (W), temp (°C), VRAM used/total, GFX clock, utilization
- Runs as systemd service for persistence
3. Prometheus (CT 116)
- Container:
prom/prometheus:latest - Port:
9090(internal Docker network) - Scrape interval: 15s
- Config:
/opt/monitoring/prometheus.yml - Storage: Docker volume
prometheus-data
4. Grafana (CT 116)
- Container:
grafana/grafana:latest - Port:
3001(mapped to host) - Data source: Prometheus at
http://prometheus:9090 - Provisioned dashboards for GPU fleet + LiteLLM
- Accessible at
http://192.168.68.116:3001
Parameters
- gpu_nvidia_hosts: ["192.168.68.8", "192.168.68.110"]
- gpu_amd_hosts: ["192.168.68.15"]
- monitoring_host: "192.168.68.116"
- prometheus_port: 9090
- grafana_port: 3001
- gpu_exporter_port: 9400
Requires
- SSH access to all GPU hosts for exporter deployment
- Docker on CT 116 for Prometheus + Grafana containers
- Python 3 on AMD host for custom exporter
- nvidia-smi on NVIDIA hosts
Maintains
- All 3 GPU hosts export metrics at :9400/metrics in Prometheus format
- Prometheus scrapes all targets every 15s
- Grafana dashboards show real-time GPU utilization, temp, VRAM, power
- LiteLLM metrics (requests, tokens, latency, errors) visible alongside GPU metrics
- Stack persists across reboots (systemd for exporters, Docker restart policy)
Execution
Execution model: This contract is executed by a host-scheduled cron job (see
scripts/contract-run.sh). The cron job runs the monitoring script directly on the
target host and appends the result to /var/log/contract-runs/<contract>.log. An
agent-session acknowledgement (a done: line in the ops status log) is NOT
execution — it only proves the agent read the result and reported it. The actual
monitoring work happens in the host cron job.
Liveness rule (scoped)
The any-HTTP-response rule applies ONLY to unauthenticated/auth-gated endpoints,
where any HTTP answer proves a listener is up: the PVE API
(https://<node>:8006/api2/json/version) and LiteLLM health
(/litellm/health, 301 → /litellm/health/liveliness). For those endpoints a
probe is ALIVE on ANY HTTP status — 401/403 auth challenges and 3xx
redirects included — and DOWN = connection refused (000) or timeout only.
The PVE API legitimately answers 401 to an unauthenticated probe — that is the
healthy signal, not a failure. Same scoped rule as zulip-health (Tanko) and
gpu-monitor.
Probes whose success condition is specifically a bare 200 are NOT covered by
the any-HTTP rule. On those — the authenticated Zulip POST and the router
/health — an unexpected status (401/403 from a bad or missing credential,
5xx, or anything other than the expected 200) is an ALERT, not "alive".
STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):
- Any HTTP status means ALIVE. 200, 301, 302, 401, 403, 404 all prove the service answered — report the code, never "down". A redirect is not a failure. Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe, and that is a statement about YOUR PROBE, not about the service.
- A failed probe is never a service verdict. Print
probe-failed: <target> <kind>naming the exact URL/host/port and the failure kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only then report. Apply the same shape as scripts/disk-gc-scan.py. - Say which probe produced each number. "Grafana: 000" is unusable; "Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s (retried at 25s: also timeout)" is actionable.
check-health
RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.
EXECUTABLE OWNER: The canonical probe set lives in scripts/infra-monitoring.sh.
A check run is a single command: bash scripts/infra-monitoring.sh (from the
repository root). Paste its raw output verbatim into the report. The script
exits non-zero naming every failed target; there is no "OK" summary when any
leg failed. Port drift is caught by scripts/test_infra_monitoring.sh which
asserts every probed port matches the documented value.
PROBE SHAPE (per standing rules above):
- Every probe prints:
✅ <name>: aliveon success, or🔴 <name>: probe-failed: <host>:<port> (expected <pattern>) (<kind>)on failure - PVE API failures include
(any-HTTP liveness, -k for self-signed)to distinguish TLS vs connection - Retry once on connection failure at longer timeout (25s connect, 30s max)
- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed
- Report the actual probe output, not a summary verdict
# Provenance — run first; paste the absolute path into the report
pwd -P
# ============================================================
# 1. ZULIP API HEALTH (POST ping) — bare-200 probe
# ============================================================
# NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via:
zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2")
if [ -z "$zulip_key" ]; then
echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env"
else
ZULIP_USER="abiba-bot@chat.sysloggh.net"
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}")
echo "Zulip API https://chat.sysloggh.net/api/v1/messages -> $code"
# Expected: 200 (bare-200 probe; any other status is an ALERT)
fi
# ============================================================
# 2. PM2 PROCESS HEALTH
# ============================================================
pm2 jlist
# Expected: 4/4 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner)
# spoton-service removed 2026-09-14 (not in live set)
# ============================================================
# 3. GPU EXPORTERS — any-HTTP probe (metrics endpoint)
# ============================================================
# Probe /metrics (the Prometheus scrape target), not bare /
for host in 192.168.68.8 192.168.68.110 192.168.68.15; do
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 "http://$host:9400/metrics")
if [ "$code" == "000" ]; then
# Retry with longer timeout
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 "http://$host:9400/metrics")
echo "GPU exporter http://$host:9400/metrics -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "GPU exporter http://$host:9400/metrics -> $code"
fi
done
# Expected: 200 on all 3 hosts (RTX 3090, RTX 5070, Strix Halo)
# ============================================================
# 4. ROUTER HEALTH (via nginx on port 80) — bare-200 probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/health)
echo "Router http://192.168.68.116/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Router http://192.168.68.116/health -> $code"
fi
# Expected: 200 (bare-200 probe; any other status is an ALERT)
# ============================================================
# 5. LITELLM HEALTH (via nginx on port 80) — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/litellm/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/litellm/health)
echo "LiteLLM http://192.168.68.116/litellm/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "LiteLLM http://192.168.68.116/litellm/health -> $code"
fi
# Expected: 301 → /litellm/health/liveliness (any HTTP status = ALIVE)
# ============================================================
# 6. PVE API LIVENESS — any-HTTP probe (auth-gated)
# ============================================================
# Probe the REAL PVE nodes on :8006, never the monitoring host CT 116.
for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 10 "https://$node:8006/api2/json/version")
if [ "$code" == "000" ]; then
code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 25 "https://$node:8006/api2/json/version")
echo "PVE API https://$node:8006/api2/json/version -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "PVE API https://$node:8006/api2/json/version -> $code"
fi
done
# Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12)
# 401 is the EXPECTED healthy response (auth-gated); 000/timeout = DOWN
# ============================================================
# 7. PROMETHEUS TARGETS — bare-200 probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:9090/api/v1/targets)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:9090/api/v1/targets)
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> $code"
fi
# Expected: 200 (bare-200 probe; any other status is an ALERT)
# ============================================================
# 8. GRAFANA HEALTH — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:3001/api/health)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:3001/api/health)
echo "Grafana http://192.168.68.116:3001/api/health -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "Grafana http://192.168.68.116:3001/api/health -> $code"
fi
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
# ============================================================
# 9. LITELLM METRICS (Prometheus endpoint) — any-HTTP probe
# ============================================================
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:4000/metrics)
if [ "$code" == "000" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:4000/metrics)
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> probe-failed: timeout (retried at 25s: still $code)"
else
echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> $code"
fi
# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN)
Report format: Begin every report with the absolute path the probe executed
from (pwd -P) so a stale-consumer report is distinguishable from a real fault
at read time. For each probe, print the target name, the full URL, and the HTTP
code (or failure kind with retry details). Apply the standing probe rules: any
HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not
a failure.
Docker Stats and PVE Exporter Ports
These two exporters bind to 127.0.0.1 on CT 116 (localhost-only) and must be probed via SSH:
| Exporter | Port | Container | Metrics |
|---|---|---|---|
| Docker Stats | 9324 | harness-docker-stats | docker_container_* (per-container CPU/mem/network) |
| PVE Exporter | 9221 | harness-pve-exporter | pve_* (5 cluster-level metrics) |
IMPORTANT: Do not confuse with port 9323, which is owned by dockerd and serves the Docker Engine's own metrics (builder_builds_*, containerd_build_info_*).
# Docker Stats (harness-docker-stats)
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9324/metrics"
# Expected: 200 or 404 (any HTTP status = ALIVE)
# PVE Exporter (harness-pve-exporter)
ssh root@192.168.68.116 "curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://127.0.0.1:9221/metrics"
# Expected: 200 or 404 (any HTTP status = ALIVE)
Phase 1: GPU Exporters
NVIDIA (.8 and .110):
- Download
nvidia_gpu_exporterbinary