diff --git a/infrastructure-monitoring.prose.md b/infrastructure-monitoring.prose.md index 68544a7..e3e5ea9 100644 --- a/infrastructure-monitoring.prose.md +++ b/infrastructure-monitoring.prose.md @@ -17,7 +17,7 @@ description: > ⚠️ This contract is target-state aspirational — but GPU export + alerting are now as-built (verified 2026-08-09). As-built GPU monitoring is via gpu-monitor contract (port 9100 poll). -version: 1.0.0 +version: 1.0.1 --- ## Architecture @@ -120,132 +120,156 @@ the any-HTTP rule. On those — the authenticated Zulip POST and the router `/health` — an unexpected status (`401`/`403` from a bad or missing credential, `5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive". +**STANDING PROBE RULES (2026-09-14, from defect report 1150.msg):** +1. **Any HTTP status means ALIVE.** 200, 301, 302, 401, 403, 404 all prove the + service answered — report the code, never "down". A redirect is not a failure. + Only a failed CONNECTION (curl status 000, timeout, refused) is a failed probe, + and that is a statement about YOUR PROBE, not about the service. +2. **A failed probe is never a service verdict.** Print + `probe-failed: ` naming the exact URL/host/port and the failure + kind (timeout, refused, no-route, dns), retry once at a longer timeout, and only + then report. Apply the same shape as scripts/disk-gc-scan.py. +3. **Say which probe produced each number.** "Grafana: 000" is unusable; + "Grafana http://192.168.68.116:3001/api/health -> connection timeout after 10s + (retried at 25s: also timeout)" is actionable. + ### check-health -**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** +**RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real +tool calls; never repeat a prior report unless a live probe fails.** + +**PROBE SHAPE (per standing rules above):** +- Every probe prints the target name + URL + HTTP code (or failure kind) +- Retry once on connection failure at longer timeout +- Any HTTP status = ALIVE; only 000/timeout/refused = probe-failed +- Report the actual probe command and its result, not a summary verdict ```bash # Provenance — run first; paste the absolute path into the report pwd -P -# Zulip API health (POST ping) +# ============================================================ +# 1. ZULIP API HEALTH (POST ping) — bare-200 probe +# ============================================================ # NOTE: /etc/litellm-monitor.env exists only on CT 116, retrieve keys from CT 116 via: zulip_key=$(ssh root@192.168.68.116 "grep ZULIP_BOT_KEY /etc/litellm-monitor.env | cut -d= -f2") if [ -z "$zulip_key" ]; then echo "credential-missing: ZULIP_BOT_KEY not found in /etc/litellm-monitor.env" else ZULIP_USER="abiba-bot@chat.sysloggh.net" - curl -s -o /dev/null -w '%{http_code}' -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}" - # Expected: 200 (HTTP 000 = unreachable/cache) + code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 -X POST https://chat.sysloggh.net/api/v1/messages -u "${ZULIP_USER}:${zulip_key}") + echo "Zulip API https://chat.sysloggh.net/api/v1/messages -> $code" + # Expected: 200 (bare-200 probe; any other status is an ALERT) fi -# PM2 process health +# ============================================================ +# 2. PM2 PROCESS HEALTH +# ============================================================ pm2 jlist -# Expected: 5/5 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner, spoton-service) +# Expected: 4/4 online (abiba-telegram, abiba-zulip, zulip-watchdog, gitea-runner) +# spoton-service removed 2026-09-14 (not in live set) -# GPU exporters (may be down per DEPLOYMENT STATUS) -curl -s http://192.168.68.8:9400/metrics && echo " - OK" || echo " - FAIL" -curl -s http://192.168.68.110:9400/metrics && echo " - OK" || echo " - FAIL" -curl -s http://192.168.68.15:9400/metrics && echo " - OK" || echo " - FAIL" +# ============================================================ +# 3. GPU EXPORTERS — any-HTTP probe (metrics endpoint) +# ============================================================ +# Probe /metrics (the Prometheus scrape target), not bare / +for host in 192.168.68.8 192.168.68.110 192.168.68.15; do + code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 "http://$host:9400/metrics") + if [ "$code" == "000" ]; then + # Retry with longer timeout + code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 "http://$host:9400/metrics") + echo "GPU exporter http://$host:9400/metrics -> probe-failed: timeout (retried at 25s: still $code)" + else + echo "GPU exporter http://$host:9400/metrics -> $code" + fi +done +# Expected: 200 on all 3 hosts (RTX 3090, RTX 5070, Strix Halo) -# Router health (via nginx on port 80) -curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health -# Expected: 200 (Router is up and responding) +# ============================================================ +# 4. ROUTER HEALTH (via nginx on port 80) — bare-200 probe +# ============================================================ +code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/health) +if [ "$code" == "000" ]; then + code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/health) + echo "Router http://192.168.68.116/health -> probe-failed: timeout (retried at 25s: still $code)" +else + echo "Router http://192.168.68.116/health -> $code" +fi +# Expected: 200 (bare-200 probe; any other status is an ALERT) -# LiteLLM health (via nginx on port 80) -curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health -# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive +# ============================================================ +# 5. LITELLM HEALTH (via nginx on port 80) — any-HTTP probe +# ============================================================ +code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116/litellm/health) +if [ "$code" == "000" ]; then + code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116/litellm/health) + echo "LiteLLM http://192.168.68.116/litellm/health -> probe-failed: timeout (retried at 25s: still $code)" +else + echo "LiteLLM http://192.168.68.116/litellm/health -> $code" +fi +# Expected: 301 → /litellm/health/liveliness (any HTTP status = ALIVE) -# PVE API liveness — probe the REAL PVE nodes on :8006, never the monitoring -# host CT 116. CT 116 runs no pveproxy, so probing it on :8006 returns 000 — -# that was the stale-vantage bug this replaces (CT 116 is the monitoring host, -# not a cluster node). Unauthenticated GET answers 401 while the API is ALIVE -# by design. Alive = ANY HTTP status (401 is the EXPECTED healthy response); -# DOWN = connection refused (000) or timeout only. +# ============================================================ +# 6. PVE API LIVENESS — any-HTTP probe (auth-gated) +# ============================================================ +# Probe the REAL PVE nodes on :8006, never the monitoring host CT 116. for node in 192.168.68.9 192.168.68.5 192.168.68.15 192.168.68.6 192.168.68.12; do - printf '%s:8006 -> %s\n' "$node" \ - "$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 5 "https://$node:8006/api2/json/version")" + code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 10 "https://$node:8006/api2/json/version") + if [ "$code" == "000" ]; then + code=$(curl -sk -o /dev/null -w '%{http_code}' --connect-timeout 25 "https://$node:8006/api2/json/version") + echo "PVE API https://$node:8006/api2/json/version -> probe-failed: timeout (retried at 25s: still $code)" + else + echo "PVE API https://$node:8006/api2/json/version -> $code" + fi done # Expected: 401 on every node (acerpve .9, ocupve .5, amdpve .15, storepve .6, minipve .12) -# A node answering 000/timeout is DOWN — flag that node. 401 is NOT a fault. +# 401 is the EXPECTED healthy response (auth-gated); 000/timeout = DOWN -# Prometheus targets -curl -s http://192.168.68.116:9090/api/v1/targets | jq '.data.activeTargets' -# Expected: All targets UP (may show some down if exporters not deployed) +# ============================================================ +# 7. PROMETHEUS TARGETS — bare-200 probe +# ============================================================ +code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:9090/api/v1/targets) +if [ "$code" == "000" ]; then + code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:9090/api/v1/targets) + echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> probe-failed: timeout (retried at 25s: still $code)" +else + echo "Prometheus http://192.168.68.116:9090/api/v1/targets -> $code" +fi +# Expected: 200 (bare-200 probe; any other status is an ALERT) -# Grafana health -curl -s http://192.168.68.116:3001/api/health | jq '{status, version}' -# Expected: {"status":"ok","version":"..."} +# ============================================================ +# 8. GRAFANA HEALTH — any-HTTP probe +# ============================================================ +code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:3001/api/health) +if [ "$code" == "000" ]; then + code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:3001/api/health) + echo "Grafana http://192.168.68.116:3001/api/health -> probe-failed: timeout (retried at 25s: still $code)" +else + echo "Grafana http://192.168.68.116:3001/api/health -> $code" +fi +# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN) -# LiteLLM metrics (Prometheus endpoint) -curl -s http://192.168.68.116:4000/metrics | head -20 -# Expected: Prometheus-formatted metrics output +# ============================================================ +# 9. LITELLM METRICS (Prometheus endpoint) — any-HTTP probe +# ============================================================ +code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 10 http://192.168.68.116:4000/metrics) +if [ "$code" == "000" ]; then + code=$(curl -s -o /dev/null -w '%{http_code}' --connect-timeout 25 http://192.168.68.116:4000/metrics) + echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> probe-failed: timeout (retried at 25s: still $code)" +else + echo "LiteLLM metrics http://192.168.68.116:4000/metrics -> $code" +fi +# Expected: 200 (any HTTP status = ALIVE; 000/timeout = DOWN) ``` **Report format**: Begin every report with the **absolute path the probe executed -from** (`pwd -P`, or the script's absolute path) so a stale-consumer report is -distinguishable from a real fault at read time. Summarize actual results from -each probe. Apply the any-HTTP-response liveness rule ONLY to the auth-gated PVE -API and LiteLLM endpoints above: only connection-refused (`000`) or timeout is -DOWN; empty output is a warning. For probes whose expected result is a bare `200` -(the authenticated Zulip POST, router `/health`), flag an alert on any unexpected -status (`401`/`403`/`5xx`) — do not summarize it as alive. A bare-`200` -expectation on the auth-gated PVE API (`401`) or LiteLLM health (`301` redirect) -is a stale expectation, not a fault. - +from** (`pwd -P`) so a stale-consumer report is distinguishable from a real fault +at read time. For each probe, print the target name, the full URL, and the HTTP +code (or failure kind with retry details). Apply the standing probe rules: any +HTTP status = ALIVE; only 000/timeout/refused = probe-failed. A redirect is not +a failure. ### Phase 1: GPU Exporters **NVIDIA (.8 and .110)**: 1. Download `nvidia_gpu_exporter` binary -2. Create systemd service `nvidia-gpu-exporter.service` -3. Start and enable - -**AMD (.15)**: -1. Create Python exporter script at `/opt/amdgpu-exporter/exporter.py` -2. Parses `amdgpu_top --json -d 1000` output -3. Exposes key metrics at `:9400/metrics` via Python http.server -4. Create systemd service -5. Start and enable - -### Phase 2: Prometheus - -1. Create `/opt/monitoring/` directory on CT 116 -2. Write `prometheus.yml` with scrape configs for all targets -3. Add to docker-compose (or separate compose file) -4. Start container - -### Phase 3: Grafana - -1. Create `/opt/monitoring/grafana/` directories -2. Provision Prometheus datasource -3. Provision GPU fleet dashboard JSON -4. Provision LiteLLM dashboard JSON -5. Add to docker-compose -6. Start container - -### Phase 4: Verification - -1. Verify all 3 GPU exporters return 200 at :9400/metrics -2. Verify Prometheus targets all UP at :9090/targets -3. Verify Grafana accessible at :3001 with dashboards -4. Verify LiteLLM metrics flowing to Prometheus -5. ~~Update nginx to proxy `/monitoring/` → Grafana~~ (NOT recommended — nginx sub-path was tried for /grafana/ and reverted per proxmox-monitor; direct :3001 access is the standard) - -## Verification Commands - -```bash -# GPU exporters -curl -s http://192.168.68.8:9400/metrics | grep nvidia -curl -s http://192.168.68.110:9400/metrics | grep nvidia -curl -s http://192.168.68.15:9400/metrics | grep amdgpu - -# Prometheus -curl -s http://192.168.68.116:9090/api/v1/targets - -# Grafana -curl -s http://192.168.68.116:3001/api/health - -# LiteLLM metrics (already live) -curl -s http://192.168.68.116:4000/metrics | head -20 -```