--- kind: responsibility name: gpu-monitor description: > Comprehensive GPU fleet monitor — polls every subsystem (sidecars, router, LiteLLM, Strix Halo, dashboard) every 15s, renders a live HTML dashboard, checks alert thresholds, and exposes a JSON API for downstream consumers. agent: abiba --- ## Maintains - gpu-fleet-data: { gpus: array, strix: object, router: object, litellm: object, summary: object, alerts: array } — Full fleet snapshot refreshed every 15s - dashboard: live HTML at `/root/dashboard/gpu-fleet.html`, served on port 9100 - gpu-data-api: JSON at `GET /gpu-data` on port 9100 - health-endpoint: `GET /health` on port 9100 → 200 if cache populated, 503 if warming up ## GPU Fleet Topology ``` ┌─────────────────────────────────────────────────────────┐ │ GPU Monitor Server (:9100) — 192.168.68.24 │ │ Polls every 15s, serves dashboard + JSON API │ └──┬──────────┬──────────┬────────────────────────────────┘ │ │ │ ▼ ▼ ▼ ┌──────┐ ┌──────┐ ┌────────┐ │.8:8080│ │.110 │ │.116:80 │ │RTX3090│ │:8080 │ │nginx │ │gemma │ │RTX5070│ │router │ └──────┘ │qwen27B│ │LiteLLM │ └──────┘ │dashboard│ └────────┘ ``` Note: JSON sidecar exporters at :8090 were never deployed on any GPU host. Router falls back to GPU /health direct probe. Monitor should use router /health/unified as source of truth for GPU status. Strix Halo :8080 is firewalled to .116 only — monitor on .24 cannot poll .15:8080 directly; must go through router on .116. **PORT RULE (verified 2026-09-10):** GPU per-host health lives on **:8080** (`http://:8080/health`); Prometheus GPU exporters live on **:9400**. There is NO listener on bare port 80 for any GPU host — `http://192.168.68.8/health` and `http://192.168.68.110/health` answer `000`. Never use a bare-port-80 probe as a GPU liveness signal: on 2026-09-09 that produced three false `DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000` rounds while `http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health` answered `200`. Port 80 is valid only on the router (.116), never on a GPU host. ### Subsystems Polled | Subsystem | Endpoint | Frequency | Metrics | |-----------|----------|-----------|---------| | GPU Status (all, via router) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (router probes each GPU /health directly) — `301` → `/gpu/gpu-data` is **alive** | | GPU .8 (RTX 3090) health | `http://192.168.68.8:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** | | GPU .110 (RTX 5070) health | `http://192.168.68.110:8080/health` | 15s | direct liveness fallback — **:8080 ONLY, never bare port 80** | | Router (unified) | `http://192.168.68.116/health/unified` | 15s | models, CB, scores, GPU status (`301` → `/gpu/gpu-data` = alive) | | Router (basic) | `http://192.168.68.116/health` | 15s | basic aliveness | | LiteLLM | `http://192.168.68.116/litellm/health` | 15s | proxy health, model count | | Strix Halo | `http://192.168.68.116/health/unified` (router) | 15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) | | Dashboard | `http://192.168.68.116/dashboard/` | 15s | harness-dashboard aliveness | ### Alert Delivery All alerts are sent to `#agent-hub` stream topics: - `alerts-gpu` — GPU fleet alerts (temp, VRAM, utilization) - `alerts-pm2` — PM2 process health (restarts, crashes) - `alerts-infra` — Infrastructure health (LiteLLM, Zulip, containers) This replaces the previous DM-only delivery. All agents on the mesh can see and respond to alerts. ## Alert Thresholds ### Liveness rule (scoped) The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness endpoints, where any HTTP answer proves a listener is up. Applied here: the router's `/health/unified` answers `301 Moved Permanently` → `/gpu/gpu-data` (the same payload) and LiteLLM's `/litellm/health` answers `301` → `/litellm/health/liveliness`. For those endpoints a probe is **ALIVE** on **ANY** HTTP status — `3xx` redirects and `401`/`403` auth challenges included — and **DOWN = connection refused (`000`) or timeout only**. Same scoped rule as zulip-health (Tanko) and infrastructure-monitoring. Probes whose success condition is specifically a bare `200` are NOT covered by the any-HTTP rule. On those — the GPU `:8080/health` endpoints, the router `/health`, and the dashboard — an unexpected status (`401`/`403`, `5xx`, or anything other than the expected `200`) is an **ALERT**, not "alive". | Metric | Warning | Critical | |--------|---------|----------| | GPU Temp | >80°C | >90°C | | VRAM Usage | >90% | >95% | | GPU Util | >95% | >98% | | Sidecar Unreachable | — | info (sidecars not deployed — use router /health/unified) | | Model Down | — | critical (circuit breaker open) | ### JSON API Response Schema (/gpu-data) ```json { "gpus": [{ "name", "gpu_name", "temp_c", "vram_used_mb", "vram_total_mb", "gpu_util_pct", "power_w", "power_limit_w", "fan_pct", "models", "hostname", "active_requests", "health_score", "circuit_open" }], "strix": { "status", "cpu_load", "llama_health" }, "router": { "available_models", "circuit_breaker", "gpus", "scores", "status", "redis", "_basic" }, "litellm": { ... }, "dashboard": { "reachable": bool }, "summary": { "fleet_status", "models_available", "models_total", "circuit_breakers_open", "gpu_count", "gpu_errors", "strix_running", "router_reachable", "litellm_reachable" }, "alerts": [{ "gpu", "metric", "level", "value", "threshold" }], "updated": "ISO8601", "monitor_version": "2.0.0" } ``` ## Operations ### list `curl http://localhost:9100/gpu-data | jq` — Full fleet status ### check-health **RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.** ```bash # Provenance — run first; paste the absolute path into the report pwd -P # GPU Monitor health curl http://localhost:9100/health | jq # Expected: 200 with {"status": "healthy", "cache_age_seconds": } # GPU host health — DIRECT on :8080. NEVER probe bare port 80 on a GPU host: # http://192.168.68.8/health has no listener and returns 000 → false DEGRADED. curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health # Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health # Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) # Router unified health (source of truth; 301 → /gpu/gpu-data is HEALTHY) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified # Expected: 301 (or 200 after following the redirect) — any HTTP status = alive # Router basic health (via nginx on port 80 — router .116 only, never a GPU host) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health # Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) # LiteLLM health (via nginx on port 80) curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health # Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive # Dashboard curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/ # Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN) ``` **Report format**: Begin every report with the **absolute path the probe executed from** (`pwd -P`, or the monitor script's absolute path) so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual results from each probe. Apply the scoped liveness rule above: on auth-gated endpoints only connection-refused (`000`) or timeout is DOWN; on bare-200 probes any other status is an alert. Never probe a GPU host on bare port 80. ### view-dashboard Open `http://localhost:9100/` in browser — Live HTML dashboard ### restart ```bash pkill -f gpu-monitor-server.py python3 /root/scripts/gpu-monitor-server.py & ``` Or via PM2: `pm2 restart gpu-monitor` ### check-router The router health is accessed through nginx on port 80 on the **router** (.116) — NOT port 9000 directly, and NOT bare port 80 on a GPU host. `curl http://192.168.68.116/health/unified` — Router unified health via nginx proxy; answers `301` → `/gpu/gpu-data` (same payload) = alive `curl http://192.168.68.116:9000/health/unified` — ❌ WILL FAIL (port bound to 127.0.0.1 only) `curl http://192.168.68.8/health` — ❌ NEVER USE (GPU host, no port-80 listener → false `000`/DEGRADED) ## Configuration Files | File | Host | Purpose | |------|------|---------| | `/root/scripts/gpu-monitor-server.py` | pi (.24) | Monitor server v2.0.0 | | `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard | | `/etc/nginx/nginx.conf` | CT 116 | Routes /health/* → router | ## Execution **Port discipline:** probe GPU hosts on `:8080` (or the router's `/health/unified`); probe port 80 only on the router (.116). Never bare port 80 on a GPU host. 1. **Poll router** (every 15s): GET .116/health/unified — single source of truth for all GPU status (router probes each GPU /health directly via sidecar fallback). `301` → `/gpu/gpu-data` counts as alive. 2. **Fallback direct GPU probe** (only if router /health/unified is DOWN): GET `http://192.168.68.8:8080/health` and `http://192.168.68.110:8080/health` — **:8080 only, never bare port 80**. 3. **Poll router** (every 15s): GET .116/health via nginx:80 4. **Poll LiteLLM** (every 15s): GET .116/litellm/health via nginx:80 5. **Poll Strix** (every 15s): via router /health/unified (cannot poll .15:8080 directly — firewalled to .116 only) 6. **Poll dashboard** (every 15s): GET .116/dashboard/ 7. **Check alerts**: Compare metrics against thresholds 8. **Compute summary**: Fleet-wide health aggregation 9. **Render dashboard**: Generate HTML at /root/dashboard/gpu-fleet.html 10. **Serve API**: HTTP server on port 9100 11. **Repeat** every 15 seconds