10 KiB
kind, name, description, agent
| kind | name | description | agent |
|---|---|---|---|
| responsibility | gpu-monitor | Comprehensive GPU fleet monitor — polls every subsystem (sidecars, router, LiteLLM, Strix Halo, dashboard) every 15s, renders a live HTML dashboard, checks alert thresholds, and exposes a JSON API for downstream consumers. | abiba |
Maintains
- gpu-fleet-data: { gpus: array, strix: object, router: object, litellm: object, summary: object, alerts: array } — Full fleet snapshot refreshed every 15s
- dashboard: live HTML at
/root/dashboard/gpu-fleet.html, served on port 9100 - gpu-data-api: JSON at
GET /gpu-dataon port 9100 - health-endpoint:
GET /healthon port 9100 → 200 if cache populated, 503 if warming up
GPU Fleet Topology
┌─────────────────────────────────────────────────────────┐
│ GPU Monitor Server (:9100) — 192.168.68.24 │
│ Polls every 15s, serves dashboard + JSON API │
└──┬──────────┬──────────┬────────────────────────────────┘
│ │ │
▼ ▼ ▼
┌──────┐ ┌──────┐ ┌────────┐
│.8:8080│ │.110 │ │.116:80 │
│RTX3090│ │:8080 │ │nginx │
│gemma │ │RTX5070│ │router │
└──────┘ │qwen27B│ │LiteLLM │
└──────┘ │dashboard│
└────────┘
Note: JSON sidecar exporters at :8090 were never deployed on any GPU host. Router falls back to GPU /health direct probe. Monitor should use router /health/unified as source of truth for GPU status. Strix Halo :8080 is firewalled to .116 only — monitor on .24 cannot poll .15:8080 directly; must go through router on .116.
PORT RULE (verified 2026-09-10): GPU per-host health lives on :8080
(http://<gpu-host>:8080/health); Prometheus GPU exporters live on :9400.
There is NO listener on bare port 80 for any GPU host — http://192.168.68.8/health
and http://192.168.68.110/health answer 000. Never use a bare-port-80 probe
as a GPU liveness signal: on 2026-09-09 that produced three false
DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000 rounds while
http://192.168.68.8:8080/health and http://192.168.68.110:8080/health
answered 200. Port 80 is valid only on the router (.116), never on a GPU host.
Subsystems Polled
| Subsystem | Endpoint | Frequency | Metrics |
|---|---|---|---|
| GPU Status (all, via router) | http://192.168.68.116/health/unified |
15s | models, CB, scores, GPU status (router probes each GPU /health directly) — 301 → /gpu/gpu-data is alive |
| GPU .8 (RTX 3090) health | http://192.168.68.8:8080/health |
15s | direct liveness fallback — :8080 ONLY, never bare port 80 |
| GPU .110 (RTX 5070) health | http://192.168.68.110:8080/health |
15s | direct liveness fallback — :8080 ONLY, never bare port 80 |
| Router (unified) | http://192.168.68.116/health/unified |
15s | models, CB, scores, GPU status (301 → /gpu/gpu-data = alive) |
| Router (basic) | http://192.168.68.116/health |
15s | basic aliveness |
| LiteLLM | http://192.168.68.116/litellm/health |
15s | proxy health, model count |
| Strix Halo | http://192.168.68.116/health/unified (router) |
15s | Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only) |
| Dashboard | http://192.168.68.116/dashboard/ |
15s | harness-dashboard aliveness |
Alert Delivery
All alerts are sent to #agent-hub stream topics:
alerts-gpu— GPU fleet alerts (temp, VRAM, utilization)alerts-pm2— PM2 process health (restarts, crashes)alerts-infra— Infrastructure health (LiteLLM, Zulip, containers)
This replaces the previous DM-only delivery. All agents on the mesh can see and respond to alerts.
Alert Thresholds
Liveness rule (scoped)
The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness
endpoints, where any HTTP answer proves a listener is up. Applied here: the
router's /health/unified answers 301 Moved Permanently → /gpu/gpu-data
(the same payload) and LiteLLM's /litellm/health answers 301 →
/litellm/health/liveliness. For those endpoints a probe is ALIVE on
ANY HTTP status — 3xx redirects and 401/403 auth challenges included —
and DOWN = connection refused (000) or timeout only. Same scoped rule as
zulip-health (Tanko) and infrastructure-monitoring.
Probes whose success condition is specifically a bare 200 are NOT covered by
the any-HTTP rule. On those — the GPU :8080/health endpoints, the router
/health, and the dashboard — an unexpected status (401/403, 5xx, or
anything other than the expected 200) is an ALERT, not "alive".
| Metric | Warning | Critical |
|---|---|---|
| GPU Temp | >80°C | >90°C |
| VRAM Usage | >90% | >95% |
| GPU Util | >95% | >98% |
| Sidecar Unreachable | — | info (sidecars not deployed — use router /health/unified) |
| Model Down | — | critical (circuit breaker open) |
JSON API Response Schema (/gpu-data)
{
"gpus": [{ "name", "gpu_name", "temp_c", "vram_used_mb", "vram_total_mb",
"gpu_util_pct", "power_w", "power_limit_w", "fan_pct",
"models", "hostname", "active_requests", "health_score",
"circuit_open" }],
"strix": { "status", "cpu_load", "llama_health" },
"router": { "available_models", "circuit_breaker", "gpus", "scores",
"status", "redis", "_basic" },
"litellm": { ... },
"dashboard": { "reachable": bool },
"summary": { "fleet_status", "models_available", "models_total",
"circuit_breakers_open", "gpu_count", "gpu_errors",
"strix_running", "router_reachable", "litellm_reachable" },
"alerts": [{ "gpu", "metric", "level", "value", "threshold" }],
"updated": "ISO8601",
"monitor_version": "2.0.0"
}
Operations
list
curl http://localhost:9100/gpu-data | jq — Full fleet status
check-health
RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.
# Provenance — run first; paste the absolute path into the report
pwd -P
# GPU Monitor health
curl http://localhost:9100/health | jq
# Expected: 200 with {"status": "healthy", "cache_age_seconds": <n>}
# GPU host health — DIRECT on :8080. NEVER probe bare port 80 on a GPU host:
# http://192.168.68.8/health has no listener and returns 000 → false DEGRADED.
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
# Router unified health (source of truth; 301 → /gpu/gpu-data is HEALTHY)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified
# Expected: 301 (or 200 after following the redirect) — any HTTP status = alive
# Router basic health (via nginx on port 80 — router .116 only, never a GPU host)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
# LiteLLM health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive
# Dashboard
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
Report format: Begin every report with the absolute path the probe executed
from (pwd -P, or the monitor script's absolute path) so a stale-consumer
report is distinguishable from a real fault at read time. Summarize actual
results from each probe. Apply the scoped liveness rule above: on auth-gated
endpoints only connection-refused (000) or timeout is DOWN; on bare-200 probes
any other status is an alert. Never probe a GPU host on bare port 80.
view-dashboard
Open http://localhost:9100/ in browser — Live HTML dashboard
restart
pkill -f gpu-monitor-server.py
python3 /root/scripts/gpu-monitor-server.py &
Or via PM2: pm2 restart gpu-monitor
check-router
The router health is accessed through nginx on port 80 on the router
(.116) — NOT port 9000 directly, and NOT bare port 80 on a GPU host.
curl http://192.168.68.116/health/unified — Router unified health via nginx proxy; answers 301 → /gpu/gpu-data (same payload) = alive
curl http://192.168.68.116:9000/health/unified — ❌ WILL FAIL (port bound to 127.0.0.1 only)
curl http://192.168.68.8/health — ❌ NEVER USE (GPU host, no port-80 listener → false 000/DEGRADED)
Configuration Files
| File | Host | Purpose |
|---|---|---|
/root/scripts/gpu-monitor-server.py |
pi (.24) | Monitor server v2.0.0 |
/root/dashboard/gpu-fleet.html |
pi (.24) | Live HTML dashboard |
/etc/nginx/nginx.conf |
CT 116 | Routes /health/* → router |
Execution
Port discipline: probe GPU hosts on :8080 (or the router's
/health/unified); probe port 80 only on the router (.116). Never bare port 80
on a GPU host.
- Poll router (every 15s): GET .116/health/unified — single source of truth for all GPU status (router probes each GPU /health directly via sidecar fallback).
301→/gpu/gpu-datacounts as alive. - Fallback direct GPU probe (only if router /health/unified is DOWN): GET
http://192.168.68.8:8080/healthandhttp://192.168.68.110:8080/health— :8080 only, never bare port 80. - Poll router (every 15s): GET .116/health via nginx:80
- Poll LiteLLM (every 15s): GET .116/litellm/health via nginx:80
- Poll Strix (every 15s): via router /health/unified (cannot poll .15:8080 directly — firewalled to .116 only)
- Poll dashboard (every 15s): GET .116/dashboard/
- Check alerts: Compare metrics against thresholds
- Compute summary: Fleet-wide health aggregation
- Render dashboard: Generate HTML at /root/dashboard/gpu-fleet.html
- Serve API: HTTP server on port 9100
- Repeat every 15 seconds