Files
prose-contracts/gpu-monitor.prose.md
T
root 78b501798f
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
no-mistakes(document): Sweep residual gemma labels; align compression rule contradiction
2026-09-12 16:54:41 +00:00

10 KiB

kind, name, description, agent
kind name description agent
responsibility gpu-monitor Comprehensive GPU fleet monitor — polls every subsystem (sidecars, LiteLLM, Strix Halo, dashboard) every 15s, renders a live HTML dashboard, checks alert thresholds, and exposes a JSON API for downstream consumers. abiba

Maintains

  • gpu-fleet-data: { gpus: array, strix: object, router: object, litellm: object, summary: object, alerts: array } — Full fleet snapshot refreshed every 15s
  • dashboard: live HTML at /root/dashboard/gpu-fleet.html, served on port 9100
  • gpu-data-api: JSON at GET /gpu-data on port 9100
  • health-endpoint: GET /health on port 9100 → 200 if cache populated, 503 if warming up

GPU Fleet Topology

┌─────────────────────────────────────────────────────────┐
│  GPU Monitor Server (:9100) — 192.168.68.24              │
│  Polls every 15s, serves dashboard + JSON API            │
└──┬──────────┬──────────┬────────────────────────────────┘
   │          │          │
   ▼          ▼          ▼
┌──────┐ ┌──────┐ ┌────────┐
│.8:8080│ │.110  │ │.116:80 │
│RTX3090│ │:8080 │ │nginx   │
│qwen   │ │RTX5070│ │router  │
└──────┘ │vision │ │LiteLLM │
         └──────┘ │dashboard│
                  └────────┘

Note: JSON sidecar exporters at :8090 were never deployed on any GPU host. Router falls back to GPU /health direct probe. Monitor should use router /health/unified as source of truth for GPU status. Strix Halo :8080 is firewalled to .116 only — monitor on .24 cannot poll .15:8080 directly; must go through router on .116.

PORT RULE (verified 2026-09-10): GPU per-host health lives on :8080 (http://<gpu-host>:8080/health); Prometheus GPU exporters live on :9400. There is NO listener on bare port 80 for any GPU host — http://192.168.68.8/health and http://192.168.68.110/health answer 000. Never use a bare-port-80 probe as a GPU liveness signal: on 2026-09-09 that produced three false DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000 rounds while http://192.168.68.8:8080/health and http://192.168.68.110:8080/health answered 200. Port 80 is valid only on the harness host (.116), never on a GPU host.

Subsystems Polled

Subsystem Endpoint Frequency Metrics
GPU Status (all, via fleet API) http://192.168.68.116/gpu/gpu-data 15s models, CB, scores, GPU status from gpu-monitor on .24:9100
GPU .8 (RTX 3090) health http://192.168.68.8:8080/health 15s direct liveness fallback — :8080 ONLY, never bare port 80
GPU .110 (RTX 5070) health http://192.168.68.110:8080/health 15s direct liveness fallback — :8080 ONLY, never bare port 80
Fleet (unified) http://192.168.68.116/health/unified 15s nginx 301 → /gpu/gpu-data served by gpu-monitor = alive (router decommissioned 2026-09-11)
Harness (basic) http://192.168.68.116/health 15s nginx → LiteLLM /health/liveliness
LiteLLM http://192.168.68.116/litellm/health 15s proxy health, model count
Strix Halo http://192.168.68.116/health/unified (nginx → fleet API) 15s Strix Halo status via gpu-monitor — cannot poll .15:8080 directly (firewalled to .116 only)
Dashboard http://192.168.68.116/dashboard/ 15s harness-dashboard aliveness

Alert Delivery

All alerts are sent to #agent-hub stream topics:

  • alerts-gpu — GPU fleet alerts (temp, VRAM, utilization)
  • alerts-pm2 — PM2 process health (restarts, crashes)
  • alerts-infra — Infrastructure health (LiteLLM, Zulip, containers)

This replaces the previous DM-only delivery. All agents on the mesh can see and respond to alerts.

Alert Thresholds

Liveness rule (scoped)

The any-HTTP-response rule applies ONLY to redirect/auth-gated liveness endpoints, where any HTTP answer proves a listener is up. Applied here: nginx's /health/unified answers 301 Moved Permanently → /gpu/gpu-data (the same payload) and LiteLLM's /litellm/health answers 301 → /litellm/health/liveliness. For those endpoints a probe is ALIVE on ANY HTTP status — 3xx redirects and 401/403 auth challenges included — and DOWN = connection refused (000) or timeout only. Same scoped rule as zulip-health (Tanko) and infrastructure-monitoring.

Probes whose success condition is specifically a bare 200 are NOT covered by the any-HTTP rule. On those — the GPU :8080/health endpoints, nginx /health, and the dashboard — an unexpected status (401/403, 5xx, or anything other than the expected 200) is an ALERT, not "alive".

Metric Warning Critical
GPU Temp >80°C >90°C
VRAM Usage >90% >95%
GPU Util >95% >98%
Sidecar Unreachable — info (sidecars not deployed — use gpu-monitor /gpu-data)
Model Down — critical (circuit breaker open)

JSON API Response Schema (/gpu-data)

{
  "gpus": [{ "name", "gpu_name", "temp_c", "vram_used_mb", "vram_total_mb",
             "gpu_util_pct", "power_w", "power_limit_w", "fan_pct",
             "models", "hostname", "active_requests", "health_score",
             "circuit_open" }],
  "strix": { "status", "cpu_load", "llama_health" },
  "router": { "available_models", "circuit_breaker", "gpus", "scores",
              "status", "redis", "_basic" },
  "litellm": { ... },
  "dashboard": { "reachable": bool },
  "summary": { "fleet_status", "models_available", "models_total",
               "circuit_breakers_open", "gpu_count", "gpu_errors",
               "strix_running", "router_reachable", "litellm_reachable" },
  "alerts": [{ "gpu", "metric", "level", "value", "threshold" }],
  "updated": "ISO8601",
  "monitor_version": "2.0.0"
}

Operations

list

curl http://localhost:9100/gpu-data | jq — Full fleet status

check-health

RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.

# Provenance — run first; paste the absolute path into the report
pwd -P

# GPU Monitor health
 curl http://localhost:9100/health | jq
# Expected: 200 with {"status": "healthy", "cache_age_seconds": <n>}

# GPU host health — DIRECT on :8080. NEVER probe bare port 80 on a GPU host:
# http://192.168.68.8/health has no listener and returns 000 → false DEGRADED.
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)

# Router unified health (source of truth; 301 → /gpu/gpu-data is HEALTHY)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified
# Expected: 301 (or 200 after following the redirect) — any HTTP status = alive

# Router basic health (via nginx on port 80 — router .116 only, never a GPU host)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)

# LiteLLM health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive

# Dashboard
 curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/
# Expected: 200 (bare-200 probe: any other status is an alert; 000/timeout = DOWN)

Report format: Begin every report with the absolute path the probe executed from (pwd -P, or the monitor script's absolute path) so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual results from each probe. Apply the scoped liveness rule above: on auth-gated endpoints only connection-refused (000) or timeout is DOWN; on bare-200 probes any other status is an alert. Never probe a GPU host on bare port 80.

view-dashboard

Open http://localhost:9100/ in browser — Live HTML dashboard

restart

pkill -f gpu-monitor-server.py
python3 /root/scripts/gpu-monitor-server.py &

Managed by systemd (verified 2026-09-11): systemctl restart gpu-monitor

check-fleet

Fleet health is accessed through nginx on port 80 on the harness host (.116) — NOT port 9000 (router decommissioned 2026-09-11), and NOT bare port 80 on a GPU host. curl http://192.168.68.116/health/unified — nginx answers 301 → /gpu/gpu-data (fleet monitor payload) = alive curl http://192.168.68.8/health — ❌ NEVER USE (GPU host, no port-80 listener → false 000/DEGRADED)

Configuration Files

File Host Purpose
/root/scripts/gpu-monitor-server.py pi (.24) Monitor server v2.0.0
/root/dashboard/gpu-fleet.html pi (.24) Live HTML dashboard
/etc/nginx/nginx.conf CT 116 Routes /health/* → router

Execution

Port discipline: probe GPU hosts on :8080 (or the router's /health/unified); probe port 80 only on the router (.116). Never bare port 80 on a GPU host.

  1. Poll router (every 15s): GET .116/health/unified — single source of truth for all GPU status (router probes each GPU /health directly via sidecar fallback). 301 → /gpu/gpu-data counts as alive.
  2. Fallback direct GPU probe (only if router /health/unified is DOWN): GET http://192.168.68.8:8080/health and http://192.168.68.110:8080/health — :8080 only, never bare port 80.
  3. Poll router (every 15s): GET .116/health via nginx:80
  4. Poll LiteLLM (every 15s): GET .116/litellm/health via nginx:80
  5. Poll Strix (every 15s): via router /health/unified (cannot poll .15:8080 directly — firewalled to .116 only)
  6. Poll dashboard (every 15s): GET .116/dashboard/
  7. Check alerts: Compare metrics against thresholds
  8. Compute summary: Fleet-wide health aggregation
  9. Render dashboard: Generate HTML at /root/dashboard/gpu-fleet.html
  10. Serve API: HTTP server on port 9100
  11. Repeat every 15 seconds