Files
prose-contracts/gpu-monitor.prose.md
T
abiba-bot c66d9c1e20 fix: probe-drift round 2 — stale monitoring expectations (health check, PVE API, GPU port 80, report provenance)
Second probe-drift correction pass after #65/#66/#68. All four legs were stale
consumer expectations, not live faults.

1. scripts/agent-health-check.py (v4)
   - abiba declared pi-only runtime (harness purge): Hermes-era gateway, config
     and wrapper legs are skipped instead of failing.
   - koby declared report_only (captain ruling 2026-08-17, Rule 17): every koby
     leg is detected and reported, never counted as a fleet failure or repaired.
   - koby's CT 111 mapping corrected to storepve (.6); the old amdpve mapping
     made `pct status 111` fail and read as ct-unreachable.
   - wrapper check no longer FAILs .env-based wrappers that legitimately never
     invoke infisical (koonimo).
   - keys load in main() (load_agent_keys) so the module is importable/testable.
   - every run prints absolute execution provenance (script + cwd), in the header
     and in --json.
   Before: 6 FAILURE(S). After: 0 failures, koby reported read-only.

2. infrastructure-monitoring.prose.md
   - PVE API probe repointed from CT 116 (no pveproxy, 000) to the five real
     nodes on https://<node>:8006/api2/json/version, all 401 = alive.
   - any-HTTP-response liveness rule added (401/3xx alive; 000/timeout = DOWN).
   - LiteLLM health documented as 301 -> /litellm/health/liveliness, not bare 200.

3. gpu-monitor.prose.md
   - GPU health probes on :8080 (or router /health/unified); bare port 80 on a
     GPU host is forbidden (no listener -> false DEGRADED).
   - router /health/unified 301 -> /gpu/gpu-data documented as alive.
   - port-discipline + liveness rule + direct-fallback execution step.

4. Report provenance (all contracts)
   - docs/AUTHORING-GUIDE.md documents the rule; scripts/prose-lint.sh enforces
     that any **Report format** contract states an absolute path (pwd -P).
   - provenance added to gpu-monitor, infrastructure-monitoring, proxmox-monitor.

Tests: tests/test_probe_drift.py (23 passed); prose-lint.sh clean + shellcheck
clean; CI frontmatter validation passes. Evidence with per-leg before/after and
absolute paths: docs/probe-drift-round2-evidence.md.
2026-09-10 01:23:54 +00:00

9.9 KiB

kind, name, description, agent
kind name description agent
responsibility gpu-monitor Comprehensive GPU fleet monitor — polls every subsystem (sidecars, router, LiteLLM, Strix Halo, dashboard) every 15s, renders a live HTML dashboard, checks alert thresholds, and exposes a JSON API for downstream consumers. abiba

Maintains

  • gpu-fleet-data: { gpus: array, strix: object, router: object, litellm: object, summary: object, alerts: array } — Full fleet snapshot refreshed every 15s
  • dashboard: live HTML at /root/dashboard/gpu-fleet.html, served on port 9100
  • gpu-data-api: JSON at GET /gpu-data on port 9100
  • health-endpoint: GET /health on port 9100 → 200 if cache populated, 503 if warming up

GPU Fleet Topology

┌─────────────────────────────────────────────────────────┐
│  GPU Monitor Server (:9100) — 192.168.68.24              │
│  Polls every 15s, serves dashboard + JSON API            │
└──┬──────────┬──────────┬────────────────────────────────┘
   │          │          │
   ▼          ▼          ▼
┌──────┐ ┌──────┐ ┌────────┐
│.8:8080│ │.110  │ │.116:80 │
│RTX3090│ │:8080 │ │nginx   │
│gemma  │ │RTX5070│ │router  │
└──────┘ │qwen27B│ │LiteLLM │
         └──────┘ │dashboard│
                  └────────┘

Note: JSON sidecar exporters at :8090 were never deployed on any GPU host. Router falls back to GPU /health direct probe. Monitor should use router /health/unified as source of truth for GPU status. Strix Halo :8080 is firewalled to .116 only — monitor on .24 cannot poll .15:8080 directly; must go through router on .116.

PORT RULE (verified 2026-09-10): GPU per-host health lives on :8080 (http://<gpu-host>:8080/health); Prometheus GPU exporters live on :9400. There is NO listener on bare port 80 for any GPU host — http://192.168.68.8/health and http://192.168.68.110/health answer 000. Never use a bare-port-80 probe as a GPU liveness signal: on 2026-09-09 that produced three false DEGRADED — GPU-rtx3090 000, GPU-rtx5070 000 rounds while http://192.168.68.8:8080/health and http://192.168.68.110:8080/health answered 200. Port 80 is valid only on the router (.116), never on a GPU host.

Subsystems Polled

Subsystem Endpoint Frequency Metrics
GPU Status (all, via router) http://192.168.68.116/health/unified 15s models, CB, scores, GPU status (router probes each GPU /health directly) — 301 → /gpu/gpu-data is alive
GPU .8 (RTX 3090) health http://192.168.68.8:8080/health 15s direct liveness fallback — :8080 ONLY, never bare port 80
GPU .110 (RTX 5070) health http://192.168.68.110:8080/health 15s direct liveness fallback — :8080 ONLY, never bare port 80
Router (unified) http://192.168.68.116/health/unified 15s models, CB, scores, GPU status (301 → /gpu/gpu-data = alive)
Router (basic) http://192.168.68.116/health 15s basic aliveness
LiteLLM http://192.168.68.116/litellm/health 15s proxy health, model count
Strix Halo http://192.168.68.116/health/unified (router) 15s Strix Halo status via router — cannot poll .15:8080 directly (firewalled to .116 only)
Dashboard http://192.168.68.116/dashboard/ 15s harness-dashboard aliveness

Alert Delivery

All alerts are sent to #agent-hub stream topics:

  • alerts-gpu — GPU fleet alerts (temp, VRAM, utilization)
  • alerts-pm2 — PM2 process health (restarts, crashes)
  • alerts-infra — Infrastructure health (LiteLLM, Zulip, containers)

This replaces the previous DM-only delivery. All agents on the mesh can see and respond to alerts.

Alert Thresholds

Liveness rule (any-HTTP-response)

A probe is ALIVE if the endpoint returns ANY HTTP status — including redirects and auth challenges. A bare 200 is not required. DOWN = connection refused (000) or timeout only. This is the same rule zulip-health adopted for Tanko (loopback :3080 + public URL). Applied here: the router's /health/unified answers 301 Moved Permanently → /gpu/gpu-data (the same payload), so 301 is healthy and a bare-200 expectation would false-alarm. Statuses outside the expected set on an otherwise-alive endpoint are reported as a warning, never as DOWN.

Metric Warning Critical
GPU Temp >80°C >90°C
VRAM Usage >90% >95%
GPU Util >95% >98%
Sidecar Unreachable — info (sidecars not deployed — use router /health/unified)
Model Down — critical (circuit breaker open)

JSON API Response Schema (/gpu-data)

{
  "gpus": [{ "name", "gpu_name", "temp_c", "vram_used_mb", "vram_total_mb",
             "gpu_util_pct", "power_w", "power_limit_w", "fan_pct",
             "models", "hostname", "active_requests", "health_score",
             "circuit_open" }],
  "strix": { "status", "cpu_load", "llama_health" },
  "router": { "available_models", "circuit_breaker", "gpus", "scores",
              "status", "redis", "_basic" },
  "litellm": { ... },
  "dashboard": { "reachable": bool },
  "summary": { "fleet_status", "models_available", "models_total",
               "circuit_breakers_open", "gpu_count", "gpu_errors",
               "strix_running", "router_reachable", "litellm_reachable" },
  "alerts": [{ "gpu", "metric", "level", "value", "threshold" }],
  "updated": "ISO8601",
  "monitor_version": "2.0.0"
}

Operations

list

curl http://localhost:9100/gpu-data | jq — Full fleet status

check-health

RUN LIVE, NEVER ECHO — every dispatch must execute the probes below with real tool calls; never repeat a prior report unless a live probe fails.

# Provenance — run first; paste the absolute path into the report
pwd -P

# GPU Monitor health
 curl http://localhost:9100/health | jq
# Expected: 200 with {"status": "healthy", "cache_age_seconds": <n>}

# GPU host health — DIRECT on :8080. NEVER probe bare port 80 on a GPU host:
# http://192.168.68.8/health has no listener and returns 000 → false DEGRADED.
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.8:8080/health
# Expected: 200 (alive = any HTTP status; 000/timeout = DOWN)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.110:8080/health
# Expected: 200 (alive = any HTTP status; 000/timeout = DOWN)

# Router unified health (source of truth; 301 → /gpu/gpu-data is HEALTHY)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health/unified
# Expected: 301 (or 200 after following the redirect) — any HTTP status = alive

# Router basic health (via nginx on port 80 — router .116 only, never a GPU host)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/health
# Expected: 200

# LiteLLM health (via nginx on port 80)
curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/litellm/health
# Expected: 301 → /litellm/health/liveliness (200 after redirect) — any HTTP status = alive

# Dashboard
 curl -s -o /dev/null -w '%{http_code}' http://192.168.68.116/dashboard/
# Expected: 200

Report format: Begin every report with the absolute path the probe executed from (pwd -P, or the monitor script's absolute path) so a stale-consumer report is distinguishable from a real fault at read time. Summarize actual results from each probe. Apply the any-HTTP-response liveness rule above: only connection-refused (000) or timeout is DOWN. Flag an alert only when a probe is DOWN, or when an alive endpoint returns an unexpected status. Never probe a GPU host on bare port 80.

view-dashboard

Open http://localhost:9100/ in browser — Live HTML dashboard

restart

pkill -f gpu-monitor-server.py
python3 /root/scripts/gpu-monitor-server.py &

Or via PM2: pm2 restart gpu-monitor

check-router

The router health is accessed through nginx on port 80 on the router (.116) — NOT port 9000 directly, and NOT bare port 80 on a GPU host. curl http://192.168.68.116/health/unified — Router unified health via nginx proxy; answers 301 → /gpu/gpu-data (same payload) = alive curl http://192.168.68.116:9000/health/unified — ❌ WILL FAIL (port bound to 127.0.0.1 only) curl http://192.168.68.8/health — ❌ NEVER USE (GPU host, no port-80 listener → false 000/DEGRADED)

Configuration Files

File Host Purpose
/root/scripts/gpu-monitor-server.py pi (.24) Monitor server v2.0.0
/root/dashboard/gpu-fleet.html pi (.24) Live HTML dashboard
/etc/nginx/nginx.conf CT 116 Routes /health/* → router

Execution

Port discipline: probe GPU hosts on :8080 (or the router's /health/unified); probe port 80 only on the router (.116). Never bare port 80 on a GPU host.

  1. Poll router (every 15s): GET .116/health/unified — single source of truth for all GPU status (router probes each GPU /health directly via sidecar fallback). 301 → /gpu/gpu-data counts as alive.
  2. Fallback direct GPU probe (only if router /health/unified is DOWN): GET http://192.168.68.8:8080/health and http://192.168.68.110:8080/health — :8080 only, never bare port 80.
  3. Poll router (every 15s): GET .116/health via nginx:80
  4. Poll LiteLLM (every 15s): GET .116/litellm/health via nginx:80
  5. Poll Strix (every 15s): via router /health/unified (cannot poll .15:8080 directly — firewalled to .116 only)
  6. Poll dashboard (every 15s): GET .116/dashboard/
  7. Check alerts: Compare metrics against thresholds
  8. Compute summary: Fleet-wide health aggregation
  9. Render dashboard: Generate HTML at /root/dashboard/gpu-fleet.html
  10. Serve API: HTTP server on port 9100
  11. Repeat every 15 seconds