GPU fleet self-healing — detects anomalies, applies remediation, tracks benchmarks, and predicts failures before they happen. Extends gpu-monitor (v2.1.0) with active remediation rules, Prometheus metrics consumption, VRAM trend analysis, and predictive alerting. UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps. Router (port 9000) DECOMMISSIONED 2026-09-11; references replaced with direct GPU routing. Benchmark baselines refreshed to live values. Prometheus exporters removed — not deployed; fall back to direct sidecar probes. Stable role-based aliases (strix-moe, gpu-dense, gpu-vision) from gpu-fleet.
abiba
gpu-monitor.prose.md (live data source on .24:9100)
gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
gpu-monitor:function — Live fleet data from localhost:9100/gpu-data
Direct sidecar probe access to all GPU hosts (:8080/health)
SSH access to GPU hosts for restart operations
Continuity
Self-driven: check every 60 seconds against GPU monitor data
Also wakes on gpu-fleet health degradation
On fix: verify with benchmark inference test before declaring resolved
Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
Current Fleet Baseline (2026-07-18)
Alias
GPU
Host
Model
VRAM
Ctx
tok/s
Role
strix-moe
Strix Halo 64GB
ct15 (.15:8080)
qwen3.6-35B-udq4
~10/64GB (16%)
128K
62.9
Compression, summarization, long docs
Key notes:
All models use direct GPU routing via LiteLLM (api_key: not-needed). Router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
Stable aliases (gpu-dense, gpu-vision, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. The retired names gpu-light and gemma-4-12b were superseded by gpu-vision on 2026-09-12 and no longer resolve (400 Invalid model name).
The RTX 5070 is the fastest endpoint per token — gpu-vision is its canonical alias. Route vision/web/light work there first. The RTX 5070 model is multimodal (image+text).
Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
Remediation Rules
Rule 1: GPU Temperature Critical (>85°C for >2 min)
Detect: Any GPU temp >85°C sustained for 2+ consecutive polls
Fix:
Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
Redirect new requests to cooler GPUs via LiteLLM fallback chains (gpu-vision → gpu-dense, gpu-dense → gpu-vision)
If all GPUs hot, alert about cooling infrastructure
Detect: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s
RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s
Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s
Fix:
Check GPU utilization — if >90%, other process is competing
Check power limit — if throttled, restore to max
Check thermal — if hot, apply Rule 1
Verify: Benchmark returns to within 10% of baseline
Escalate after: persistent regression → possible hardware degradation
Rule 5: Circuit Breaker Stuck Open
Detect: Circuit breaker open >10 minutes with GPU reporting healthy
Note: Router (port 9000) was decommissioned 2026-09-11. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
Fix:
Verify GPU /health returns 200 on direct port (:8080)
If GPU healthy, alert but do NOT reset via router API (decommissioned 2026-09-11)
Note: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
Weights are not restated here — the live syslog-auto pool weights and rpm caps live in CT 116 /opt/inference-harness/litellm_config.yaml, the single source of truth.
Fix:
Alert if any GPU is handling workload outside its designated role
Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-vision, strix-moe)
Track per-GPU request distribution via LiteLLM spend logs
Verify: Each GPU's request pattern matches its designated role within 24h
Escalate: If role mismatch persists >48h → agent alias audit needed
Execution
-- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor
endpoint: "http://localhost:9100/gpu-data"
-- Phase 2: Evaluate each GPU against remediation rules
let actions = []
for gpu in fleet.gpus:
-- Rule 1: Thermal critical
if gpu.temp_c > 85 and sustained_for(gpu, 120):
push actions apply-thermal-fix(gpu)
-- Rule 2: VRAM leak
let vram_rate = calculate-vram-trend(gpu, hours=6)
if vram_rate > 50:
push actions apply-vram-fix(gpu, vram_rate)
-- Rule 4: Benchmark regression
let bench = fleet.benchmarks[gpu.hostname]
if bench.current_tok_s < bench.baseline_tok_s * 0.8:
push actions apply-benchmark-fix(gpu, bench)
-- Rule 3: Model stuck
for model in fleet.summary.available_models:
if model.consecutive_timeouts >= 3:
push actions apply-model-restart(model)
-- Rule 5: Circuit breaker check via LiteLLM (router deprecated)
if fleet.summary.circuit_breakers_open > 0:
push actions check-litellm-circuit-breakers()
-- Rule 6: Strix Halo
if not fleet.strix.running and pingable("192.168.68.15"):
push actions apply-strix-restart()
-- Rule 7: GPU data source
if not fleet.gpus or len(fleet.gpus) < 2:
push actions check-gpu-monitor-service()
-- Rule 8: Predictive thermal
for gpu in fleet.gpus:
let rise_rate = calculate-temp-rise(gpu, minutes=5)
if rise_rate > 2.0 and gpu.temp_c < 80:
push actions apply-proactive-cooling(gpu)
-- Phase 3: Execute actions, verify, log
for action in actions:
let result = execute-with-verify(action)
log-to-kg(action, result)
if result.failed:
escalate-if-needed(action)
-- Phase 4: Update health state
call update-gpu-health
gpus: fleet.gpus
actions: actions
status: derive-overall-status(fleet, actions)
-- Wait 60s and repeat
Fan control: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
Model restart: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
Strix direct access: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
VRAM thresholds: Tiered — 300MB/h (RTX 3090), 300MB/h (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
CB auto-reset: ✅ Router decommissioned 2026-09-11 — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
Benchmark baseline: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
Predictive alerts: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
Prometheus: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
Lessons Learned (2026-07-12, Updated 2026-07-18)
L1: API Key Standardization Is Critical
All GPU llama-servers MUST use the same api-key as the LiteLLM config.
RTX 5070 had --api-key sk-loc...5678 while LiteLLM sent not-needed.
This caused cascading 401 → fallback → timeout → 401 loops.
Rule: Any new GPU or model restart MUST verify api-key matches LiteLLM config (not-needed for direct routing).
L2: Fallback Chain Cascading Failures
When one model returns 401 (auth) and another is slow (timeout), the fallback
chain creates an infinite loop.
Rule: If a model returns 401 (auth error), do NOT fall back to it again.
Mark it as permanently failed for this request.
L3: Verify Running State, Not Docs
RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18).
Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models).
Rule: Before making decisions, check /proc/PID/cmdline on GPU hosts.
L4: Infisical Is Not Always Available
Keep a local .env fallback for LITELLM_API_KEY.
Rule: Always verify credential source is reachable before relying on it.
L5: GPU Monitor Response Size Can Cause Self-Heal Crash
gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response.
Root cause: router poll returns accumulated data → cache balloons.
Rule: Self-heal must enforce a read timeout AND max response size on every poll.
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
L6: Stable Aliases Replace Model Names
gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15; gpu-light was superseded by gpu-vision on 2026-09-12.
Self-heal must use aliases for reporting and alerting, not model-specific names.
Rule: All alert messages and KG nodes use the stable alias as the GPU identifier.