--- kind: responsibility name: gpu-self-heal description: > GPU fleet self-healing — detects anomalies, applies remediation, tracks benchmarks, and predicts failures before they happen. Extends gpu-monitor (v2.1.0) with active remediation rules, Prometheus metrics consumption, VRAM trend analysis, and predictive alerting. agent: abiba depends_on: - gpu-monitor.prose.md (live data source on .24:9100) - gpu-fleet.prose.md (source of truth for topology) --- ## Maintains - gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array } - gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail - benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts } - vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours } - circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset } ## Requires - gpu-monitor:function — Live fleet data from .24:9100/gpu-data - Prometheus exporters on all 3 GPUs (:9400/metrics) - SSH access to GPU hosts for restart operations ## Continuity - Self-driven: check every 60 seconds against GPU monitor data - Also wakes on gpu-fleet health degradation - On fix: verify with benchmark inference test before declaring resolved - Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert --- ## Remediation Rules ### Rule 1: GPU Temperature Critical (>85°C for >2 min) - **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls - **Fix**: 1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control) 2. Redirect new requests to cooler GPUs via LiteLLM fallback chains 3. If all GPUs hot, alert about cooling infrastructure - **Verify**: Temp drops below 80°C within 5 minutes - **Escalate after**: 3 verification failures → Zulip alert ### Rule 2: VRAM Leak Detection (tiered by GPU capacity) - **Detect**: VRAM growing at sustained rate over 6+ hour window - RTX 3090 (24GB): ≥100MB/hour - RTX 5070 (12GB): ≥50MB/hour - Strix Halo (64GB UMA): ≥200MB/hour - **Fix**: 1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux) 2. If llama-server is the growth source → restart with memory cap flag 3. If unknown process → kill and alert - **Verify**: VRAM growth rate drops below threshold - **Escalate after**: persistent leak after restart → hardware investigation ### Rule 3: Model Inference Timeout / GPU Stuck - **Detect**: >50% failure rate over 60s window + 30s grace period (not just single stuck request) - **Fix**: 1. Restart llama-server on affected GPU host 2. Wait 15s for model to reload 3. Run benchmark inference test - **Verify**: Model returns 200 with <30s response, failure rate drops to 0% - **Escalate after**: 3 restarts in 1 hour → GPU hardware check ### Rule 4: Benchmark Regression (>20% drop) - **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks - **Fix**: 1. Check GPU utilization — if >90%, other process is competing 2. Check power limit — if throttled, restore to max 3. Check thermal — if hot, apply Rule 1 - **Verify**: Benchmark returns to within 10% of baseline - **Escalate after**: persistent regression → possible hardware degradation ### Rule 5: Circuit Breaker Stuck Open - **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy - **Fix**: 1. Verify GPU /health returns 200 2. If GPU healthy, send 1 test inference 3. If test succeeds → reset circuit breaker via router API 4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset 5. Max 1 auto-reset per GPU per hour - **Verify**: CB closes, inference succeeds, CB stays closed for 60s+ - **Escalate after**: CB won't close after reset → router issue ### Rule 6: Strix Halo Unreachable - **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15) - **Fix**: 1. SSH to .15 → check llama-server process 2. Restart llama-server if not running 3. Verify through both direct probe AND router - **Verify**: Direct health probe returns 200, router reports Strix healthy - **Escalate**: If host .15 itself is unreachable → infrastructure alert ### Rule 7: Prometheus Exporter Down - **Detect**: Any GPU :9400/metrics unreachable for >2 polls - **Fix**: 1. SSH to GPU host → check prometheus-exporter process 2. Restart exporter if dead 3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes 4. If exporter is running but unreachable → check firewall/host networking - **Verify**: :9400/metrics returns 200 - **Escalate after**: 3 failed restarts → networking issue ### Rule 8: Predictive Thermal Warning (two-tier) - **Detect**: - Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert - Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding - **Fix**: - Tier 1: silently reduce parallel requests to that GPU by 50% - Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic - **Verify**: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2) - **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure ### Rule 9: Context Window Optimization - **Detect**: Benchmark tok/s vs baseline for each GPU at current context - RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline - RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role - Strix Halo (256K ctx, ornith-1.0-35b): target 70+ tok/s — currently above baseline - **Fix**: - If tok/s > baseline → context has headroom, consider increasing - If tok/s < 90% baseline → reduce context by 25% and retest - If tok/s within 10% of baseline → optimal, no change - **Verify**: Re-benchmark after context change, confirm within 10% of target - **Escalate**: If context can't be adjusted without significant perf loss ### Rule 10: Workload Distribution Optimization - **Detect**: GPU roles misaligned with hardware capabilities - **Target distribution**: - RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations - RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks - Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs - **Fix**: - Alert if any GPU is handling workload outside its designated role - Recommend Hermes agent profile updates to match workload to GPU - Track per-GPU request distribution via LiteLLM spend logs - **Verify**: Each GPU's request pattern matches its designated role within 24h - **Escalate**: If role mismatch persists >48h → agent profile audit needed --- ## Execution ```prose -- Phase 1: Fetch live GPU data let fleet = call gpu-monitor endpoint: "http://192.168.68.24:9100/gpu-data" -- Phase 2: Evaluate each GPU against remediation rules let actions = [] for gpu in fleet.gpus: -- Rule 1: Thermal critical if gpu.temp_c > 85 and sustained_for(gpu, 120): push actions apply-thermal-fix(gpu) -- Rule 2: VRAM leak let vram_rate = calculate-vram-trend(gpu, hours=6) if vram_rate > 50: push actions apply-vram-fix(gpu, vram_rate) -- Rule 4: Benchmark regression let bench = fleet.benchmarks[gpu.hostname] if bench.current_tok_sec < bench.baseline_tok_sec * 0.8: push actions apply-benchmark-fix(gpu, bench) -- Rule 3: Model stuck for model in fleet.router.available_models: if model.consecutive_timeouts >= 3: push actions apply-model-restart(model) -- Rule 5: Circuit breaker for cb in fleet.router.circuit_breaker: if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu): push actions apply-cb-reset(cb) -- Rule 6: Strix Halo if not fleet.strix.running and pingable("192.168.68.15"): push actions apply-strix-restart() -- Rule 7: Prometheus exporters for gpu in fleet.gpus: if not prometheus_reachable(gpu.hostname, 9400): push actions apply-exporter-restart(gpu) -- Rule 8: Predictive thermal for gpu in fleet.gpus: let rise_rate = calculate-temp-rise(gpu, minutes=5) if rise_rate > 2.0 and gpu.temp_c < 80: push actions apply-proactive-cooling(gpu) -- Phase 3: Execute actions, verify, log for action in actions: let result = execute-with-verify(action) log-to-kg(action, result) if result.failed: escalate-if-needed(action) -- Phase 4: Update health state call update-gpu-health gpus: fleet.gpus actions: actions status: derive-overall-status(fleet, actions) -- Wait 60s and repeat ``` ## Audit Trail Format ```json { "run_id": "gpu-self-heal-20260712-001", "timestamp": "2026-07-12T16:00:00Z", "gpu": "ct8-rtx3090", "issue": "thermal-critical", "detected": { "temp_c": 87, "duration_s": 180 }, "action": "set-fan-100pct", "result": "resolved", "verification": { "temp_c": 76, "after_s": 300 }, "escalated": false } ``` --- ## Reporting ### 1. Knowledge Graph Every action logged as `[GPU-SELF-HEAL] ` node with full audit trail. ### 2. Zulip Alerts (#agent-hub → alerts-gpu) - `issues_fixed > 0` → "🛠 GPU Self-Heal — resolved" - `issues_escalated > 0` → "⚠ GPU Self-Heal — needs attention" - Every 100th clean cycle → "✅ GPU Fleet: All Clear" ### 3. Prometheus/Grafana Integration - GPU self-heal actions exposed as Prometheus counter metrics - Dashboard panel: "GPU Interventions (24h)" showing count/type/result ### 4. Weekly Benchmark Report - Per-GPU tok/s trend over 7 days - Regression alerts if any GPU degrades >10% week-over-week --- ## Design Decisions (Grilled & Confirmed — 2026-07-12) 1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect). 2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request. 3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart. 4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix). 5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU. 6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana. 7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising. 8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.