--- kind: responsibility name: gpu-self-heal description: > GPU fleet self-healing — detects anomalies, applies remediation, tracks benchmarks, and predicts failures before they happen. Extends gpu-monitor (v2.1.0) with active remediation rules, Prometheus metrics consumption, VRAM trend analysis, and predictive alerting. UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps. Router (port 9000) references replaced with direct GPU routing. Benchmark baselines refreshed to live values. Prometheus exporters removed — not deployed; fall back to direct sidecar probes. Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet. agent: abiba depends_on: - gpu-monitor.prose.md (live data source on .24:9100) - gpu-fleet.prose.md (source of truth for topology, aliases, model assignments) --- ## Maintains - gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array } - gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail - benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts } - vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours } - circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset } ## Requires - gpu-monitor:function — Live fleet data from localhost:9100/gpu-data - Direct sidecar probe access to all GPU hosts (:8080/health) - SSH access to GPU hosts for restart operations ## Continuity - Self-driven: check every 60 seconds against GPU monitor data - Also wakes on gpu-fleet health degradation - On fix: verify with benchmark inference test before declaring resolved - Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert --- ## Current Fleet Baseline (2026-07-18) | Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role | |-------|-----|------|-------|------|-----|-------|------| | `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen | | `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks | | `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | qwen3.6-35B-udq4 | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs | Key notes: - All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path. - Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. - RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first. - Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads. - RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom). - RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes). ## Remediation Rules ### Rule 1: GPU Temperature Critical (>85°C for >2 min) - **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls - **Fix**: 1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control) 2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma) 3. If all GPUs hot, alert about cooling infrastructure - **Verify**: Temp drops below 80°C within 5 minutes - **Escalate after**: 3 verification failures → Zulip alert ### Rule 2: VRAM Leak Detection (tiered by GPU capacity) - **Detect**: VRAM growing at sustained rate over 6+ hour window - RTX 3090 (24GB): ≥300MB/hour - RTX 5070 (12GB): ≥300MB/hour - Strix Halo (64GB UMA): ≥200MB/hour - **Fix**: 1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux) 2. If llama-server is the growth source → restart with memory cap flag 3. If unknown process → kill and alert - **Verify**: VRAM growth rate drops below threshold - **Escalate after**: persistent leak after restart → hardware investigation ### Rule 3: Model Inference Timeout / GPU Stuck - **Detect**: >50% failure rate over 60s window + 30s grace period (not just single stuck request) - **Fix**: 1. Restart llama-server on affected GPU host 2. Wait 15s for model to reload 3. Run benchmark inference test - **Verify**: Model returns 200 with <30s response, failure rate drops to 0% - **Escalate after**: 3 restarts in 1 hour → GPU hardware check ### Rule 4: Benchmark Regression (>20% drop) - **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks - RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s - RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s - Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s - **Fix**: 1. Check GPU utilization — if >90%, other process is competing 2. Check power limit — if throttled, restore to max 3. Check thermal — if hot, apply Rule 1 - **Verify**: Benchmark returns to within 10% of baseline - **Escalate after**: persistent regression → possible hardware degradation ### Rule 5: Circuit Breaker Stuck Open - **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy - **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router. - **Fix**: 1. Verify GPU /health returns 200 on direct port (:8080) 2. If GPU healthy, alert but do NOT reset via router API (deprecated) 3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness 4. Restart LiteLLM container on CT 116 if circuit breakers are stuck - **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s - **Escalate after**: LiteLLM restart doesn't clear → human investigation ### Rule 6: Strix Halo Unreachable - **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15) - **Fix**: 1. SSH to .15 → check llama-server process 2. Restart llama-server if not running 3. Verify through both direct probe AND LiteLLM health - **Verify**: Direct health probe returns 200, LiteLLM reports model healthy - **Escalate**: If host .15 itself is unreachable → infrastructure alert ### Rule 7: GPU Data Source Unreachable (replaces old Prometheus rule) - **Detect**: gpu-monitor endpoint (localhost:9100/gpu-data) or sidecar port (:8080) on any GPU unreachable for >2 polls - **Fix**: 1. If gpu-monitor is down: restart systemd service `gpu-monitor.service` on this host 2. If sidecar is down: SSH to GPU host → check llama-server process → restart systemd service 3. Fall back to direct nvidia-smi/rocm-smi probe via SSH if all API paths fail - **Verify**: gpu-monitor returns healthy + all sidecars reachable - **Escalate after**: 3 failed restarts → networking issue ### Rule 8: Predictive Thermal Warning (two-tier) - **Detect**: - Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert - Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding - **Fix**: - Tier 1: silently reduce parallel requests to that GPU by 50% - Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic - **Verify**: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2) - **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure ### Rule 9: Context Window Optimization - **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K) - RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%) - RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%) - Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%) - **Fix**: - If tok/s > baseline → context has headroom, consider increasing - If tok/s < 90% baseline → reduce context by 25% and retest - If tok/s within 10% of baseline → optimal, no change - Strix Halo at 89% of baseline → MONITOR but do not reduce yet (recent model swap may still be settling) - **Verify**: Re-benchmark after context change, confirm within 10% of target - **Escalate**: If context can't be adjusted without significant perf loss ### Rule 10: Workload Distribution Optimization (updated 2026-07-18) - **Detect**: GPU roles misaligned with hardware capabilities - **Target distribution**: - RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM). - RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM). - Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM). - **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first. - **Fix**: - Alert if any GPU is handling workload outside its designated role - Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe) - Track per-GPU request distribution via LiteLLM spend logs - **Verify**: Each GPU's request pattern matches its designated role within 24h - **Escalate**: If role mismatch persists >48h → agent alias audit needed --- ## Execution ```prose -- Phase 1: Fetch live GPU data let fleet = call gpu-monitor endpoint: "http://localhost:9100/gpu-data" -- Phase 2: Evaluate each GPU against remediation rules let actions = [] for gpu in fleet.gpus: -- Rule 1: Thermal critical if gpu.temp_c > 85 and sustained_for(gpu, 120): push actions apply-thermal-fix(gpu) -- Rule 2: VRAM leak let vram_rate = calculate-vram-trend(gpu, hours=6) if vram_rate > 50: push actions apply-vram-fix(gpu, vram_rate) -- Rule 4: Benchmark regression let bench = fleet.benchmarks[gpu.hostname] if bench.current_tok_s < bench.baseline_tok_s * 0.8: push actions apply-benchmark-fix(gpu, bench) -- Rule 3: Model stuck for model in fleet.summary.available_models: if model.consecutive_timeouts >= 3: push actions apply-model-restart(model) -- Rule 5: Circuit breaker check via LiteLLM (router deprecated) if fleet.summary.circuit_breakers_open > 0: push actions check-litellm-circuit-breakers() -- Rule 6: Strix Halo if not fleet.strix.running and pingable("192.168.68.15"): push actions apply-strix-restart() -- Rule 7: GPU data source if not fleet.gpus or len(fleet.gpus) < 2: push actions check-gpu-monitor-service() -- Rule 8: Predictive thermal for gpu in fleet.gpus: let rise_rate = calculate-temp-rise(gpu, minutes=5) if rise_rate > 2.0 and gpu.temp_c < 80: push actions apply-proactive-cooling(gpu) -- Phase 3: Execute actions, verify, log for action in actions: let result = execute-with-verify(action) log-to-kg(action, result) if result.failed: escalate-if-needed(action) -- Phase 4: Update health state call update-gpu-health gpus: fleet.gpus actions: actions status: derive-overall-status(fleet, actions) -- Wait 60s and repeat ``` ## Audit Trail Format ```json { "run_id": "gpu-self-heal-20260718-001", "timestamp": "2026-07-18T08:00:00Z", "gpu": "ct8-rtx3090", "issue": "thermal-critical", "detected": { "temp_c": 87, "duration_s": 180 }, "action": "load-shedding", "result": "resolved", "verification": { "temp_c": 76, "after_s": 300 }, "escalated": false } ``` --- ## Reporting ### 1. Knowledge Graph Every action logged as `[GPU-SELF-HEAL] ` node with full audit trail. ### 2. Zulip Alerts (#agent-hub → alerts-gpu) - `issues_fixed > 0` → "🛠 GPU Self-Heal — resolved" - `issues_escalated > 0` → "⚠ GPU Self-Heal — needs attention" - Every 100th clean cycle → "✅ GPU Fleet: All Clear" ### 3. Weekly Benchmark Report - Per-GPU tok/s trend over 7 days - Regression alerts if any GPU degrades >10% week-over-week --- ## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18) 1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect). 2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request. 3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart. 4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data. 5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset. 6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor. 7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising. 8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts. ## Lessons Learned (2026-07-12, Updated 2026-07-18) ### L1: API Key Standardization Is Critical - All GPU llama-servers MUST use the same api-key as the LiteLLM config. - RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`. This caused cascading 401 → fallback → timeout → 401 loops. - **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config (`not-needed` for direct routing). ### L2: Fallback Chain Cascading Failures - When one model returns 401 (auth) and another is slow (timeout), the fallback chain creates an infinite loop. - **Rule**: If a model returns 401 (auth error), do NOT fall back to it again. Mark it as permanently failed for this request. ### L3: Verify Running State, Not Docs - RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18). - Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models). - **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts. ### L4: Infisical Is Not Always Available - Keep a local `.env` fallback for `LITELLM_API_KEY`. - **Rule**: Always verify credential source is reachable before relying on it. ### L5: GPU Monitor Response Size Can Cause Self-Heal Crash - gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response. - Root cause: router poll returns accumulated data → cache balloons. - **Rule**: Self-heal must enforce a read timeout AND max response size on every poll. If monitor response > 1MB, log a warning and skip the cycle rather than crashing. ### L6: Stable Aliases Replace Model Names - gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15. - Self-heal must use aliases for reporting and alerting, not model-specific names. - **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.