- Move compression model from gemma-4-12b (RTX 5070) to ornith-1.0-35b (Strix Halo) - Add Rule 8: GPU Workload Distribution — per-GPU role assignment - Add Rule 9: Compression Threshold for 256K models - Update Rule 7: Auxiliary Model Consistency with new compression routing - Add gpu-self-heal.prose.md contract with 10 remediation rules - Strix Halo (64GB, 256K, 72.4 tok/s) → compression specialist - RTX 5070 (12GB) → vision/web search specialist - RTX 3090 (24GB, 256K) → heavy reasoning specialist - All rules grilled and confirmed with Kwame 2026-07-12
259 lines
10 KiB
Markdown
259 lines
10 KiB
Markdown
---
|
|
kind: responsibility
|
|
name: gpu-self-heal
|
|
description: >
|
|
GPU fleet self-healing — detects anomalies, applies remediation, tracks
|
|
benchmarks, and predicts failures before they happen. Extends gpu-monitor
|
|
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
|
|
VRAM trend analysis, and predictive alerting.
|
|
agent: abiba
|
|
depends_on:
|
|
- gpu-monitor.prose.md (live data source on .24:9100)
|
|
- gpu-fleet.prose.md (source of truth for topology)
|
|
---
|
|
|
|
## Maintains
|
|
|
|
- gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array }
|
|
- gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail
|
|
- benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts }
|
|
- vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours }
|
|
- circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset }
|
|
|
|
## Requires
|
|
|
|
- gpu-monitor:function — Live fleet data from .24:9100/gpu-data
|
|
- Prometheus exporters on all 3 GPUs (:9400/metrics)
|
|
- SSH access to GPU hosts for restart operations
|
|
|
|
## Continuity
|
|
|
|
- Self-driven: check every 60 seconds against GPU monitor data
|
|
- Also wakes on gpu-fleet health degradation
|
|
- On fix: verify with benchmark inference test before declaring resolved
|
|
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
|
|
|
|
---
|
|
|
|
## Remediation Rules
|
|
|
|
### Rule 1: GPU Temperature Critical (>85°C for >2 min)
|
|
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
|
|
- **Fix**:
|
|
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
|
|
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains
|
|
3. If all GPUs hot, alert about cooling infrastructure
|
|
- **Verify**: Temp drops below 80°C within 5 minutes
|
|
- **Escalate after**: 3 verification failures → Zulip alert
|
|
|
|
### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
|
|
- **Detect**: VRAM growing at sustained rate over 6+ hour window
|
|
- RTX 3090 (24GB): ≥100MB/hour
|
|
- RTX 5070 (12GB): ≥50MB/hour
|
|
- Strix Halo (64GB UMA): ≥200MB/hour
|
|
- **Fix**:
|
|
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
|
|
2. If llama-server is the growth source → restart with memory cap flag
|
|
3. If unknown process → kill and alert
|
|
- **Verify**: VRAM growth rate drops below threshold
|
|
- **Escalate after**: persistent leak after restart → hardware investigation
|
|
|
|
### Rule 3: Model Inference Timeout / GPU Stuck
|
|
- **Detect**: >50% failure rate over 60s window + 30s grace period (not just single stuck request)
|
|
- **Fix**:
|
|
1. Restart llama-server on affected GPU host
|
|
2. Wait 15s for model to reload
|
|
3. Run benchmark inference test
|
|
- **Verify**: Model returns 200 with <30s response, failure rate drops to 0%
|
|
- **Escalate after**: 3 restarts in 1 hour → GPU hardware check
|
|
|
|
### Rule 4: Benchmark Regression (>20% drop)
|
|
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
|
|
- **Fix**:
|
|
1. Check GPU utilization — if >90%, other process is competing
|
|
2. Check power limit — if throttled, restore to max
|
|
3. Check thermal — if hot, apply Rule 1
|
|
- **Verify**: Benchmark returns to within 10% of baseline
|
|
- **Escalate after**: persistent regression → possible hardware degradation
|
|
|
|
### Rule 5: Circuit Breaker Stuck Open
|
|
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
|
|
- **Fix**:
|
|
1. Verify GPU /health returns 200
|
|
2. If GPU healthy, send 1 test inference
|
|
3. If test succeeds → reset circuit breaker via router API
|
|
4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset
|
|
5. Max 1 auto-reset per GPU per hour
|
|
- **Verify**: CB closes, inference succeeds, CB stays closed for 60s+
|
|
- **Escalate after**: CB won't close after reset → router issue
|
|
|
|
### Rule 6: Strix Halo Unreachable
|
|
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
|
|
- **Fix**:
|
|
1. SSH to .15 → check llama-server process
|
|
2. Restart llama-server if not running
|
|
3. Verify through both direct probe AND router
|
|
- **Verify**: Direct health probe returns 200, router reports Strix healthy
|
|
- **Escalate**: If host .15 itself is unreachable → infrastructure alert
|
|
|
|
### Rule 7: Prometheus Exporter Down
|
|
- **Detect**: Any GPU :9400/metrics unreachable for >2 polls
|
|
- **Fix**:
|
|
1. SSH to GPU host → check prometheus-exporter process
|
|
2. Restart exporter if dead
|
|
3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes
|
|
4. If exporter is running but unreachable → check firewall/host networking
|
|
- **Verify**: :9400/metrics returns 200
|
|
- **Escalate after**: 3 failed restarts → networking issue
|
|
|
|
### Rule 8: Predictive Thermal Warning (two-tier)
|
|
- **Detect**:
|
|
- Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert
|
|
- Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding
|
|
- **Fix**:
|
|
- Tier 1: silently reduce parallel requests to that GPU by 50%
|
|
- Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic
|
|
- **Verify**: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2)
|
|
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
|
|
|
|
### Rule 9: Context Window Optimization
|
|
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context
|
|
- RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
|
|
- RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
|
|
- Strix Halo (256K ctx, ornith-1.0-35b): target 70+ tok/s — currently above baseline
|
|
- **Fix**:
|
|
- If tok/s > baseline → context has headroom, consider increasing
|
|
- If tok/s < 90% baseline → reduce context by 25% and retest
|
|
- If tok/s within 10% of baseline → optimal, no change
|
|
- **Verify**: Re-benchmark after context change, confirm within 10% of target
|
|
- **Escalate**: If context can't be adjusted without significant perf loss
|
|
|
|
### Rule 10: Workload Distribution Optimization
|
|
- **Detect**: GPU roles misaligned with hardware capabilities
|
|
- **Target distribution**:
|
|
- RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations
|
|
- RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
|
|
- Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs
|
|
- **Fix**:
|
|
- Alert if any GPU is handling workload outside its designated role
|
|
- Recommend Hermes agent profile updates to match workload to GPU
|
|
- Track per-GPU request distribution via LiteLLM spend logs
|
|
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
|
- **Escalate**: If role mismatch persists >48h → agent profile audit needed
|
|
|
|
---
|
|
|
|
## Execution
|
|
|
|
```prose
|
|
-- Phase 1: Fetch live GPU data
|
|
let fleet = call gpu-monitor
|
|
endpoint: "http://192.168.68.24:9100/gpu-data"
|
|
|
|
-- Phase 2: Evaluate each GPU against remediation rules
|
|
let actions = []
|
|
for gpu in fleet.gpus:
|
|
-- Rule 1: Thermal critical
|
|
if gpu.temp_c > 85 and sustained_for(gpu, 120):
|
|
push actions apply-thermal-fix(gpu)
|
|
|
|
-- Rule 2: VRAM leak
|
|
let vram_rate = calculate-vram-trend(gpu, hours=6)
|
|
if vram_rate > 50:
|
|
push actions apply-vram-fix(gpu, vram_rate)
|
|
|
|
-- Rule 4: Benchmark regression
|
|
let bench = fleet.benchmarks[gpu.hostname]
|
|
if bench.current_tok_sec < bench.baseline_tok_sec * 0.8:
|
|
push actions apply-benchmark-fix(gpu, bench)
|
|
|
|
-- Rule 3: Model stuck
|
|
for model in fleet.router.available_models:
|
|
if model.consecutive_timeouts >= 3:
|
|
push actions apply-model-restart(model)
|
|
|
|
-- Rule 5: Circuit breaker
|
|
for cb in fleet.router.circuit_breaker:
|
|
if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu):
|
|
push actions apply-cb-reset(cb)
|
|
|
|
-- Rule 6: Strix Halo
|
|
if not fleet.strix.running and pingable("192.168.68.15"):
|
|
push actions apply-strix-restart()
|
|
|
|
-- Rule 7: Prometheus exporters
|
|
for gpu in fleet.gpus:
|
|
if not prometheus_reachable(gpu.hostname, 9400):
|
|
push actions apply-exporter-restart(gpu)
|
|
|
|
-- Rule 8: Predictive thermal
|
|
for gpu in fleet.gpus:
|
|
let rise_rate = calculate-temp-rise(gpu, minutes=5)
|
|
if rise_rate > 2.0 and gpu.temp_c < 80:
|
|
push actions apply-proactive-cooling(gpu)
|
|
|
|
-- Phase 3: Execute actions, verify, log
|
|
for action in actions:
|
|
let result = execute-with-verify(action)
|
|
log-to-kg(action, result)
|
|
if result.failed:
|
|
escalate-if-needed(action)
|
|
|
|
-- Phase 4: Update health state
|
|
call update-gpu-health
|
|
gpus: fleet.gpus
|
|
actions: actions
|
|
status: derive-overall-status(fleet, actions)
|
|
|
|
-- Wait 60s and repeat
|
|
```
|
|
|
|
## Audit Trail Format
|
|
|
|
```json
|
|
{
|
|
"run_id": "gpu-self-heal-20260712-001",
|
|
"timestamp": "2026-07-12T16:00:00Z",
|
|
"gpu": "ct8-rtx3090",
|
|
"issue": "thermal-critical",
|
|
"detected": { "temp_c": 87, "duration_s": 180 },
|
|
"action": "set-fan-100pct",
|
|
"result": "resolved",
|
|
"verification": { "temp_c": 76, "after_s": 300 },
|
|
"escalated": false
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
## Reporting
|
|
|
|
### 1. Knowledge Graph
|
|
Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail.
|
|
|
|
### 2. Zulip Alerts (#agent-hub → alerts-gpu)
|
|
- `issues_fixed > 0` → "🛠 GPU Self-Heal — <gpu> <issue> resolved"
|
|
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
|
|
- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
|
|
|
|
### 3. Prometheus/Grafana Integration
|
|
- GPU self-heal actions exposed as Prometheus counter metrics
|
|
- Dashboard panel: "GPU Interventions (24h)" showing count/type/result
|
|
|
|
### 4. Weekly Benchmark Report
|
|
- Per-GPU tok/s trend over 7 days
|
|
- Regression alerts if any GPU degrades >10% week-over-week
|
|
|
|
---
|
|
|
|
## Design Decisions (Grilled & Confirmed — 2026-07-12)
|
|
|
|
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
|
|
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
|
|
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
|
|
4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix).
|
|
5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU.
|
|
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana.
|
|
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
|
|
8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.
|