Verified on ground 2026-07-16 against CT 116 litellm_config.yaml + GPU hosts: - AMD host serves qwen3.6-35B-udq4 (LiteLLM alias strix-moe); ornith-1.0-35b does NOT exist - All 3 GPUs at 256K ctx, parallel 2 (RTX 3090 was listed 128K/parallel 1) - LiteLLM timeouts: qwen 300s, gemma 120s, strix 300s (were stale 90s/120s) - Added LiteLLM model surface + key scoping to litellm-self-heal - Patched health-check script path ref Files: litellm-self-heal, litellm-health, gpu-fleet, gpu-self-heal, zulip-adapter-lessons, abiba-zulip-restore, hermes-agent-baseline, delegation-prose-contract, mumuni-delegation-prose-contract
12 KiB
12 KiB
kind, name, description, agent, depends_on
| kind | name | description | agent | depends_on | ||
|---|---|---|---|---|---|---|
| responsibility | gpu-self-heal | GPU fleet self-healing — detects anomalies, applies remediation, tracks benchmarks, and predicts failures before they happen. Extends gpu-monitor (v2.1.0) with active remediation rules, Prometheus metrics consumption, VRAM trend analysis, and predictive alerting. | abiba |
|
Maintains
- gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array }
- gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail
- benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts }
- vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours }
- circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset }
Requires
- gpu-monitor:function — Live fleet data from .24:9100/gpu-data
- Prometheus exporters on all 3 GPUs (:9400/metrics)
- SSH access to GPU hosts for restart operations
Continuity
- Self-driven: check every 60 seconds against GPU monitor data
- Also wakes on gpu-fleet health degradation
- On fix: verify with benchmark inference test before declaring resolved
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
Remediation Rules
Rule 1: GPU Temperature Critical (>85°C for >2 min)
- Detect: Any GPU temp >85°C sustained for 2+ consecutive polls
- Fix:
- Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
- Redirect new requests to cooler GPUs via LiteLLM fallback chains
- If all GPUs hot, alert about cooling infrastructure
- Verify: Temp drops below 80°C within 5 minutes
- Escalate after: 3 verification failures → Zulip alert
Rule 2: VRAM Leak Detection (tiered by GPU capacity)
- Detect: VRAM growing at sustained rate over 6+ hour window
- RTX 3090 (24GB): ≥100MB/hour
- RTX 5070 (12GB): ≥50MB/hour
- Strix Halo (64GB UMA): ≥200MB/hour
- Fix:
- Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
- If llama-server is the growth source → restart with memory cap flag
- If unknown process → kill and alert
- Verify: VRAM growth rate drops below threshold
- Escalate after: persistent leak after restart → hardware investigation
Rule 3: Model Inference Timeout / GPU Stuck
- Detect: >50% failure rate over 60s window + 30s grace period (not just single stuck request)
- Fix:
- Restart llama-server on affected GPU host
- Wait 15s for model to reload
- Run benchmark inference test
- Verify: Model returns 200 with <30s response, failure rate drops to 0%
- Escalate after: 3 restarts in 1 hour → GPU hardware check
Rule 4: Benchmark Regression (>20% drop)
- Detect: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
- Fix:
- Check GPU utilization — if >90%, other process is competing
- Check power limit — if throttled, restore to max
- Check thermal — if hot, apply Rule 1
- Verify: Benchmark returns to within 10% of baseline
- Escalate after: persistent regression → possible hardware degradation
Rule 5: Circuit Breaker Stuck Open
- Detect: Circuit breaker open >10 minutes with GPU reporting healthy
- Fix:
- Verify GPU /health returns 200
- If GPU healthy, send 1 test inference
- If test succeeds → reset circuit breaker via router API
- 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset
- Max 1 auto-reset per GPU per hour
- Verify: CB closes, inference succeeds, CB stays closed for 60s+
- Escalate after: CB won't close after reset → router issue
Rule 6: Strix Halo Unreachable
- Detect: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
- Fix:
- SSH to .15 → check llama-server process
- Restart llama-server if not running
- Verify through both direct probe AND router
- Verify: Direct health probe returns 200, router reports Strix healthy
- Escalate: If host .15 itself is unreachable → infrastructure alert
Rule 7: Prometheus Exporter Down
- Detect: Any GPU :9400/metrics unreachable for >2 polls
- Fix:
- SSH to GPU host → check prometheus-exporter process
- Restart exporter if dead
- While exporter is down, fall back to nvidia-smi/rocm-smi direct probes
- If exporter is running but unreachable → check firewall/host networking
- Verify: :9400/metrics returns 200
- Escalate after: 3 failed restarts → networking issue
Rule 8: Predictive Thermal Warning (two-tier)
- Detect:
- Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert
- Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding
- Fix:
- Tier 1: silently reduce parallel requests to that GPU by 50%
- Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic
- Verify: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2)
- Escalate: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
Rule 9: Context Window Optimization
- Detect: Benchmark tok/s vs baseline for each GPU at current context
- RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
- RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
- Strix Halo (256K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
- Fix:
- If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest
- If tok/s within 10% of baseline → optimal, no change
- Verify: Re-benchmark after context change, confirm within 10% of target
- Escalate: If context can't be adjusted without significant perf loss
Rule 10: Workload Distribution Optimization
- Detect: GPU roles misaligned with hardware capabilities
- Target distribution:
- RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations
- RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
- Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs
- Fix:
- Alert if any GPU is handling workload outside its designated role
- Recommend Hermes agent profile updates to match workload to GPU
- Track per-GPU request distribution via LiteLLM spend logs
- Verify: Each GPU's request pattern matches its designated role within 24h
- Escalate: If role mismatch persists >48h → agent profile audit needed
Execution
-- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor
endpoint: "http://192.168.68.24:9100/gpu-data"
-- Phase 2: Evaluate each GPU against remediation rules
let actions = []
for gpu in fleet.gpus:
-- Rule 1: Thermal critical
if gpu.temp_c > 85 and sustained_for(gpu, 120):
push actions apply-thermal-fix(gpu)
-- Rule 2: VRAM leak
let vram_rate = calculate-vram-trend(gpu, hours=6)
if vram_rate > 50:
push actions apply-vram-fix(gpu, vram_rate)
-- Rule 4: Benchmark regression
let bench = fleet.benchmarks[gpu.hostname]
if bench.current_tok_sec < bench.baseline_tok_sec * 0.8:
push actions apply-benchmark-fix(gpu, bench)
-- Rule 3: Model stuck
for model in fleet.router.available_models:
if model.consecutive_timeouts >= 3:
push actions apply-model-restart(model)
-- Rule 5: Circuit breaker
for cb in fleet.router.circuit_breaker:
if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu):
push actions apply-cb-reset(cb)
-- Rule 6: Strix Halo
if not fleet.strix.running and pingable("192.168.68.15"):
push actions apply-strix-restart()
-- Rule 7: Prometheus exporters
for gpu in fleet.gpus:
if not prometheus_reachable(gpu.hostname, 9400):
push actions apply-exporter-restart(gpu)
-- Rule 8: Predictive thermal
for gpu in fleet.gpus:
let rise_rate = calculate-temp-rise(gpu, minutes=5)
if rise_rate > 2.0 and gpu.temp_c < 80:
push actions apply-proactive-cooling(gpu)
-- Phase 3: Execute actions, verify, log
for action in actions:
let result = execute-with-verify(action)
log-to-kg(action, result)
if result.failed:
escalate-if-needed(action)
-- Phase 4: Update health state
call update-gpu-health
gpus: fleet.gpus
actions: actions
status: derive-overall-status(fleet, actions)
-- Wait 60s and repeat
Audit Trail Format
{
"run_id": "gpu-self-heal-20260712-001",
"timestamp": "2026-07-12T16:00:00Z",
"gpu": "ct8-rtx3090",
"issue": "thermal-critical",
"detected": { "temp_c": 87, "duration_s": 180 },
"action": "set-fan-100pct",
"result": "resolved",
"verification": { "temp_c": 76, "after_s": 300 },
"escalated": false
}
Reporting
1. Knowledge Graph
Every action logged as [GPU-SELF-HEAL] <run_id> node with full audit trail.
2. Zulip Alerts (#agent-hub → alerts-gpu)
issues_fixed > 0→ "🛠 GPU Self-Heal — resolved"issues_escalated > 0→ "⚠ GPU Self-Heal — needs attention"- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
3. Prometheus/Grafana Integration
- GPU self-heal actions exposed as Prometheus counter metrics
- Dashboard panel: "GPU Interventions (24h)" showing count/type/result
4. Weekly Benchmark Report
- Per-GPU tok/s trend over 7 days
- Regression alerts if any GPU degrades >10% week-over-week
Design Decisions (Grilled & Confirmed — 2026-07-12)
- Fan control: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
- Model restart: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
- Strix direct access: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
- VRAM thresholds: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix).
- CB auto-reset: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU.
- Benchmark baseline: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana.
- Predictive alerts: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
- Prometheus: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.
Lessons Learned (2026-07-12)
L1: API Key Standardization Is Critical
- All GPU llama-servers MUST use the same api-key as the LiteLLM config.
- RTX 5070 had
--api-key sk-loc...5678while LiteLLM sentnot-needed. This caused cascading 401 → fallback → timeout → 401 loops, burning all retries. - Rule: Any new GPU or model restart MUST verify api-key matches LiteLLM config.
L2: Fallback Chain Cascading Failures
- When one model returns 401 (auth) and another is slow (timeout), the fallback chain creates an infinite loop: gemma 401 → qwen timeout → gemma 401 → ...
- Rule: If a model returns 401 (auth error), do NOT fall back to it again. Mark it as permanently failed for this request.
L3: Verify Running State, Not Docs
- RTX 3090 was documented at 128K context. Actually running at 256K.
- Parallel count wrong (docs said 2, actual is 1 on RTX 3090).
- Rule: Before making decisions, check
/proc/PID/cmdlineon GPU hosts.
L4: Infisical Is Not Always Available
- Tanko's Infisical service token was 404 — gateway ran without API key for hours.
- Rule: Always keep a local
.envfallback forLITELLM_API_KEY. - Contract hermes-config-template Rule 3 updated.
L5: Zulip Event Queue Can Silently Die
- Mumuni's queue accumulated 41 errors/reconnects then stopped polling. Gateway was running but ignoring all messages.
- Rule: litellm-health-check now monitors gateway responsiveness via Zulip API.