PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Synced model assignments to 2026-07-17 swaps (ThinkingCap, HauhauCS QAT, Genesis Hermes V3) - Added stable role-based aliases from gpu-fleet (gpu-dense, gpu-light, strix-moe) - Updated benchmark baselines to live values (74.9/169.6/62.9 tok/s) - Replaced router (port 9000) references with LiteLLM + direct routing - Replaced Prometheus exporter rule with sidecar health probe - Updated VRAM thresholds to match operational data (300/300/200 MB/h) - Added response size limit (1MB) to prevent OOM crashes - Added Lessons L5 (response size crash) and L6 (stable aliases) - Removed deprecated Rules 11-12 (router-specific distribution balance)
15 KiB
15 KiB
kind, name, description, agent, depends_on
| kind | name | description | agent | depends_on | ||
|---|---|---|---|---|---|---|
| responsibility | gpu-self-heal | GPU fleet self-healing — detects anomalies, applies remediation, tracks benchmarks, and predicts failures before they happen. Extends gpu-monitor (v2.1.0) with active remediation rules, Prometheus metrics consumption, VRAM trend analysis, and predictive alerting. UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps. Router (port 9000) references replaced with direct GPU routing. Benchmark baselines refreshed to live values. Prometheus exporters removed — not deployed; fall back to direct sidecar probes. Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet. | abiba |
|
Maintains
- gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array }
- gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail
- benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts }
- vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours }
- circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset }
Requires
- gpu-monitor:function — Live fleet data from localhost:9100/gpu-data
- Direct sidecar probe access to all GPU hosts (:8080/health)
- SSH access to GPU hosts for restart operations
Continuity
- Self-driven: check every 60 seconds against GPU monitor data
- Also wakes on gpu-fleet health degradation
- On fix: verify with benchmark inference test before declaring resolved
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
Current Fleet Baseline (2026-07-18)
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|---|---|---|---|---|---|---|---|
gpu-dense |
RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
gpu-light |
RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
strix-moe |
Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (
api_key: not-needed). Router (port 9000) is deprecated and NOT in the inference path. - Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
Remediation Rules
Rule 1: GPU Temperature Critical (>85°C for >2 min)
- Detect: Any GPU temp >85°C sustained for 2+ consecutive polls
- Fix:
- Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
- Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
- If all GPUs hot, alert about cooling infrastructure
- Verify: Temp drops below 80°C within 5 minutes
- Escalate after: 3 verification failures → Zulip alert
Rule 2: VRAM Leak Detection (tiered by GPU capacity)
- Detect: VRAM growing at sustained rate over 6+ hour window
- RTX 3090 (24GB): ≥300MB/hour
- RTX 5070 (12GB): ≥300MB/hour
- Strix Halo (64GB UMA): ≥200MB/hour
- Fix:
- Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
- If llama-server is the growth source → restart with memory cap flag
- If unknown process → kill and alert
- Verify: VRAM growth rate drops below threshold
- Escalate after: persistent leak after restart → hardware investigation
Rule 3: Model Inference Timeout / GPU Stuck
- Detect: >50% failure rate over 60s window + 30s grace period (not just single stuck request)
- Fix:
- Restart llama-server on affected GPU host
- Wait 15s for model to reload
- Run benchmark inference test
- Verify: Model returns 200 with <30s response, failure rate drops to 0%
- Escalate after: 3 restarts in 1 hour → GPU hardware check
Rule 4: Benchmark Regression (>20% drop)
- Detect: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
- RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s
- RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s
- Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s
- Fix:
- Check GPU utilization — if >90%, other process is competing
- Check power limit — if throttled, restore to max
- Check thermal — if hot, apply Rule 1
- Verify: Benchmark returns to within 10% of baseline
- Escalate after: persistent regression → possible hardware degradation
Rule 5: Circuit Breaker Stuck Open
- Detect: Circuit breaker open >10 minutes with GPU reporting healthy
- Note: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
- Fix:
- Verify GPU /health returns 200 on direct port (:8080)
- If GPU healthy, alert but do NOT reset via router API (deprecated)
- Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
- Restart LiteLLM container on CT 116 if circuit breakers are stuck
- Verify: LiteLLM returns healthy, circuit breaker clears within 60s
- Escalate after: LiteLLM restart doesn't clear → human investigation
Rule 6: Strix Halo Unreachable
- Detect: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
- Fix:
- SSH to .15 → check llama-server process
- Restart llama-server if not running
- Verify through both direct probe AND LiteLLM health
- Verify: Direct health probe returns 200, LiteLLM reports model healthy
- Escalate: If host .15 itself is unreachable → infrastructure alert
Rule 7: GPU Data Source Unreachable (replaces old Prometheus rule)
- Detect: gpu-monitor endpoint (localhost:9100/gpu-data) or sidecar port (:8080) on any GPU unreachable for >2 polls
- Fix:
- If gpu-monitor is down: restart systemd service
gpu-monitor.serviceon this host - If sidecar is down: SSH to GPU host → check llama-server process → restart systemd service
- Fall back to direct nvidia-smi/rocm-smi probe via SSH if all API paths fail
- If gpu-monitor is down: restart systemd service
- Verify: gpu-monitor returns healthy + all sidecars reachable
- Escalate after: 3 failed restarts → networking issue
Rule 8: Predictive Thermal Warning (two-tier)
- Detect:
- Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert
- Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding
- Fix:
- Tier 1: silently reduce parallel requests to that GPU by 50%
- Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic
- Verify: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2)
- Escalate: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
Rule 9: Context Window Optimization
- Detect: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
- Strix Halo (128K ctx, Genesis Hermes V3): baseline 70.5 tok/s — currently at 62.9 (89%)
- Fix:
- If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest
- If tok/s within 10% of baseline → optimal, no change
- Strix Halo at 89% of baseline → MONITOR but do not reduce yet (recent model swap may still be settling)
- Verify: Re-benchmark after context change, confirm within 10% of target
- Escalate: If context can't be adjusted without significant perf loss
Rule 10: Workload Distribution Optimization (updated 2026-07-18)
- Detect: GPU roles misaligned with hardware capabilities
- Target distribution:
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
- Note: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
- Fix:
- Alert if any GPU is handling workload outside its designated role
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe)
- Track per-GPU request distribution via LiteLLM spend logs
- Verify: Each GPU's request pattern matches its designated role within 24h
- Escalate: If role mismatch persists >48h → agent alias audit needed
Execution
-- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor
endpoint: "http://localhost:9100/gpu-data"
-- Phase 2: Evaluate each GPU against remediation rules
let actions = []
for gpu in fleet.gpus:
-- Rule 1: Thermal critical
if gpu.temp_c > 85 and sustained_for(gpu, 120):
push actions apply-thermal-fix(gpu)
-- Rule 2: VRAM leak
let vram_rate = calculate-vram-trend(gpu, hours=6)
if vram_rate > 50:
push actions apply-vram-fix(gpu, vram_rate)
-- Rule 4: Benchmark regression
let bench = fleet.benchmarks[gpu.hostname]
if bench.current_tok_s < bench.baseline_tok_s * 0.8:
push actions apply-benchmark-fix(gpu, bench)
-- Rule 3: Model stuck
for model in fleet.summary.available_models:
if model.consecutive_timeouts >= 3:
push actions apply-model-restart(model)
-- Rule 5: Circuit breaker check via LiteLLM (router deprecated)
if fleet.summary.circuit_breakers_open > 0:
push actions check-litellm-circuit-breakers()
-- Rule 6: Strix Halo
if not fleet.strix.running and pingable("192.168.68.15"):
push actions apply-strix-restart()
-- Rule 7: GPU data source
if not fleet.gpus or len(fleet.gpus) < 2:
push actions check-gpu-monitor-service()
-- Rule 8: Predictive thermal
for gpu in fleet.gpus:
let rise_rate = calculate-temp-rise(gpu, minutes=5)
if rise_rate > 2.0 and gpu.temp_c < 80:
push actions apply-proactive-cooling(gpu)
-- Phase 3: Execute actions, verify, log
for action in actions:
let result = execute-with-verify(action)
log-to-kg(action, result)
if result.failed:
escalate-if-needed(action)
-- Phase 4: Update health state
call update-gpu-health
gpus: fleet.gpus
actions: actions
status: derive-overall-status(fleet, actions)
-- Wait 60s and repeat
Audit Trail Format
{
"run_id": "gpu-self-heal-20260718-001",
"timestamp": "2026-07-18T08:00:00Z",
"gpu": "ct8-rtx3090",
"issue": "thermal-critical",
"detected": { "temp_c": 87, "duration_s": 180 },
"action": "load-shedding",
"result": "resolved",
"verification": { "temp_c": 76, "after_s": 300 },
"escalated": false
}
Reporting
1. Knowledge Graph
Every action logged as [GPU-SELF-HEAL] <run_id> node with full audit trail.
2. Zulip Alerts (#agent-hub → alerts-gpu)
issues_fixed > 0→ "🛠 GPU Self-Heal — resolved"issues_escalated > 0→ "⚠ GPU Self-Heal — needs attention"- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
3. Weekly Benchmark Report
- Per-GPU tok/s trend over 7 days
- Regression alerts if any GPU degrades >10% week-over-week
Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
- Fan control: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
- Model restart: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
- Strix direct access: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
- VRAM thresholds: Tiered — 300MB/h (RTX 3090), 300MB/h (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
- CB auto-reset: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
- Benchmark baseline: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
- Predictive alerts: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
- Prometheus: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
Lessons Learned (2026-07-12, Updated 2026-07-18)
L1: API Key Standardization Is Critical
- All GPU llama-servers MUST use the same api-key as the LiteLLM config.
- RTX 5070 had
--api-key sk-loc...5678while LiteLLM sentnot-needed. This caused cascading 401 → fallback → timeout → 401 loops. - Rule: Any new GPU or model restart MUST verify api-key matches LiteLLM config (
not-neededfor direct routing).
L2: Fallback Chain Cascading Failures
- When one model returns 401 (auth) and another is slow (timeout), the fallback chain creates an infinite loop.
- Rule: If a model returns 401 (auth error), do NOT fall back to it again. Mark it as permanently failed for this request.
L3: Verify Running State, Not Docs
- RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18).
- Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models).
- Rule: Before making decisions, check
/proc/PID/cmdlineon GPU hosts.
L4: Infisical Is Not Always Available
- Keep a local
.envfallback forLITELLM_API_KEY. - Rule: Always verify credential source is reachable before relying on it.
L5: GPU Monitor Response Size Can Cause Self-Heal Crash
- gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response.
- Root cause: router poll returns accumulated data → cache balloons.
- Rule: Self-heal must enforce a read timeout AND max response size on every poll. If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
L6: Stable Aliases Replace Model Names
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
- Self-heal must use aliases for reporting and alerting, not model-specific names.
- Rule: All alert messages and KG nodes use the stable alias as the GPU identifier.