Files
prose-contracts/gpu-self-heal.prose.md
T
Abiba dc572889f8 contracts: sync to ground truth — ornith-1.0-35b→strix-moe, 256K all GPUs, real LiteLLM timeouts
Verified on ground 2026-07-16 against CT 116 litellm_config.yaml + GPU hosts:
- AMD host serves qwen3.6-35B-udq4 (LiteLLM alias strix-moe); ornith-1.0-35b does NOT exist
- All 3 GPUs at 256K ctx, parallel 2 (RTX 3090 was listed 128K/parallel 1)
- LiteLLM timeouts: qwen 300s, gemma 120s, strix 300s (were stale 90s/120s)
- Added LiteLLM model surface + key scoping to litellm-self-heal
- Patched health-check script path ref

Files: litellm-self-heal, litellm-health, gpu-fleet, gpu-self-heal,
zulip-adapter-lessons, abiba-zulip-restore, hermes-agent-baseline,
delegation-prose-contract, mumuni-delegation-prose-contract
2026-07-16 17:02:44 +00:00

12 KiB

kind, name, description, agent, depends_on
kind name description agent depends_on
responsibility gpu-self-heal GPU fleet self-healing — detects anomalies, applies remediation, tracks benchmarks, and predicts failures before they happen. Extends gpu-monitor (v2.1.0) with active remediation rules, Prometheus metrics consumption, VRAM trend analysis, and predictive alerting. abiba
gpu-monitor.prose.md (live data source on .24:9100)
gpu-fleet.prose.md (source of truth for topology)

Maintains

  • gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array }
  • gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail
  • benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts }
  • vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours }
  • circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset }

Requires

  • gpu-monitor:function — Live fleet data from .24:9100/gpu-data
  • Prometheus exporters on all 3 GPUs (:9400/metrics)
  • SSH access to GPU hosts for restart operations

Continuity

  • Self-driven: check every 60 seconds against GPU monitor data
  • Also wakes on gpu-fleet health degradation
  • On fix: verify with benchmark inference test before declaring resolved
  • Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert

Remediation Rules

Rule 1: GPU Temperature Critical (>85°C for >2 min)

  • Detect: Any GPU temp >85°C sustained for 2+ consecutive polls
  • Fix:
    1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
    2. Redirect new requests to cooler GPUs via LiteLLM fallback chains
    3. If all GPUs hot, alert about cooling infrastructure
  • Verify: Temp drops below 80°C within 5 minutes
  • Escalate after: 3 verification failures → Zulip alert

Rule 2: VRAM Leak Detection (tiered by GPU capacity)

  • Detect: VRAM growing at sustained rate over 6+ hour window
    • RTX 3090 (24GB): ≥100MB/hour
    • RTX 5070 (12GB): ≥50MB/hour
    • Strix Halo (64GB UMA): ≥200MB/hour
  • Fix:
    1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
    2. If llama-server is the growth source → restart with memory cap flag
    3. If unknown process → kill and alert
  • Verify: VRAM growth rate drops below threshold
  • Escalate after: persistent leak after restart → hardware investigation

Rule 3: Model Inference Timeout / GPU Stuck

  • Detect: >50% failure rate over 60s window + 30s grace period (not just single stuck request)
  • Fix:
    1. Restart llama-server on affected GPU host
    2. Wait 15s for model to reload
    3. Run benchmark inference test
  • Verify: Model returns 200 with <30s response, failure rate drops to 0%
  • Escalate after: 3 restarts in 1 hour → GPU hardware check

Rule 4: Benchmark Regression (>20% drop)

  • Detect: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
  • Fix:
    1. Check GPU utilization — if >90%, other process is competing
    2. Check power limit — if throttled, restore to max
    3. Check thermal — if hot, apply Rule 1
  • Verify: Benchmark returns to within 10% of baseline
  • Escalate after: persistent regression → possible hardware degradation

Rule 5: Circuit Breaker Stuck Open

  • Detect: Circuit breaker open >10 minutes with GPU reporting healthy
  • Fix:
    1. Verify GPU /health returns 200
    2. If GPU healthy, send 1 test inference
    3. If test succeeds → reset circuit breaker via router API
    4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset
    5. Max 1 auto-reset per GPU per hour
  • Verify: CB closes, inference succeeds, CB stays closed for 60s+
  • Escalate after: CB won't close after reset → router issue

Rule 6: Strix Halo Unreachable

  • Detect: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
  • Fix:
    1. SSH to .15 → check llama-server process
    2. Restart llama-server if not running
    3. Verify through both direct probe AND router
  • Verify: Direct health probe returns 200, router reports Strix healthy
  • Escalate: If host .15 itself is unreachable → infrastructure alert

Rule 7: Prometheus Exporter Down

  • Detect: Any GPU :9400/metrics unreachable for >2 polls
  • Fix:
    1. SSH to GPU host → check prometheus-exporter process
    2. Restart exporter if dead
    3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes
    4. If exporter is running but unreachable → check firewall/host networking
  • Verify: :9400/metrics returns 200
  • Escalate after: 3 failed restarts → networking issue

Rule 8: Predictive Thermal Warning (two-tier)

  • Detect:
    • Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert
    • Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding
  • Fix:
    • Tier 1: silently reduce parallel requests to that GPU by 50%
    • Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic
  • Verify: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2)
  • Escalate: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure

Rule 9: Context Window Optimization

  • Detect: Benchmark tok/s vs baseline for each GPU at current context
    • RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
    • RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
    • Strix Halo (256K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
  • Fix:
    • If tok/s > baseline → context has headroom, consider increasing
    • If tok/s < 90% baseline → reduce context by 25% and retest
    • If tok/s within 10% of baseline → optimal, no change
  • Verify: Re-benchmark after context change, confirm within 10% of target
  • Escalate: If context can't be adjusted without significant perf loss

Rule 10: Workload Distribution Optimization

  • Detect: GPU roles misaligned with hardware capabilities
  • Target distribution:
    • RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations
    • RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
    • Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs
  • Fix:
    • Alert if any GPU is handling workload outside its designated role
    • Recommend Hermes agent profile updates to match workload to GPU
    • Track per-GPU request distribution via LiteLLM spend logs
  • Verify: Each GPU's request pattern matches its designated role within 24h
  • Escalate: If role mismatch persists >48h → agent profile audit needed

Execution

-- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor
  endpoint: "http://192.168.68.24:9100/gpu-data"

-- Phase 2: Evaluate each GPU against remediation rules
let actions = []
for gpu in fleet.gpus:
  -- Rule 1: Thermal critical
  if gpu.temp_c > 85 and sustained_for(gpu, 120):
    push actions apply-thermal-fix(gpu)
  
  -- Rule 2: VRAM leak
  let vram_rate = calculate-vram-trend(gpu, hours=6)
  if vram_rate > 50:
    push actions apply-vram-fix(gpu, vram_rate)
  
  -- Rule 4: Benchmark regression
  let bench = fleet.benchmarks[gpu.hostname]
  if bench.current_tok_sec < bench.baseline_tok_sec * 0.8:
    push actions apply-benchmark-fix(gpu, bench)

-- Rule 3: Model stuck
for model in fleet.router.available_models:
  if model.consecutive_timeouts >= 3:
    push actions apply-model-restart(model)

-- Rule 5: Circuit breaker
for cb in fleet.router.circuit_breaker:
  if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu):
    push actions apply-cb-reset(cb)

-- Rule 6: Strix Halo
if not fleet.strix.running and pingable("192.168.68.15"):
  push actions apply-strix-restart()

-- Rule 7: Prometheus exporters
for gpu in fleet.gpus:
  if not prometheus_reachable(gpu.hostname, 9400):
    push actions apply-exporter-restart(gpu)

-- Rule 8: Predictive thermal
for gpu in fleet.gpus:
  let rise_rate = calculate-temp-rise(gpu, minutes=5)
  if rise_rate > 2.0 and gpu.temp_c < 80:
    push actions apply-proactive-cooling(gpu)

-- Phase 3: Execute actions, verify, log
for action in actions:
  let result = execute-with-verify(action)
  log-to-kg(action, result)
  if result.failed:
    escalate-if-needed(action)

-- Phase 4: Update health state
call update-gpu-health
  gpus: fleet.gpus
  actions: actions
  status: derive-overall-status(fleet, actions)

-- Wait 60s and repeat

Audit Trail Format

{
  "run_id": "gpu-self-heal-20260712-001",
  "timestamp": "2026-07-12T16:00:00Z",
  "gpu": "ct8-rtx3090",
  "issue": "thermal-critical",
  "detected": { "temp_c": 87, "duration_s": 180 },
  "action": "set-fan-100pct",
  "result": "resolved",
  "verification": { "temp_c": 76, "after_s": 300 },
  "escalated": false
}

Reporting

1. Knowledge Graph

Every action logged as [GPU-SELF-HEAL] <run_id> node with full audit trail.

2. Zulip Alerts (#agent-hub → alerts-gpu)

  • issues_fixed > 0 → "🛠 GPU Self-Heal — resolved"
  • issues_escalated > 0 → "⚠ GPU Self-Heal — needs attention"
  • Every 100th clean cycle → " GPU Fleet: All Clear"

3. Prometheus/Grafana Integration

  • GPU self-heal actions exposed as Prometheus counter metrics
  • Dashboard panel: "GPU Interventions (24h)" showing count/type/result

4. Weekly Benchmark Report

  • Per-GPU tok/s trend over 7 days
  • Regression alerts if any GPU degrades >10% week-over-week

Design Decisions (Grilled & Confirmed — 2026-07-12)

  1. Fan control: NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
  2. Model restart: Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
  3. Strix direct access: Open firewall .15:8080 → .24 for direct health probe + restart.
  4. VRAM thresholds: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix).
  5. CB auto-reset: With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU.
  6. Benchmark baseline: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana.
  7. Predictive alerts: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
  8. Prometheus: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.

Lessons Learned (2026-07-12)

L1: API Key Standardization Is Critical

  • All GPU llama-servers MUST use the same api-key as the LiteLLM config.
  • RTX 5070 had --api-key sk-loc...5678 while LiteLLM sent not-needed. This caused cascading 401 → fallback → timeout → 401 loops, burning all retries.
  • Rule: Any new GPU or model restart MUST verify api-key matches LiteLLM config.

L2: Fallback Chain Cascading Failures

  • When one model returns 401 (auth) and another is slow (timeout), the fallback chain creates an infinite loop: gemma 401 → qwen timeout → gemma 401 → ...
  • Rule: If a model returns 401 (auth error), do NOT fall back to it again. Mark it as permanently failed for this request.

L3: Verify Running State, Not Docs

  • RTX 3090 was documented at 128K context. Actually running at 256K.
  • Parallel count wrong (docs said 2, actual is 1 on RTX 3090).
  • Rule: Before making decisions, check /proc/PID/cmdline on GPU hosts.

L4: Infisical Is Not Always Available

  • Tanko's Infisical service token was 404 — gateway ran without API key for hours.
  • Rule: Always keep a local .env fallback for LITELLM_API_KEY.
  • Contract hermes-config-template Rule 3 updated.

L5: Zulip Event Queue Can Silently Die

  • Mumuni's queue accumulated 41 errors/reconnects then stopped polling. Gateway was running but ignoring all messages.
  • Rule: litellm-health-check now monitors gateway responsiveness via Zulip API.