Files
prose-contracts/gpu-self-heal.prose.md
root c9359e1808
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
fix: redirect health check logs from knowledge graph to Gitea (hard rule)
Logs (LITELLM-HEALTH, GPU-SELF-HEAL, PM2) now pushed to SyslogSolution/health-logs
instead of creating orphan nodes in the shared knowledge graph.

- litellm-self-heal: Phase 4 now calls gitea-logger instead of kg-logger
- gpu-self-heal: Reporting section updated to Gitea path
- pm2-self-heal: Log step redirected to Gitea
- 271 existing orphan nodes remain in graph (no delete tool available)
2026-07-28 21:20:56 +00:00

15 KiB

kind, name, description, agent, depends_on
kind name description agent depends_on
responsibility gpu-self-heal GPU fleet self-healing — detects anomalies, applies remediation, tracks benchmarks, and predicts failures before they happen. Extends gpu-monitor (v2.1.0) with active remediation rules, Prometheus metrics consumption, VRAM trend analysis, and predictive alerting. UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps. Router (port 9000) references replaced with direct GPU routing. Benchmark baselines refreshed to live values. Prometheus exporters removed — not deployed; fall back to direct sidecar probes. Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet. abiba
gpu-monitor.prose.md (live data source on .24:9100)
gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)

Maintains

  • gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array }
  • gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail
  • benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts }
  • vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours }
  • circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset }

Requires

  • gpu-monitor:function — Live fleet data from localhost:9100/gpu-data
  • Direct sidecar probe access to all GPU hosts (:8080/health)
  • SSH access to GPU hosts for restart operations

Continuity

  • Self-driven: check every 60 seconds against GPU monitor data
  • Also wakes on gpu-fleet health degradation
  • On fix: verify with benchmark inference test before declaring resolved
  • Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert

Current Fleet Baseline (2026-07-18)

Alias GPU Host Model VRAM Ctx tok/s Role
gpu-dense RTX 3090 24GB ct8 (.8:8080) ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision 21.6/24.6GB (88%) 128K 74.9 Heavy reasoning, code gen
gpu-light RTX 5070 12GB ct110 (.110:8080) HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft 10.1/12.2GB (83%) 128K 169.6 Vision, web extract, light tasks
strix-moe Strix Halo 64GB ct15 (.15:8080) qwen3.6-35B-udq4 ~10/64GB (16%) 128K 62.9 Compression, summarization, long docs

Key notes:

  • All models use direct GPU routing via LiteLLM (api_key: not-needed). Router (port 9000) is deprecated and NOT in the inference path.
  • Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
  • RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
  • Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
  • RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
  • RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).

Remediation Rules

Rule 1: GPU Temperature Critical (>85°C for >2 min)

  • Detect: Any GPU temp >85°C sustained for 2+ consecutive polls
  • Fix:
    1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
    2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
    3. If all GPUs hot, alert about cooling infrastructure
  • Verify: Temp drops below 80°C within 5 minutes
  • Escalate after: 3 verification failures → Zulip alert

Rule 2: VRAM Leak Detection (tiered by GPU capacity)

  • Detect: VRAM growing at sustained rate over 6+ hour window
    • RTX 3090 (24GB): ≥300MB/hour
    • RTX 5070 (12GB): ≥300MB/hour
    • Strix Halo (64GB UMA): ≥200MB/hour
  • Fix:
    1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
    2. If llama-server is the growth source → restart with memory cap flag
    3. If unknown process → kill and alert
  • Verify: VRAM growth rate drops below threshold
  • Escalate after: persistent leak after restart → hardware investigation

Rule 3: Model Inference Timeout / GPU Stuck

  • Detect: >50% failure rate over 60s window + 30s grace period (not just single stuck request)
  • Fix:
    1. Restart llama-server on affected GPU host
    2. Wait 15s for model to reload
    3. Run benchmark inference test
  • Verify: Model returns 200 with <30s response, failure rate drops to 0%
  • Escalate after: 3 restarts in 1 hour → GPU hardware check

Rule 4: Benchmark Regression (>20% drop)

  • Detect: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
    • RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s
    • RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s
    • Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s
  • Fix:
    1. Check GPU utilization — if >90%, other process is competing
    2. Check power limit — if throttled, restore to max
    3. Check thermal — if hot, apply Rule 1
  • Verify: Benchmark returns to within 10% of baseline
  • Escalate after: persistent regression → possible hardware degradation

Rule 5: Circuit Breaker Stuck Open

  • Detect: Circuit breaker open >10 minutes with GPU reporting healthy
  • Note: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
  • Fix:
    1. Verify GPU /health returns 200 on direct port (:8080)
    2. If GPU healthy, alert but do NOT reset via router API (deprecated)
    3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
    4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
  • Verify: LiteLLM returns healthy, circuit breaker clears within 60s
  • Escalate after: LiteLLM restart doesn't clear → human investigation

Rule 6: Strix Halo Unreachable

  • Detect: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
  • Fix:
    1. SSH to .15 → check llama-server process
    2. Restart llama-server if not running
    3. Verify through both direct probe AND LiteLLM health
  • Verify: Direct health probe returns 200, LiteLLM reports model healthy
  • Escalate: If host .15 itself is unreachable → infrastructure alert

Rule 7: GPU Data Source Unreachable (replaces old Prometheus rule)

  • Detect: gpu-monitor endpoint (localhost:9100/gpu-data) or sidecar port (:8080) on any GPU unreachable for >2 polls
  • Fix:
    1. If gpu-monitor is down: restart systemd service gpu-monitor.service on this host
    2. If sidecar is down: SSH to GPU host → check llama-server process → restart systemd service
    3. Fall back to direct nvidia-smi/rocm-smi probe via SSH if all API paths fail
  • Verify: gpu-monitor returns healthy + all sidecars reachable
  • Escalate after: 3 failed restarts → networking issue

Rule 8: Predictive Thermal Warning (two-tier)

  • Detect:
    • Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert
    • Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding
  • Fix:
    • Tier 1: silently reduce parallel requests to that GPU by 50%
    • Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic
  • Verify: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2)
  • Escalate: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure

Rule 9: Context Window Optimization

  • Detect: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
    • RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
    • RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
    • Strix Halo (128K ctx, qwen3.6-35B-udq4): baseline 70.5 tok/s — currently at 62.9 (89%)
  • Fix:
    • If tok/s > baseline → context has headroom, consider increasing
    • If tok/s < 90% baseline → reduce context by 25% and retest
    • If tok/s within 10% of baseline → optimal, no change
    • Strix Halo at 89% of baseline → MONITOR but do not reduce yet (recent model swap may still be settling)
  • Verify: Re-benchmark after context change, confirm within 10% of target
  • Escalate: If context can't be adjusted without significant perf loss

Rule 10: Workload Distribution Optimization (updated 2026-07-18)

  • Detect: GPU roles misaligned with hardware capabilities
  • Target distribution:
    • RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
    • RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
    • Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
  • Note: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
  • Fix:
    • Alert if any GPU is handling workload outside its designated role
    • Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe)
    • Track per-GPU request distribution via LiteLLM spend logs
  • Verify: Each GPU's request pattern matches its designated role within 24h
  • Escalate: If role mismatch persists >48h → agent alias audit needed

Execution

-- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor
  endpoint: "http://localhost:9100/gpu-data"

-- Phase 2: Evaluate each GPU against remediation rules
let actions = []
for gpu in fleet.gpus:
  -- Rule 1: Thermal critical
  if gpu.temp_c > 85 and sustained_for(gpu, 120):
    push actions apply-thermal-fix(gpu)
  
  -- Rule 2: VRAM leak
  let vram_rate = calculate-vram-trend(gpu, hours=6)
  if vram_rate > 50:
    push actions apply-vram-fix(gpu, vram_rate)
  
  -- Rule 4: Benchmark regression
  let bench = fleet.benchmarks[gpu.hostname]
  if bench.current_tok_s < bench.baseline_tok_s * 0.8:
    push actions apply-benchmark-fix(gpu, bench)

-- Rule 3: Model stuck
for model in fleet.summary.available_models:
  if model.consecutive_timeouts >= 3:
    push actions apply-model-restart(model)

-- Rule 5: Circuit breaker check via LiteLLM (router deprecated)
if fleet.summary.circuit_breakers_open > 0:
  push actions check-litellm-circuit-breakers()

-- Rule 6: Strix Halo
if not fleet.strix.running and pingable("192.168.68.15"):
  push actions apply-strix-restart()

-- Rule 7: GPU data source
if not fleet.gpus or len(fleet.gpus) < 2:
  push actions check-gpu-monitor-service()

-- Rule 8: Predictive thermal
for gpu in fleet.gpus:
  let rise_rate = calculate-temp-rise(gpu, minutes=5)
  if rise_rate > 2.0 and gpu.temp_c < 80:
    push actions apply-proactive-cooling(gpu)

-- Phase 3: Execute actions, verify, log
for action in actions:
  let result = execute-with-verify(action)
  log-to-kg(action, result)
  if result.failed:
    escalate-if-needed(action)

-- Phase 4: Update health state
call update-gpu-health
  gpus: fleet.gpus
  actions: actions
  status: derive-overall-status(fleet, actions)

-- Wait 60s and repeat

Audit Trail Format

{
  "run_id": "gpu-self-heal-20260718-001",
  "timestamp": "2026-07-18T08:00:00Z",
  "gpu": "ct8-rtx3090",
  "issue": "thermal-critical",
  "detected": { "temp_c": 87, "duration_s": 180 },
  "action": "load-shedding",
  "result": "resolved",
  "verification": { "temp_c": 76, "after_s": 300 },
  "escalated": false
}

Reporting

1. Gitea Log (not knowledge graph — hard rule)

Pushed to SyslogSolution/health-logs/gpu/{run_id}.json — versioned, searchable, not in graph.

2. Zulip Alerts (#agent-hub → alerts-gpu)

  • issues_fixed > 0 → "🛠 GPU Self-Heal — resolved"
  • issues_escalated > 0 → "⚠ GPU Self-Heal — needs attention"
  • Every 100th clean cycle → " GPU Fleet: All Clear"

3. Weekly Benchmark Report

  • Per-GPU tok/s trend over 7 days
  • Regression alerts if any GPU degrades >10% week-over-week

Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)

  1. Fan control: NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
  2. Model restart: Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
  3. Strix direct access: Open firewall .15:8080 → .24 for direct health probe + restart.
  4. VRAM thresholds: Tiered — 300MB/h (RTX 3090), 300MB/h (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
  5. CB auto-reset: Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
  6. Benchmark baseline: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
  7. Predictive alerts: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
  8. Prometheus: Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.

Lessons Learned (2026-07-12, Updated 2026-07-18)

L1: API Key Standardization Is Critical

  • All GPU llama-servers MUST use the same api-key as the LiteLLM config.
  • RTX 5070 had --api-key sk-loc...5678 while LiteLLM sent not-needed. This caused cascading 401 → fallback → timeout → 401 loops.
  • Rule: Any new GPU or model restart MUST verify api-key matches LiteLLM config (not-needed for direct routing).

L2: Fallback Chain Cascading Failures

  • When one model returns 401 (auth) and another is slow (timeout), the fallback chain creates an infinite loop.
  • Rule: If a model returns 401 (auth error), do NOT fall back to it again. Mark it as permanently failed for this request.

L3: Verify Running State, Not Docs

  • RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18).
  • Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models).
  • Rule: Before making decisions, check /proc/PID/cmdline on GPU hosts.

L4: Infisical Is Not Always Available

  • Keep a local .env fallback for LITELLM_API_KEY.
  • Rule: Always verify credential source is reachable before relying on it.

L5: GPU Monitor Response Size Can Cause Self-Heal Crash

  • gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response.
  • Root cause: router poll returns accumulated data → cache balloons.
  • Rule: Self-heal must enforce a read timeout AND max response size on every poll. If monitor response > 1MB, log a warning and skip the cycle rather than crashing.

L6: Stable Aliases Replace Model Names

  • gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
  • Self-heal must use aliases for reporting and alerting, not model-specific names.
  • Rule: All alert messages and KG nodes use the stable alias as the GPU identifier.