From bddbb22f03fadb8ef14e1f3f68f48bc126e5405f Mon Sep 17 00:00:00 2001 From: root Date: Sat, 18 Jul 2026 08:07:03 +0000 Subject: [PATCH] gpu-self-heal: refresh to current fleet baseline and topology - Synced model assignments to 2026-07-17 swaps (ThinkingCap, HauhauCS QAT, Genesis Hermes V3) - Added stable role-based aliases from gpu-fleet (gpu-dense, gpu-light, strix-moe) - Updated benchmark baselines to live values (74.9/169.6/62.9 tok/s) - Replaced router (port 9000) references with LiteLLM + direct routing - Replaced Prometheus exporter rule with sidecar health probe - Updated VRAM thresholds to match operational data (300/300/200 MB/h) - Added response size limit (1MB) to prevent OOM crashes - Added Lessons L5 (response size crash) and L6 (stable aliases) - Removed deprecated Rules 11-12 (router-specific distribution balance) --- gpu-self-heal.prose.md | 162 +++++++++++++++++++++++------------------ 1 file changed, 93 insertions(+), 69 deletions(-) diff --git a/gpu-self-heal.prose.md b/gpu-self-heal.prose.md index 5388a40..62cf205 100644 --- a/gpu-self-heal.prose.md +++ b/gpu-self-heal.prose.md @@ -6,10 +6,15 @@ description: > benchmarks, and predicts failures before they happen. Extends gpu-monitor (v2.1.0) with active remediation rules, Prometheus metrics consumption, VRAM trend analysis, and predictive alerting. + UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps. + Router (port 9000) references replaced with direct GPU routing. + Benchmark baselines refreshed to live values. + Prometheus exporters removed — not deployed; fall back to direct sidecar probes. + Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet. agent: abiba depends_on: - gpu-monitor.prose.md (live data source on .24:9100) - - gpu-fleet.prose.md (source of truth for topology) + - gpu-fleet.prose.md (source of truth for topology, aliases, model assignments) --- ## Maintains @@ -22,8 +27,8 @@ depends_on: ## Requires -- gpu-monitor:function — Live fleet data from .24:9100/gpu-data -- Prometheus exporters on all 3 GPUs (:9400/metrics) +- gpu-monitor:function — Live fleet data from localhost:9100/gpu-data +- Direct sidecar probe access to all GPU hosts (:8080/health) - SSH access to GPU hosts for restart operations ## Continuity @@ -35,21 +40,37 @@ depends_on: --- +## Current Fleet Baseline (2026-07-18) + +| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role | +|-------|-----|------|-------|------|-----|-------|------| +| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen | +| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks | +| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs | + +Key notes: +- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path. +- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. +- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first. +- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads. +- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom). +- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes). + ## Remediation Rules ### Rule 1: GPU Temperature Critical (>85°C for >2 min) - **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls - **Fix**: 1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control) - 2. Redirect new requests to cooler GPUs via LiteLLM fallback chains + 2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma) 3. If all GPUs hot, alert about cooling infrastructure - **Verify**: Temp drops below 80°C within 5 minutes - **Escalate after**: 3 verification failures → Zulip alert ### Rule 2: VRAM Leak Detection (tiered by GPU capacity) - **Detect**: VRAM growing at sustained rate over 6+ hour window - - RTX 3090 (24GB): ≥100MB/hour - - RTX 5070 (12GB): ≥50MB/hour + - RTX 3090 (24GB): ≥300MB/hour + - RTX 5070 (12GB): ≥300MB/hour - Strix Halo (64GB UMA): ≥200MB/hour - **Fix**: 1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux) @@ -69,6 +90,9 @@ depends_on: ### Rule 4: Benchmark Regression (>20% drop) - **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks + - RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s + - RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s + - Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s - **Fix**: 1. Check GPU utilization — if >90%, other process is competing 2. Check power limit — if throttled, restore to max @@ -78,32 +102,31 @@ depends_on: ### Rule 5: Circuit Breaker Stuck Open - **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy +- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router. - **Fix**: - 1. Verify GPU /health returns 200 - 2. If GPU healthy, send 1 test inference - 3. If test succeeds → reset circuit breaker via router API - 4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset - 5. Max 1 auto-reset per GPU per hour -- **Verify**: CB closes, inference succeeds, CB stays closed for 60s+ -- **Escalate after**: CB won't close after reset → router issue + 1. Verify GPU /health returns 200 on direct port (:8080) + 2. If GPU healthy, alert but do NOT reset via router API (deprecated) + 3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness + 4. Restart LiteLLM container on CT 116 if circuit breakers are stuck +- **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s +- **Escalate after**: LiteLLM restart doesn't clear → human investigation ### Rule 6: Strix Halo Unreachable - **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15) - **Fix**: 1. SSH to .15 → check llama-server process 2. Restart llama-server if not running - 3. Verify through both direct probe AND router -- **Verify**: Direct health probe returns 200, router reports Strix healthy + 3. Verify through both direct probe AND LiteLLM health +- **Verify**: Direct health probe returns 200, LiteLLM reports model healthy - **Escalate**: If host .15 itself is unreachable → infrastructure alert -### Rule 7: Prometheus Exporter Down -- **Detect**: Any GPU :9400/metrics unreachable for >2 polls +### Rule 7: GPU Data Source Unreachable (replaces old Prometheus rule) +- **Detect**: gpu-monitor endpoint (localhost:9100/gpu-data) or sidecar port (:8080) on any GPU unreachable for >2 polls - **Fix**: - 1. SSH to GPU host → check prometheus-exporter process - 2. Restart exporter if dead - 3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes - 4. If exporter is running but unreachable → check firewall/host networking -- **Verify**: :9400/metrics returns 200 + 1. If gpu-monitor is down: restart systemd service `gpu-monitor.service` on this host + 2. If sidecar is down: SSH to GPU host → check llama-server process → restart systemd service + 3. Fall back to direct nvidia-smi/rocm-smi probe via SSH if all API paths fail +- **Verify**: gpu-monitor returns healthy + all sidecars reachable - **Escalate after**: 3 failed restarts → networking issue ### Rule 8: Predictive Thermal Warning (two-tier) @@ -117,29 +140,31 @@ depends_on: - **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure ### Rule 9: Context Window Optimization -- **Detect**: Benchmark tok/s vs baseline for each GPU at current context - - RTX 3090 (128K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline - - RTX 5070 (128K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role - - Strix Halo (128K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline +- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K) + - RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%) + - RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%) + - Strix Halo (128K ctx, Genesis Hermes V3): baseline 70.5 tok/s — currently at 62.9 (89%) - **Fix**: - If tok/s > baseline → context has headroom, consider increasing - If tok/s < 90% baseline → reduce context by 25% and retest - If tok/s within 10% of baseline → optimal, no change + - Strix Halo at 89% of baseline → MONITOR but do not reduce yet (recent model swap may still be settling) - **Verify**: Re-benchmark after context change, confirm within 10% of target - **Escalate**: If context can't be adjusted without significant perf loss -### Rule 10: Workload Distribution Optimization +### Rule 10: Workload Distribution Optimization (updated 2026-07-18) - **Detect**: GPU roles misaligned with hardware capabilities - **Target distribution**: - - RTX 3090 (24GB, 128K, 75 tok/s) → Heavy reasoning, code gen, long conversations - - RTX 5070 (12GB, 128K, 76 tok/s) → Vision/image, web search, quick lightweight tasks - - Strix Halo (64GB, 128K, 72 tok/s) → Context compression, summarization, long docs + - RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM). + - RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM). + - Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM). +- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first. - **Fix**: - Alert if any GPU is handling workload outside its designated role - - Recommend Hermes agent profile updates to match workload to GPU + - Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe) - Track per-GPU request distribution via LiteLLM spend logs - **Verify**: Each GPU's request pattern matches its designated role within 24h -- **Escalate**: If role mismatch persists >48h → agent profile audit needed +- **Escalate**: If role mismatch persists >48h → agent alias audit needed --- @@ -148,7 +173,7 @@ depends_on: ```prose -- Phase 1: Fetch live GPU data let fleet = call gpu-monitor - endpoint: "http://192.168.68.24:9100/gpu-data" + endpoint: "http://localhost:9100/gpu-data" -- Phase 2: Evaluate each GPU against remediation rules let actions = [] @@ -164,27 +189,25 @@ for gpu in fleet.gpus: -- Rule 4: Benchmark regression let bench = fleet.benchmarks[gpu.hostname] - if bench.current_tok_sec < bench.baseline_tok_sec * 0.8: + if bench.current_tok_s < bench.baseline_tok_s * 0.8: push actions apply-benchmark-fix(gpu, bench) -- Rule 3: Model stuck -for model in fleet.router.available_models: +for model in fleet.summary.available_models: if model.consecutive_timeouts >= 3: push actions apply-model-restart(model) --- Rule 5: Circuit breaker -for cb in fleet.router.circuit_breaker: - if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu): - push actions apply-cb-reset(cb) +-- Rule 5: Circuit breaker check via LiteLLM (router deprecated) +if fleet.summary.circuit_breakers_open > 0: + push actions check-litellm-circuit-breakers() -- Rule 6: Strix Halo if not fleet.strix.running and pingable("192.168.68.15"): push actions apply-strix-restart() --- Rule 7: Prometheus exporters -for gpu in fleet.gpus: - if not prometheus_reachable(gpu.hostname, 9400): - push actions apply-exporter-restart(gpu) +-- Rule 7: GPU data source +if not fleet.gpus or len(fleet.gpus) < 2: + push actions check-gpu-monitor-service() -- Rule 8: Predictive thermal for gpu in fleet.gpus: @@ -212,12 +235,12 @@ call update-gpu-health ```json { - "run_id": "gpu-self-heal-20260712-001", - "timestamp": "2026-07-12T16:00:00Z", + "run_id": "gpu-self-heal-20260718-001", + "timestamp": "2026-07-18T08:00:00Z", "gpu": "ct8-rtx3090", "issue": "thermal-critical", "detected": { "temp_c": 87, "duration_s": 180 }, - "action": "set-fan-100pct", + "action": "load-shedding", "result": "resolved", "verification": { "temp_c": 76, "after_s": 300 }, "escalated": false @@ -236,52 +259,53 @@ Every action logged as `[GPU-SELF-HEAL] ` node with full audit trail. - `issues_escalated > 0` → "⚠ GPU Self-Heal — needs attention" - Every 100th clean cycle → "✅ GPU Fleet: All Clear" -### 3. Prometheus/Grafana Integration -- GPU self-heal actions exposed as Prometheus counter metrics -- Dashboard panel: "GPU Interventions (24h)" showing count/type/result - -### 4. Weekly Benchmark Report +### 3. Weekly Benchmark Report - Per-GPU tok/s trend over 7 days - Regression alerts if any GPU degrades >10% week-over-week --- -## Design Decisions (Grilled & Confirmed — 2026-07-12) +## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18) 1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect). 2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request. 3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart. -4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix). -5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU. -6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana. +4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data. +5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset. +6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor. 7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising. -8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down. +8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts. -## Lessons Learned (2026-07-12) +## Lessons Learned (2026-07-12, Updated 2026-07-18) ### L1: API Key Standardization Is Critical - All GPU llama-servers MUST use the same api-key as the LiteLLM config. - RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`. - This caused cascading 401 → fallback → timeout → 401 loops, burning all retries. -- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config. + This caused cascading 401 → fallback → timeout → 401 loops. +- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config (`not-needed` for direct routing). ### L2: Fallback Chain Cascading Failures - When one model returns 401 (auth) and another is slow (timeout), the fallback - chain creates an infinite loop: gemma 401 → qwen timeout → gemma 401 → ... + chain creates an infinite loop. - **Rule**: If a model returns 401 (auth error), do NOT fall back to it again. Mark it as permanently failed for this request. ### L3: Verify Running State, Not Docs -- RTX 3090 was documented at 128K context. Actually running at 256K. -- Parallel count wrong (docs said 2, actual is 1 on RTX 3090). +- RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18). +- Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models). - **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts. ### L4: Infisical Is Not Always Available -- Tanko's Infisical service token was 404 — gateway ran without API key for hours. -- **Rule**: Always keep a local `.env` fallback for `LITELLM_API_KEY`. -- Contract hermes-config-template Rule 3 updated. +- Keep a local `.env` fallback for `LITELLM_API_KEY`. +- **Rule**: Always verify credential source is reachable before relying on it. -### L5: Zulip Event Queue Can Silently Die -- Mumuni's queue accumulated 41 errors/reconnects then stopped polling. - Gateway was running but ignoring all messages. -- **Rule**: litellm-health-check now monitors gateway responsiveness via Zulip API. +### L5: GPU Monitor Response Size Can Cause Self-Heal Crash +- gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response. +- Root cause: router poll returns accumulated data → cache balloons. +- **Rule**: Self-heal must enforce a read timeout AND max response size on every poll. + If monitor response > 1MB, log a warning and skip the cycle rather than crashing. + +### L6: Stable Aliases Replace Model Names +- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15. +- Self-heal must use aliases for reporting and alerting, not model-specific names. +- **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.