gpu-self-heal: refresh to current fleet baseline and topology #21
+93
-69
@@ -6,10 +6,15 @@ description: >
|
|||||||
benchmarks, and predicts failures before they happen. Extends gpu-monitor
|
benchmarks, and predicts failures before they happen. Extends gpu-monitor
|
||||||
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
|
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
|
||||||
VRAM trend analysis, and predictive alerting.
|
VRAM trend analysis, and predictive alerting.
|
||||||
|
UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps.
|
||||||
|
Router (port 9000) references replaced with direct GPU routing.
|
||||||
|
Benchmark baselines refreshed to live values.
|
||||||
|
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
|
||||||
|
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
|
||||||
agent: abiba
|
agent: abiba
|
||||||
depends_on:
|
depends_on:
|
||||||
- gpu-monitor.prose.md (live data source on .24:9100)
|
- gpu-monitor.prose.md (live data source on .24:9100)
|
||||||
- gpu-fleet.prose.md (source of truth for topology)
|
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
|
||||||
---
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -22,8 +27,8 @@ depends_on:
|
|||||||
|
|
||||||
## Requires
|
## Requires
|
||||||
|
|
||||||
- gpu-monitor:function — Live fleet data from .24:9100/gpu-data
|
- gpu-monitor:function — Live fleet data from localhost:9100/gpu-data
|
||||||
- Prometheus exporters on all 3 GPUs (:9400/metrics)
|
- Direct sidecar probe access to all GPU hosts (:8080/health)
|
||||||
- SSH access to GPU hosts for restart operations
|
- SSH access to GPU hosts for restart operations
|
||||||
|
|
||||||
## Continuity
|
## Continuity
|
||||||
@@ -35,21 +40,37 @@ depends_on:
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## Current Fleet Baseline (2026-07-18)
|
||||||
|
|
||||||
|
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|
||||||
|
|-------|-----|------|-------|------|-----|-------|------|
|
||||||
|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
|
||||||
|
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
|
||||||
|
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
|
||||||
|
|
||||||
|
Key notes:
|
||||||
|
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
||||||
|
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
|
||||||
|
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
|
||||||
|
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
||||||
|
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
||||||
|
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
|
||||||
|
|
||||||
## Remediation Rules
|
## Remediation Rules
|
||||||
|
|
||||||
### Rule 1: GPU Temperature Critical (>85°C for >2 min)
|
### Rule 1: GPU Temperature Critical (>85°C for >2 min)
|
||||||
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
|
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
|
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
|
||||||
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains
|
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
|
||||||
3. If all GPUs hot, alert about cooling infrastructure
|
3. If all GPUs hot, alert about cooling infrastructure
|
||||||
- **Verify**: Temp drops below 80°C within 5 minutes
|
- **Verify**: Temp drops below 80°C within 5 minutes
|
||||||
- **Escalate after**: 3 verification failures → Zulip alert
|
- **Escalate after**: 3 verification failures → Zulip alert
|
||||||
|
|
||||||
### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
|
### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
|
||||||
- **Detect**: VRAM growing at sustained rate over 6+ hour window
|
- **Detect**: VRAM growing at sustained rate over 6+ hour window
|
||||||
- RTX 3090 (24GB): ≥100MB/hour
|
- RTX 3090 (24GB): ≥300MB/hour
|
||||||
- RTX 5070 (12GB): ≥50MB/hour
|
- RTX 5070 (12GB): ≥300MB/hour
|
||||||
- Strix Halo (64GB UMA): ≥200MB/hour
|
- Strix Halo (64GB UMA): ≥200MB/hour
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
|
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
|
||||||
@@ -69,6 +90,9 @@ depends_on:
|
|||||||
|
|
||||||
### Rule 4: Benchmark Regression (>20% drop)
|
### Rule 4: Benchmark Regression (>20% drop)
|
||||||
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
|
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
|
||||||
|
- RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s
|
||||||
|
- RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s
|
||||||
|
- Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. Check GPU utilization — if >90%, other process is competing
|
1. Check GPU utilization — if >90%, other process is competing
|
||||||
2. Check power limit — if throttled, restore to max
|
2. Check power limit — if throttled, restore to max
|
||||||
@@ -78,32 +102,31 @@ depends_on:
|
|||||||
|
|
||||||
### Rule 5: Circuit Breaker Stuck Open
|
### Rule 5: Circuit Breaker Stuck Open
|
||||||
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
|
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
|
||||||
|
- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. Verify GPU /health returns 200
|
1. Verify GPU /health returns 200 on direct port (:8080)
|
||||||
2. If GPU healthy, send 1 test inference
|
2. If GPU healthy, alert but do NOT reset via router API (deprecated)
|
||||||
3. If test succeeds → reset circuit breaker via router API
|
3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
|
||||||
4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset
|
4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
|
||||||
5. Max 1 auto-reset per GPU per hour
|
- **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s
|
||||||
- **Verify**: CB closes, inference succeeds, CB stays closed for 60s+
|
- **Escalate after**: LiteLLM restart doesn't clear → human investigation
|
||||||
- **Escalate after**: CB won't close after reset → router issue
|
|
||||||
|
|
||||||
### Rule 6: Strix Halo Unreachable
|
### Rule 6: Strix Halo Unreachable
|
||||||
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
|
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. SSH to .15 → check llama-server process
|
1. SSH to .15 → check llama-server process
|
||||||
2. Restart llama-server if not running
|
2. Restart llama-server if not running
|
||||||
3. Verify through both direct probe AND router
|
3. Verify through both direct probe AND LiteLLM health
|
||||||
- **Verify**: Direct health probe returns 200, router reports Strix healthy
|
- **Verify**: Direct health probe returns 200, LiteLLM reports model healthy
|
||||||
- **Escalate**: If host .15 itself is unreachable → infrastructure alert
|
- **Escalate**: If host .15 itself is unreachable → infrastructure alert
|
||||||
|
|
||||||
### Rule 7: Prometheus Exporter Down
|
### Rule 7: GPU Data Source Unreachable (replaces old Prometheus rule)
|
||||||
- **Detect**: Any GPU :9400/metrics unreachable for >2 polls
|
- **Detect**: gpu-monitor endpoint (localhost:9100/gpu-data) or sidecar port (:8080) on any GPU unreachable for >2 polls
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. SSH to GPU host → check prometheus-exporter process
|
1. If gpu-monitor is down: restart systemd service `gpu-monitor.service` on this host
|
||||||
2. Restart exporter if dead
|
2. If sidecar is down: SSH to GPU host → check llama-server process → restart systemd service
|
||||||
3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes
|
3. Fall back to direct nvidia-smi/rocm-smi probe via SSH if all API paths fail
|
||||||
4. If exporter is running but unreachable → check firewall/host networking
|
- **Verify**: gpu-monitor returns healthy + all sidecars reachable
|
||||||
- **Verify**: :9400/metrics returns 200
|
|
||||||
- **Escalate after**: 3 failed restarts → networking issue
|
- **Escalate after**: 3 failed restarts → networking issue
|
||||||
|
|
||||||
### Rule 8: Predictive Thermal Warning (two-tier)
|
### Rule 8: Predictive Thermal Warning (two-tier)
|
||||||
@@ -117,29 +140,31 @@ depends_on:
|
|||||||
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
|
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
|
||||||
|
|
||||||
### Rule 9: Context Window Optimization
|
### Rule 9: Context Window Optimization
|
||||||
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context
|
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
|
||||||
- RTX 3090 (128K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
|
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
|
||||||
- RTX 5070 (128K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
|
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
|
||||||
- Strix Halo (128K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
|
- Strix Halo (128K ctx, Genesis Hermes V3): baseline 70.5 tok/s — currently at 62.9 (89%)
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
- If tok/s > baseline → context has headroom, consider increasing
|
- If tok/s > baseline → context has headroom, consider increasing
|
||||||
- If tok/s < 90% baseline → reduce context by 25% and retest
|
- If tok/s < 90% baseline → reduce context by 25% and retest
|
||||||
- If tok/s within 10% of baseline → optimal, no change
|
- If tok/s within 10% of baseline → optimal, no change
|
||||||
|
- Strix Halo at 89% of baseline → MONITOR but do not reduce yet (recent model swap may still be settling)
|
||||||
- **Verify**: Re-benchmark after context change, confirm within 10% of target
|
- **Verify**: Re-benchmark after context change, confirm within 10% of target
|
||||||
- **Escalate**: If context can't be adjusted without significant perf loss
|
- **Escalate**: If context can't be adjusted without significant perf loss
|
||||||
|
|
||||||
### Rule 10: Workload Distribution Optimization
|
### Rule 10: Workload Distribution Optimization (updated 2026-07-18)
|
||||||
- **Detect**: GPU roles misaligned with hardware capabilities
|
- **Detect**: GPU roles misaligned with hardware capabilities
|
||||||
- **Target distribution**:
|
- **Target distribution**:
|
||||||
- RTX 3090 (24GB, 128K, 75 tok/s) → Heavy reasoning, code gen, long conversations
|
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
|
||||||
- RTX 5070 (12GB, 128K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
|
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
|
||||||
- Strix Halo (64GB, 128K, 72 tok/s) → Context compression, summarization, long docs
|
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
|
||||||
|
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
- Alert if any GPU is handling workload outside its designated role
|
- Alert if any GPU is handling workload outside its designated role
|
||||||
- Recommend Hermes agent profile updates to match workload to GPU
|
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe)
|
||||||
- Track per-GPU request distribution via LiteLLM spend logs
|
- Track per-GPU request distribution via LiteLLM spend logs
|
||||||
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
||||||
- **Escalate**: If role mismatch persists >48h → agent profile audit needed
|
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -148,7 +173,7 @@ depends_on:
|
|||||||
```prose
|
```prose
|
||||||
-- Phase 1: Fetch live GPU data
|
-- Phase 1: Fetch live GPU data
|
||||||
let fleet = call gpu-monitor
|
let fleet = call gpu-monitor
|
||||||
endpoint: "http://192.168.68.24:9100/gpu-data"
|
endpoint: "http://localhost:9100/gpu-data"
|
||||||
|
|
||||||
-- Phase 2: Evaluate each GPU against remediation rules
|
-- Phase 2: Evaluate each GPU against remediation rules
|
||||||
let actions = []
|
let actions = []
|
||||||
@@ -164,27 +189,25 @@ for gpu in fleet.gpus:
|
|||||||
|
|
||||||
-- Rule 4: Benchmark regression
|
-- Rule 4: Benchmark regression
|
||||||
let bench = fleet.benchmarks[gpu.hostname]
|
let bench = fleet.benchmarks[gpu.hostname]
|
||||||
if bench.current_tok_sec < bench.baseline_tok_sec * 0.8:
|
if bench.current_tok_s < bench.baseline_tok_s * 0.8:
|
||||||
push actions apply-benchmark-fix(gpu, bench)
|
push actions apply-benchmark-fix(gpu, bench)
|
||||||
|
|
||||||
-- Rule 3: Model stuck
|
-- Rule 3: Model stuck
|
||||||
for model in fleet.router.available_models:
|
for model in fleet.summary.available_models:
|
||||||
if model.consecutive_timeouts >= 3:
|
if model.consecutive_timeouts >= 3:
|
||||||
push actions apply-model-restart(model)
|
push actions apply-model-restart(model)
|
||||||
|
|
||||||
-- Rule 5: Circuit breaker
|
-- Rule 5: Circuit breaker check via LiteLLM (router deprecated)
|
||||||
for cb in fleet.router.circuit_breaker:
|
if fleet.summary.circuit_breakers_open > 0:
|
||||||
if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu):
|
push actions check-litellm-circuit-breakers()
|
||||||
push actions apply-cb-reset(cb)
|
|
||||||
|
|
||||||
-- Rule 6: Strix Halo
|
-- Rule 6: Strix Halo
|
||||||
if not fleet.strix.running and pingable("192.168.68.15"):
|
if not fleet.strix.running and pingable("192.168.68.15"):
|
||||||
push actions apply-strix-restart()
|
push actions apply-strix-restart()
|
||||||
|
|
||||||
-- Rule 7: Prometheus exporters
|
-- Rule 7: GPU data source
|
||||||
for gpu in fleet.gpus:
|
if not fleet.gpus or len(fleet.gpus) < 2:
|
||||||
if not prometheus_reachable(gpu.hostname, 9400):
|
push actions check-gpu-monitor-service()
|
||||||
push actions apply-exporter-restart(gpu)
|
|
||||||
|
|
||||||
-- Rule 8: Predictive thermal
|
-- Rule 8: Predictive thermal
|
||||||
for gpu in fleet.gpus:
|
for gpu in fleet.gpus:
|
||||||
@@ -212,12 +235,12 @@ call update-gpu-health
|
|||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
"run_id": "gpu-self-heal-20260712-001",
|
"run_id": "gpu-self-heal-20260718-001",
|
||||||
"timestamp": "2026-07-12T16:00:00Z",
|
"timestamp": "2026-07-18T08:00:00Z",
|
||||||
"gpu": "ct8-rtx3090",
|
"gpu": "ct8-rtx3090",
|
||||||
"issue": "thermal-critical",
|
"issue": "thermal-critical",
|
||||||
"detected": { "temp_c": 87, "duration_s": 180 },
|
"detected": { "temp_c": 87, "duration_s": 180 },
|
||||||
"action": "set-fan-100pct",
|
"action": "load-shedding",
|
||||||
"result": "resolved",
|
"result": "resolved",
|
||||||
"verification": { "temp_c": 76, "after_s": 300 },
|
"verification": { "temp_c": 76, "after_s": 300 },
|
||||||
"escalated": false
|
"escalated": false
|
||||||
@@ -236,52 +259,53 @@ Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail.
|
|||||||
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
|
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
|
||||||
- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
|
- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
|
||||||
|
|
||||||
### 3. Prometheus/Grafana Integration
|
### 3. Weekly Benchmark Report
|
||||||
- GPU self-heal actions exposed as Prometheus counter metrics
|
|
||||||
- Dashboard panel: "GPU Interventions (24h)" showing count/type/result
|
|
||||||
|
|
||||||
### 4. Weekly Benchmark Report
|
|
||||||
- Per-GPU tok/s trend over 7 days
|
- Per-GPU tok/s trend over 7 days
|
||||||
- Regression alerts if any GPU degrades >10% week-over-week
|
- Regression alerts if any GPU degrades >10% week-over-week
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Design Decisions (Grilled & Confirmed — 2026-07-12)
|
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
|
||||||
|
|
||||||
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
|
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
|
||||||
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
|
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
|
||||||
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
|
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
|
||||||
4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix).
|
4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
|
||||||
5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU.
|
5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
|
||||||
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana.
|
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
|
||||||
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
|
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
|
||||||
8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.
|
8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
|
||||||
|
|
||||||
## Lessons Learned (2026-07-12)
|
## Lessons Learned (2026-07-12, Updated 2026-07-18)
|
||||||
|
|
||||||
### L1: API Key Standardization Is Critical
|
### L1: API Key Standardization Is Critical
|
||||||
- All GPU llama-servers MUST use the same api-key as the LiteLLM config.
|
- All GPU llama-servers MUST use the same api-key as the LiteLLM config.
|
||||||
- RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`.
|
- RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`.
|
||||||
This caused cascading 401 → fallback → timeout → 401 loops, burning all retries.
|
This caused cascading 401 → fallback → timeout → 401 loops.
|
||||||
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config.
|
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config (`not-needed` for direct routing).
|
||||||
|
|
||||||
### L2: Fallback Chain Cascading Failures
|
### L2: Fallback Chain Cascading Failures
|
||||||
- When one model returns 401 (auth) and another is slow (timeout), the fallback
|
- When one model returns 401 (auth) and another is slow (timeout), the fallback
|
||||||
chain creates an infinite loop: gemma 401 → qwen timeout → gemma 401 → ...
|
chain creates an infinite loop.
|
||||||
- **Rule**: If a model returns 401 (auth error), do NOT fall back to it again.
|
- **Rule**: If a model returns 401 (auth error), do NOT fall back to it again.
|
||||||
Mark it as permanently failed for this request.
|
Mark it as permanently failed for this request.
|
||||||
|
|
||||||
### L3: Verify Running State, Not Docs
|
### L3: Verify Running State, Not Docs
|
||||||
- RTX 3090 was documented at 128K context. Actually running at 256K.
|
- RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18).
|
||||||
- Parallel count wrong (docs said 2, actual is 1 on RTX 3090).
|
- Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models).
|
||||||
- **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts.
|
- **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts.
|
||||||
|
|
||||||
### L4: Infisical Is Not Always Available
|
### L4: Infisical Is Not Always Available
|
||||||
- Tanko's Infisical service token was 404 — gateway ran without API key for hours.
|
- Keep a local `.env` fallback for `LITELLM_API_KEY`.
|
||||||
- **Rule**: Always keep a local `.env` fallback for `LITELLM_API_KEY`.
|
- **Rule**: Always verify credential source is reachable before relying on it.
|
||||||
- Contract hermes-config-template Rule 3 updated.
|
|
||||||
|
|
||||||
### L5: Zulip Event Queue Can Silently Die
|
### L5: GPU Monitor Response Size Can Cause Self-Heal Crash
|
||||||
- Mumuni's queue accumulated 41 errors/reconnects then stopped polling.
|
- gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response.
|
||||||
Gateway was running but ignoring all messages.
|
- Root cause: router poll returns accumulated data → cache balloons.
|
||||||
- **Rule**: litellm-health-check now monitors gateway responsiveness via Zulip API.
|
- **Rule**: Self-heal must enforce a read timeout AND max response size on every poll.
|
||||||
|
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
|
||||||
|
|
||||||
|
### L6: Stable Aliases Replace Model Names
|
||||||
|
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
|
||||||
|
- Self-heal must use aliases for reporting and alerting, not model-specific names.
|
||||||
|
- **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.
|
||||||
|
|||||||
Reference in New Issue
Block a user