Compare commits

...
3 Commits
Author SHA1 Message Date
jerome a74229ee74 Merge branch 'master' into fix/fleet-config-issues-20260718
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
2026-07-18 18:06:19 +00:00
jerome 14d27a09b5 Merge pull request 'gpu-self-heal: refresh to current fleet baseline and topology' (#21) from feat/gpu-self-heal-refresh-20260718 into master
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
Reviewed-on: #21
2026-07-18 08:36:37 +00:00
root bddbb22f03 gpu-self-heal: refresh to current fleet baseline and topology
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
- Synced model assignments to 2026-07-17 swaps (ThinkingCap, HauhauCS QAT, Genesis Hermes V3)
- Added stable role-based aliases from gpu-fleet (gpu-dense, gpu-light, strix-moe)
- Updated benchmark baselines to live values (74.9/169.6/62.9 tok/s)
- Replaced router (port 9000) references with LiteLLM + direct routing
- Replaced Prometheus exporter rule with sidecar health probe
- Updated VRAM thresholds to match operational data (300/300/200 MB/h)
- Added response size limit (1MB) to prevent OOM crashes
- Added Lessons L5 (response size crash) and L6 (stable aliases)
- Removed deprecated Rules 11-12 (router-specific distribution balance)
2026-07-18 08:07:03 +00:00
+93 -69
View File
@@ -6,10 +6,15 @@ description: >
benchmarks, and predicts failures before they happen. Extends gpu-monitor benchmarks, and predicts failures before they happen. Extends gpu-monitor
(v2.1.0) with active remediation rules, Prometheus metrics consumption, (v2.1.0) with active remediation rules, Prometheus metrics consumption,
VRAM trend analysis, and predictive alerting. VRAM trend analysis, and predictive alerting.
UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps.
Router (port 9000) references replaced with direct GPU routing.
Benchmark baselines refreshed to live values.
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
agent: abiba agent: abiba
depends_on: depends_on:
- gpu-monitor.prose.md (live data source on .24:9100) - gpu-monitor.prose.md (live data source on .24:9100)
- gpu-fleet.prose.md (source of truth for topology) - gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
--- ---
## Maintains ## Maintains
@@ -22,8 +27,8 @@ depends_on:
## Requires ## Requires
- gpu-monitor:function — Live fleet data from .24:9100/gpu-data - gpu-monitor:function — Live fleet data from localhost:9100/gpu-data
- Prometheus exporters on all 3 GPUs (:9400/metrics) - Direct sidecar probe access to all GPU hosts (:8080/health)
- SSH access to GPU hosts for restart operations - SSH access to GPU hosts for restart operations
## Continuity ## Continuity
@@ -35,21 +40,37 @@ depends_on:
--- ---
## Current Fleet Baseline (2026-07-18)
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|-------|-----|------|-------|------|-----|-------|------|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Compression, summarization, long docs |
Key notes:
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
## Remediation Rules ## Remediation Rules
### Rule 1: GPU Temperature Critical (>85°C for >2 min) ### Rule 1: GPU Temperature Critical (>85°C for >2 min)
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls - **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
- **Fix**: - **Fix**:
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control) 1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains 2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
3. If all GPUs hot, alert about cooling infrastructure 3. If all GPUs hot, alert about cooling infrastructure
- **Verify**: Temp drops below 80°C within 5 minutes - **Verify**: Temp drops below 80°C within 5 minutes
- **Escalate after**: 3 verification failures → Zulip alert - **Escalate after**: 3 verification failures → Zulip alert
### Rule 2: VRAM Leak Detection (tiered by GPU capacity) ### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
- **Detect**: VRAM growing at sustained rate over 6+ hour window - **Detect**: VRAM growing at sustained rate over 6+ hour window
- RTX 3090 (24GB): ≥100MB/hour - RTX 3090 (24GB): ≥300MB/hour
- RTX 5070 (12GB): ≥50MB/hour - RTX 5070 (12GB): ≥300MB/hour
- Strix Halo (64GB UMA): ≥200MB/hour - Strix Halo (64GB UMA): ≥200MB/hour
- **Fix**: - **Fix**:
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux) 1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
@@ -69,6 +90,9 @@ depends_on:
### Rule 4: Benchmark Regression (>20% drop) ### Rule 4: Benchmark Regression (>20% drop)
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks - **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
- RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s
- RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s
- Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s
- **Fix**: - **Fix**:
1. Check GPU utilization — if >90%, other process is competing 1. Check GPU utilization — if >90%, other process is competing
2. Check power limit — if throttled, restore to max 2. Check power limit — if throttled, restore to max
@@ -78,32 +102,31 @@ depends_on:
### Rule 5: Circuit Breaker Stuck Open ### Rule 5: Circuit Breaker Stuck Open
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy - **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
- **Fix**: - **Fix**:
1. Verify GPU /health returns 200 1. Verify GPU /health returns 200 on direct port (:8080)
2. If GPU healthy, send 1 test inference 2. If GPU healthy, alert but do NOT reset via router API (deprecated)
3. If test succeeds → reset circuit breaker via router API 3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset 4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
5. Max 1 auto-reset per GPU per hour - **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s
- **Verify**: CB closes, inference succeeds, CB stays closed for 60s+ - **Escalate after**: LiteLLM restart doesn't clear → human investigation
- **Escalate after**: CB won't close after reset → router issue
### Rule 6: Strix Halo Unreachable ### Rule 6: Strix Halo Unreachable
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15) - **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
- **Fix**: - **Fix**:
1. SSH to .15 → check llama-server process 1. SSH to .15 → check llama-server process
2. Restart llama-server if not running 2. Restart llama-server if not running
3. Verify through both direct probe AND router 3. Verify through both direct probe AND LiteLLM health
- **Verify**: Direct health probe returns 200, router reports Strix healthy - **Verify**: Direct health probe returns 200, LiteLLM reports model healthy
- **Escalate**: If host .15 itself is unreachable → infrastructure alert - **Escalate**: If host .15 itself is unreachable → infrastructure alert
### Rule 7: Prometheus Exporter Down ### Rule 7: GPU Data Source Unreachable (replaces old Prometheus rule)
- **Detect**: Any GPU :9400/metrics unreachable for >2 polls - **Detect**: gpu-monitor endpoint (localhost:9100/gpu-data) or sidecar port (:8080) on any GPU unreachable for >2 polls
- **Fix**: - **Fix**:
1. SSH to GPU host → check prometheus-exporter process 1. If gpu-monitor is down: restart systemd service `gpu-monitor.service` on this host
2. Restart exporter if dead 2. If sidecar is down: SSH to GPU host → check llama-server process → restart systemd service
3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes 3. Fall back to direct nvidia-smi/rocm-smi probe via SSH if all API paths fail
4. If exporter is running but unreachable → check firewall/host networking - **Verify**: gpu-monitor returns healthy + all sidecars reachable
- **Verify**: :9400/metrics returns 200
- **Escalate after**: 3 failed restarts → networking issue - **Escalate after**: 3 failed restarts → networking issue
### Rule 8: Predictive Thermal Warning (two-tier) ### Rule 8: Predictive Thermal Warning (two-tier)
@@ -117,29 +140,31 @@ depends_on:
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure - **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
### Rule 9: Context Window Optimization ### Rule 9: Context Window Optimization
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context - **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
- RTX 3090 (128K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline - RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
- RTX 5070 (128K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role - RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
- Strix Halo (128K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline - Strix Halo (128K ctx, Genesis Hermes V3): baseline 70.5 tok/s — currently at 62.9 (89%)
- **Fix**: - **Fix**:
- If tok/s > baseline → context has headroom, consider increasing - If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest - If tok/s < 90% baseline → reduce context by 25% and retest
- If tok/s within 10% of baseline → optimal, no change - If tok/s within 10% of baseline → optimal, no change
- Strix Halo at 89% of baseline → MONITOR but do not reduce yet (recent model swap may still be settling)
- **Verify**: Re-benchmark after context change, confirm within 10% of target - **Verify**: Re-benchmark after context change, confirm within 10% of target
- **Escalate**: If context can't be adjusted without significant perf loss - **Escalate**: If context can't be adjusted without significant perf loss
### Rule 10: Workload Distribution Optimization ### Rule 10: Workload Distribution Optimization (updated 2026-07-18)
- **Detect**: GPU roles misaligned with hardware capabilities - **Detect**: GPU roles misaligned with hardware capabilities
- **Target distribution**: - **Target distribution**:
- RTX 3090 (24GB, 128K, 75 tok/s) → Heavy reasoning, code gen, long conversations - RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
- RTX 5070 (12GB, 128K, 76 tok/s) → Vision/image, web search, quick lightweight tasks - RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
- Strix Halo (64GB, 128K, 72 tok/s) → Context compression, summarization, long docs - Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Context compression, summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
- **Fix**: - **Fix**:
- Alert if any GPU is handling workload outside its designated role - Alert if any GPU is handling workload outside its designated role
- Recommend Hermes agent profile updates to match workload to GPU - Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe)
- Track per-GPU request distribution via LiteLLM spend logs - Track per-GPU request distribution via LiteLLM spend logs
- **Verify**: Each GPU's request pattern matches its designated role within 24h - **Verify**: Each GPU's request pattern matches its designated role within 24h
- **Escalate**: If role mismatch persists >48h → agent profile audit needed - **Escalate**: If role mismatch persists >48h → agent alias audit needed
--- ---
@@ -148,7 +173,7 @@ depends_on:
```prose ```prose
-- Phase 1: Fetch live GPU data -- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor let fleet = call gpu-monitor
endpoint: "http://192.168.68.24:9100/gpu-data" endpoint: "http://localhost:9100/gpu-data"
-- Phase 2: Evaluate each GPU against remediation rules -- Phase 2: Evaluate each GPU against remediation rules
let actions = [] let actions = []
@@ -164,27 +189,25 @@ for gpu in fleet.gpus:
-- Rule 4: Benchmark regression -- Rule 4: Benchmark regression
let bench = fleet.benchmarks[gpu.hostname] let bench = fleet.benchmarks[gpu.hostname]
if bench.current_tok_sec < bench.baseline_tok_sec * 0.8: if bench.current_tok_s < bench.baseline_tok_s * 0.8:
push actions apply-benchmark-fix(gpu, bench) push actions apply-benchmark-fix(gpu, bench)
-- Rule 3: Model stuck -- Rule 3: Model stuck
for model in fleet.router.available_models: for model in fleet.summary.available_models:
if model.consecutive_timeouts >= 3: if model.consecutive_timeouts >= 3:
push actions apply-model-restart(model) push actions apply-model-restart(model)
-- Rule 5: Circuit breaker -- Rule 5: Circuit breaker check via LiteLLM (router deprecated)
for cb in fleet.router.circuit_breaker: if fleet.summary.circuit_breakers_open > 0:
if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu): push actions check-litellm-circuit-breakers()
push actions apply-cb-reset(cb)
-- Rule 6: Strix Halo -- Rule 6: Strix Halo
if not fleet.strix.running and pingable("192.168.68.15"): if not fleet.strix.running and pingable("192.168.68.15"):
push actions apply-strix-restart() push actions apply-strix-restart()
-- Rule 7: Prometheus exporters -- Rule 7: GPU data source
for gpu in fleet.gpus: if not fleet.gpus or len(fleet.gpus) < 2:
if not prometheus_reachable(gpu.hostname, 9400): push actions check-gpu-monitor-service()
push actions apply-exporter-restart(gpu)
-- Rule 8: Predictive thermal -- Rule 8: Predictive thermal
for gpu in fleet.gpus: for gpu in fleet.gpus:
@@ -212,12 +235,12 @@ call update-gpu-health
```json ```json
{ {
"run_id": "gpu-self-heal-20260712-001", "run_id": "gpu-self-heal-20260718-001",
"timestamp": "2026-07-12T16:00:00Z", "timestamp": "2026-07-18T08:00:00Z",
"gpu": "ct8-rtx3090", "gpu": "ct8-rtx3090",
"issue": "thermal-critical", "issue": "thermal-critical",
"detected": { "temp_c": 87, "duration_s": 180 }, "detected": { "temp_c": 87, "duration_s": 180 },
"action": "set-fan-100pct", "action": "load-shedding",
"result": "resolved", "result": "resolved",
"verification": { "temp_c": 76, "after_s": 300 }, "verification": { "temp_c": 76, "after_s": 300 },
"escalated": false "escalated": false
@@ -236,52 +259,53 @@ Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail.
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention" - `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
- Every 100th clean cycle → "✅ GPU Fleet: All Clear" - Every 100th clean cycle → "✅ GPU Fleet: All Clear"
### 3. Prometheus/Grafana Integration ### 3. Weekly Benchmark Report
- GPU self-heal actions exposed as Prometheus counter metrics
- Dashboard panel: "GPU Interventions (24h)" showing count/type/result
### 4. Weekly Benchmark Report
- Per-GPU tok/s trend over 7 days - Per-GPU tok/s trend over 7 days
- Regression alerts if any GPU degrades >10% week-over-week - Regression alerts if any GPU degrades >10% week-over-week
--- ---
## Design Decisions (Grilled & Confirmed 2026-07-12) ## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect). 1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request. 2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart. 3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix). 4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU. 5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana. 6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising. 7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down. 8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
## Lessons Learned (2026-07-12) ## Lessons Learned (2026-07-12, Updated 2026-07-18)
### L1: API Key Standardization Is Critical ### L1: API Key Standardization Is Critical
- All GPU llama-servers MUST use the same api-key as the LiteLLM config. - All GPU llama-servers MUST use the same api-key as the LiteLLM config.
- RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`. - RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`.
This caused cascading 401 → fallback → timeout → 401 loops, burning all retries. This caused cascading 401 → fallback → timeout → 401 loops.
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config. - **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config (`not-needed` for direct routing).
### L2: Fallback Chain Cascading Failures ### L2: Fallback Chain Cascading Failures
- When one model returns 401 (auth) and another is slow (timeout), the fallback - When one model returns 401 (auth) and another is slow (timeout), the fallback
chain creates an infinite loop: gemma 401 → qwen timeout → gemma 401 → ... chain creates an infinite loop.
- **Rule**: If a model returns 401 (auth error), do NOT fall back to it again. - **Rule**: If a model returns 401 (auth error), do NOT fall back to it again.
Mark it as permanently failed for this request. Mark it as permanently failed for this request.
### L3: Verify Running State, Not Docs ### L3: Verify Running State, Not Docs
- RTX 3090 was documented at 128K context. Actually running at 256K. - RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18).
- Parallel count wrong (docs said 2, actual is 1 on RTX 3090). - Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models).
- **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts. - **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts.
### L4: Infisical Is Not Always Available ### L4: Infisical Is Not Always Available
- Tanko's Infisical service token was 404 — gateway ran without API key for hours. - Keep a local `.env` fallback for `LITELLM_API_KEY`.
- **Rule**: Always keep a local `.env` fallback for `LITELLM_API_KEY`. - **Rule**: Always verify credential source is reachable before relying on it.
- Contract hermes-config-template Rule 3 updated.
### L5: Zulip Event Queue Can Silently Die ### L5: GPU Monitor Response Size Can Cause Self-Heal Crash
- Mumuni's queue accumulated 41 errors/reconnects then stopped polling. - gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response.
Gateway was running but ignoring all messages. - Root cause: router poll returns accumulated data → cache balloons.
- **Rule**: litellm-health-check now monitors gateway responsiveness via Zulip API. - **Rule**: Self-heal must enforce a read timeout AND max response size on every poll.
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
### L6: Stable Aliases Replace Model Names
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
- Self-heal must use aliases for reporting and alerting, not model-specific names.
- **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.