docs: lessons learned from 2026-07-12 session
- gpu-fleet: Corrected architecture (direct GPU, no router in path). Updated context values (RTX 3090=256K, not 128K). Added api-key standardization requirement. - gpu-self-heal: Added Lessons Learned section with 5 critical findings: L1: API key standardization (RTX 5070 sk-loc...5678 vs not-needed) L2: Fallback chain cascading failure loop detection L3: Verify running state, not documentation L4: Infisical fallback requirement (.env must have uncommented key) L5: Zulip event queue can silently die after ~40 reconnects - litellm-self-heal: Updated status manual-only→deployed, cron schedule - litellm-api-keys: Added Infisical token expiry warning + .env fallback - hermes-config-template: Rule 3 updated with .env fallback requirement
This commit is contained in:
@@ -256,3 +256,32 @@ Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail.
|
||||
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana.
|
||||
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
|
||||
8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.
|
||||
|
||||
## Lessons Learned (2026-07-12)
|
||||
|
||||
### L1: API Key Standardization Is Critical
|
||||
- All GPU llama-servers MUST use the same api-key as the LiteLLM config.
|
||||
- RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`.
|
||||
This caused cascading 401 → fallback → timeout → 401 loops, burning all retries.
|
||||
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config.
|
||||
|
||||
### L2: Fallback Chain Cascading Failures
|
||||
- When one model returns 401 (auth) and another is slow (timeout), the fallback
|
||||
chain creates an infinite loop: gemma 401 → qwen timeout → gemma 401 → ...
|
||||
- **Rule**: If a model returns 401 (auth error), do NOT fall back to it again.
|
||||
Mark it as permanently failed for this request.
|
||||
|
||||
### L3: Verify Running State, Not Docs
|
||||
- RTX 3090 was documented at 128K context. Actually running at 256K.
|
||||
- Parallel count wrong (docs said 2, actual is 1 on RTX 3090).
|
||||
- **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts.
|
||||
|
||||
### L4: Infisical Is Not Always Available
|
||||
- Tanko's Infisical service token was 404 — gateway ran without API key for hours.
|
||||
- **Rule**: Always keep a local `.env` fallback for `LITELLM_API_KEY`.
|
||||
- Contract hermes-config-template Rule 3 updated.
|
||||
|
||||
### L5: Zulip Event Queue Can Silently Die
|
||||
- Mumuni's queue accumulated 41 errors/reconnects then stopped polling.
|
||||
Gateway was running but ignoring all messages.
|
||||
- **Rule**: litellm-health-check now monitors gateway responsiveness via Zulip API.
|
||||
|
||||
Reference in New Issue
Block a user