fix(audit): stop requiring the retired gpu-light alias; derive model fields; sweep retired names #80
@@ -64,7 +64,7 @@ Key notes:
|
||||
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
|
||||
- **Fix**:
|
||||
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
|
||||
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
|
||||
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gpu-vision → gpu-dense, gpu-dense → gpu-vision)
|
||||
3. If all GPUs hot, alert about cooling infrastructure
|
||||
- **Verify**: Temp drops below 80°C within 5 minutes
|
||||
- **Escalate after**: 3 verification failures → Zulip alert
|
||||
|
||||
@@ -253,8 +253,8 @@ The following MUST be identical across ALL profiles:
|
||||
|
||||
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
|
||||
- Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized)
|
||||
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
|
||||
- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
|
||||
- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 128K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it)
|
||||
- **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it
|
||||
(do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml`
|
||||
is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
|
||||
- **OPERATIONAL DECISION (2026-07-23): Use `syslog-auto` for compression across all agents.**
|
||||
@@ -265,7 +265,7 @@ The following MUST be identical across ALL profiles:
|
||||
- All auxiliary services MUST use identical routing:
|
||||
- `base_url: http://192.168.68.116/litellm/v1` (Rule 5, 2026-08-09: canonical authenticated; `/v1` also OK)
|
||||
- `api_key_env: LITELLM_API_KEY`
|
||||
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
|
||||
- **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above)
|
||||
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
|
||||
(64GB UMA, 128K context) — the designated compression GPU. This frees the
|
||||
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
|
||||
|
||||
@@ -23,7 +23,7 @@ management, and prompt caching — without sacrificing agent capability.
|
||||
.123, any others on .129/.122) including compression, model, context_window,
|
||||
prompt_caching, memory settings
|
||||
- `gpu-health`: health check response from all 3 GPU backends (strix-moe .15:8080,
|
||||
qwen .8:8080, gemma .110:8080)
|
||||
gpu-dense .8:8080, gpu-vision .110:8080)
|
||||
|
||||
### Maintains
|
||||
|
||||
@@ -56,8 +56,8 @@ duration.
|
||||
**Context is the root cause.** Every ~46K prompt token costs ~87s of
|
||||
prefill time at 532 tok/s. Fix context first, routing second.
|
||||
|
||||
- **Route by task**: qwen for code/standard queries; gemma for
|
||||
compression/auxiliary; strix-moe for compression tasks.
|
||||
- **Route by task**: gpu-dense for code/standard queries; gpu-vision for
|
||||
vision/web-auxiliary; syslog-auto for compression.
|
||||
- **Compress aggressively**: threshold at 40% (not 65%) — a 128K window should
|
||||
compact at 51K, not 85K. Target 15% tail (not 30%).
|
||||
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
|
||||
|
||||
Reference in New Issue
Block a user