From 1f1b47f59df2da7a6d89185271edf29ebf6dd5a1 Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 16:51:21 +0000 Subject: [PATCH] no-mistakes(document): Sweep residual gemma aliases; align compression rule contradiction --- gpu-self-heal.prose.md | 2 +- hermes-config-template.prose.md | 6 +++--- inference-optimization.prose.md | 6 +++--- 3 files changed, 7 insertions(+), 7 deletions(-) diff --git a/gpu-self-heal.prose.md b/gpu-self-heal.prose.md index cdc6264..32019ca 100644 --- a/gpu-self-heal.prose.md +++ b/gpu-self-heal.prose.md @@ -64,7 +64,7 @@ Key notes: - **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls - **Fix**: 1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control) - 2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma) + 2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gpu-vision → gpu-dense, gpu-dense → gpu-vision) 3. If all GPUs hot, alert about cooling infrastructure - **Verify**: Temp drops below 80°C within 5 minutes - **Escalate after**: 3 verification failures → Zulip alert diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index 485f2ee..ea700c8 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -253,8 +253,8 @@ The following MUST be identical across ALL profiles: ### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16) - Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized) -- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized) -- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b` +- Compression uses `syslog-auto` — the Strix Halo weighted pool (64GB, 128K ctx, compression-optimized); do NOT pin `compression.model` to `strix-moe` (audit Rule 7 rejects it) +- **`ornith-1.0-35b` is NOT a valid compression model name** — LiteLLM does not serve it (do not restate the served model list here — CT 116 `/opt/inference-harness/litellm_config.yaml` is the single source of truth for models, aliases, weights and fallbacks). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls. - **OPERATIONAL DECISION (2026-07-23): Use `syslog-auto` for compression across all agents.** @@ -265,7 +265,7 @@ The following MUST be identical across ALL profiles: - All auxiliary services MUST use identical routing: - `base_url: http://192.168.68.116/litellm/v1` (Rule 5, 2026-08-09: canonical authenticated; `/v1` also OK) - `api_key_env: LITELLM_API_KEY` -- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably +- **Do NOT use `syslog-auto` for `vision`/`web_extract`** — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above) - **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo (64GB UMA, 128K context) — the designated compression GPU. This frees the RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning. diff --git a/inference-optimization.prose.md b/inference-optimization.prose.md index beb18a2..69a0d96 100644 --- a/inference-optimization.prose.md +++ b/inference-optimization.prose.md @@ -23,7 +23,7 @@ management, and prompt caching — without sacrificing agent capability. .123, any others on .129/.122) including compression, model, context_window, prompt_caching, memory settings - `gpu-health`: health check response from all 3 GPU backends (strix-moe .15:8080, - qwen .8:8080, gemma .110:8080) + gpu-dense .8:8080, gpu-vision .110:8080) ### Maintains @@ -56,8 +56,8 @@ duration. **Context is the root cause.** Every ~46K prompt token costs ~87s of prefill time at 532 tok/s. Fix context first, routing second. -- **Route by task**: qwen for code/standard queries; gemma for - compression/auxiliary; strix-moe for compression tasks. +- **Route by task**: gpu-dense for code/standard queries; gpu-vision for + vision/web-auxiliary; syslog-auto for compression. - **Compress aggressively**: threshold at 40% (not 65%) — a 128K window should compact at 51K, not 85K. Target 15% tail (not 30%). - **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these