From b9149bce47cd4748966cf136a1e42db603bcaaa0 Mon Sep 17 00:00:00 2001 From: root Date: Fri, 17 Jul 2026 10:59:45 +0000 Subject: [PATCH 1/3] =?UTF-8?q?feat:=20GPU=20context=20256K=E2=86=92128K?= =?UTF-8?q?=20fleet-wide=20+=20Genesis=20Hermes=20V3=20on=20Strix=20Halo?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit GPU Changes: - All 3 GPUs reduced from 256K (-c 262144) to 128K (-c 131072) for stability - Observed instability near 100K at 256K — 128K is the stable ceiling - VRAM improved: RTX 3090 ~70% (was 90%), RTX 5070 ~65% (was 88%) - Strix Halo swapped to LuffyTheFox/Genesis Hermes V3 APEX - Hermes agent fine-tune, tensor repair (3 SSM layers, 76% W1 improvement) - Uncensored (0/465 refusals), multimodal (mmproj F16) - Speed: 65 tok/s gen, 140 tok/s prompt - Alias strix-moe maintained Agent Updates: - Mumuni: max_context_window 262144→131072, already aligned on strix-moe/0.65 - Tanko: max_context_window 262144→131072 - Koonimo: max_context_window + context_length 262144→131072 - CT114 SSH access confirmed (was 'Zulip only') LiteLLM (CT116): - Updated backend model references qwen3.6-35B-udq4→strix-moe - Removed stale ornith-1.0-35b from model_cost - Fallback chains updated Contracts Updated: - gpu-fleet.prose.md: topology, VRAM, benchmarks, config lines, model assignments - gpu-self-heal.prose.md: Rule 9/10 context targets - hermes-config-template.prose.md: template values, Rules 7-9, compression thresholds - inference-optimization.prose.md: added to repo, 128K recommendation Compression: 0.65 fires at ~85K (~43K headroom before 128K ceiling) For >128K workloads: route to external providers (deepseek) --- gpu-fleet.prose.md | 50 ++++++++-------- gpu-self-heal.prose.md | 12 ++-- hermes-config-template.prose.md | 41 +++++++------ inference-optimization.prose.md | 100 ++++++++++++++++++++++++++++++++ 4 files changed, 153 insertions(+), 50 deletions(-) create mode 100644 inference-optimization.prose.md diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 8946c9f..7e938d9 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -9,8 +9,12 @@ description: > gpu-dense, gpu-light. These never change — only the underlying model does. Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster). - RTX 5070 context: 131K → 256K. VRAM: 88% (10.8/12.2GB). - Compression timeout: 300s (was 120s). Mumuni context: 128K (was 256K). + UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability. + Strix Halo model swapped to Genesis Hermes V3 APEX (LuffyTheFox, 24GB, uncensored, + Hermes agent fine-tune, tensor repair, multimodal with mmproj). + Instability observed near 100K at 256K. 128K is the stable ceiling. + For larger context needs → fall back to external providers (deepseek). + VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. agent: abiba triggers: - on model add/remove @@ -67,7 +71,7 @@ triggers: │ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │ │ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │ │ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │ -│ 256K ctx │ │ 256K ctx │ │ 256K ctx │ │ Watchdog │ +│ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │ │ qwen3.6 │ │ gemma-4-12b │ │ qwen3.6 │ │ Prometheus │ │ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │ │ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │ @@ -93,9 +97,9 @@ but are deprecated for agent configs. Only the stable aliases survive model swap | Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | |-------|-----|------|------|-----|----------|----------|-------------|--------| -| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | 22.2/24.6GB (90%) | **256K** 🚀 | turbo4 | 2 | default | ✅ 63 tok/s | -| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 10.0/12.2GB (82%) | 256K | q4_0 | 2 | 2048/1024 | ✅ healthy | -| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~9GB/64GB | 256K | q8_0 | 2 | 2048/512 | ✅ healthy | +| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s | +| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy | +| Genesis Hermes V3 APEX | Strix Halo Vulkan | .15 (amdpve) | ~10GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s | ## Routing Configuration (LiteLLM — July 2026) @@ -104,7 +108,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap | Model | GPU | Weight | RPM Cap | Timeout | |-------|-----|--------|---------|---------| | qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | -| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** | +| Genesis Hermes V3 APEX | Strix Halo (.15:8080) | **0.30** | 60 | **300s** | | gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** | Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path. @@ -113,7 +117,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. | Model | RPM Cap | Notes | |-------|---------|-------| -| qwen3.6-35B-udq4 | 40 | Tight cap — prevents Strix overload | +| strix-moe (Hermes V3) | 40 | Tight cap — prevents Strix overload | | qwen3.6-27B-code | 500 | High cap — primary workhorse | | gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s | @@ -128,11 +132,11 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. ### Fallback Chains - gemma → qwen - qwen → gemma -- qwen3.6-35B-udq4 → qwen → gemma +- strix-moe → qwen → gemma - syslog-auto → qwen → gemma → qwen3.6-35B-udq4 ### Why Strix Halo RPM Is Capped -- Direct (qwen3.6-35B-udq4): 40 RPM (tight) — Strix Halo is shared with compression tasks +- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks - Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously - Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C @@ -243,12 +247,12 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions. - **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first. Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster(). - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. -- **VRAM (2026-07-15)**: RTX 3090 at 22.2/24.6GB (90%) with **256K context** (corrected from 131K). RTX 5070 at 10.8/12.2GB (88%) with 256K context + MTP. Strix Halo at ~9GB/64GB. +- **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB. - **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2). -- **RTX 3090 config**: `-c 262144 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context corrected to 256K (2026-07-15). VRAM: 90%. Service: `/home/llmuser/llama-wrapper.sh`. -- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 256K context. Gen speed: 122 tok/s (was 70). VRAM: 10.8/12.2GB (88%). No draft model pre-upgrade due to VRAM constraints. Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 262144`. +- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`. +- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`. - **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. -- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080 (was `ornith-server.service`), model changed to `unsloth/Qwen3.6-35B-A3B-MTP-GGUF` (UD-Q4_K_M), alias `qwen3.6-35B-udq4`, 256K context, flash-attn + q8 KV. MTP support enabled for 1.4-2.2x faster inference. +- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF` (APEX quant), alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). Hermes agent fine-tune, tensor repair (SSM layers fixed via SVD), uncensored (0/465 refusals). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. - **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116. @@ -262,12 +266,12 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions. | GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context | |-----|-------|-----------|--------------|----------|---------| -| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **256K** | -| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **256K** | -| Strix Halo (.15) | qwen3.6-35B-udq4 | **71** | — | — | **256K** | +| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** | +| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** | +| Strix Halo (.15) | Genesis Hermes V3 APEX | **65** | 140 | — | **128K** | -Benchmarks from 2026-07-15 verification run. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s. -All 3 GPUs now at 256K context (2026-07-15). +Benchmarks from 2026-07-17. Strix Halo swapped to Genesis Hermes V3 APEX (LuffyTheFox). RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s. +All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability). Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes. Degradation alerts fire at 30% (warning) and 50% (critical) below baseline. @@ -288,9 +292,9 @@ All agent configs MUST use stable role-based aliases, never model-specific names When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched. ### Context Windows -- RTX 3090: **256K** (was 131K, bumped 2026-07-15) | RTX 5070: **256K** (up from 131K) | Strix Halo: **256K** -- **Mumuni compression context**: 128K (down from 256K) — ensures compression model doesn't timeout -- Compression threshold 0.65: fires at ~85K for 128K context window +- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K** +- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek) +- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling) - Mumuni compression model alias: `strix-moe` with 300s timeout ### Mumuni Agent Profile @@ -306,7 +310,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i | `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) | | `aux.web_extract.model` | `gpu-light` | Web extraction | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | -| `context.max_context_window` | 262144 (256K) | Fixed 2026-07-16 (was 131072 — caused premature compression, WAL #1300) | +| `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling | | `compression.threshold` | 0.65 | Triggers at ~85K | | `compression.target_ratio` | 0.3 | Compresses to ~38K | | `compression.protect_last_n` | 40 | Preserves last 40 messages | diff --git a/gpu-self-heal.prose.md b/gpu-self-heal.prose.md index 9209219..5388a40 100644 --- a/gpu-self-heal.prose.md +++ b/gpu-self-heal.prose.md @@ -118,9 +118,9 @@ depends_on: ### Rule 9: Context Window Optimization - **Detect**: Benchmark tok/s vs baseline for each GPU at current context - - RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline - - RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role - - Strix Halo (256K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline + - RTX 3090 (128K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline + - RTX 5070 (128K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role + - Strix Halo (128K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline - **Fix**: - If tok/s > baseline → context has headroom, consider increasing - If tok/s < 90% baseline → reduce context by 25% and retest @@ -131,9 +131,9 @@ depends_on: ### Rule 10: Workload Distribution Optimization - **Detect**: GPU roles misaligned with hardware capabilities - **Target distribution**: - - RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations - - RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks - - Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs + - RTX 3090 (24GB, 128K, 75 tok/s) → Heavy reasoning, code gen, long conversations + - RTX 5070 (12GB, 128K, 76 tok/s) → Vision/image, web search, quick lightweight tasks + - Strix Halo (64GB, 128K, 72 tok/s) → Context compression, summarization, long docs - **Fix**: - Alert if any GPU is handling workload outside its designated role - Recommend Hermes agent profile updates to match workload to GPU diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index cd95030..693026a 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -6,11 +6,10 @@ description: > Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, RA-H OS MCP) while keeping agent-specific API keys and model choices. UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`, - which LiteLLM does not serve). All 3 GPUs verified 256K (RTX 5070 bumped 131K→256K on Jul 15). + which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability). Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the 2026-07-16 Mumuni root-cause investigation (WAL #1300). - UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context - verified at 256K. Infisical .env fallback required (Rule 3/13). + UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13). --- ## Maintains @@ -36,7 +35,7 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed | Mumuni | `mumuni` | 192.168.68.123 | root@.123 | 6 profiles ✱ | | Abiba | `abiba-pi` | 192.168.68.24 | local | — | | Koby | `koby` | CT 111 (tdunna) | Zulip | — | -| Koonimo | `koonimo` | CT 113 (baggy) | Zulip | — | +| Koonimo | `koonimo` | CT 114 (baggy) | SSH root | — | | Kagenz0 | `kagenz0-*` | ? | Zulip | — | > CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo). @@ -101,8 +100,8 @@ model: base_url: http://192.168.68.116/v1 api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation - context_length: 262144 # For syslog-auto (all GPUs support 256K). - # Set 131072 if using gemma-4-12b directly (12GB VRAM constraint). + context_length: 131072 # For syslog-auto (all GPUs at 128K for stability). + # Set 65536 if using gemma-4-12b directly (tight VRAM). fallback_providers: provider: deepseek @@ -133,7 +132,7 @@ compression: enabled: true model: strix-moe # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name). provider: harness - max_context_window: 262144 # MUST match actual GPU capacity. All 3 GPUs are 256K (Jul 15). + max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17). threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K target_ratio: 0.30 protect_last_n: 40 @@ -250,7 +249,7 @@ The following MUST be identical across ALL profiles: ### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16) - Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized) -- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 256K ctx, compression-optimized) +- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized) - **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b` (it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls. - All auxiliary services MUST use identical routing: @@ -258,31 +257,31 @@ The following MUST be identical across ALL profiles: - `api_key_env: LITELLM_API_KEY` - **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably - **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo - (64GB UMA, 256K context) — the designated compression GPU. This frees the + (64GB UMA, 128K context) — the designated compression GPU. This frees the RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning. - The `compression:` block's `model` MUST match `auxiliary: compression: model` -- The `compression: max_context_window: 262144` MUST match Strix Halo's actual capacity +- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K) ### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16) -- **RTX 3090 (24GB, 256K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations -- **RTX 5070 (12GB, 256K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (bumped 131K→256K Jul 15; IQ4_NL+MTP, 88% VRAM) -- **Strix Halo (64GB, 256K ctx, strix-moe)**: Context compression, summarization, long docs +- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations +- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K) +- **Strix Halo (64GB, 128K ctx, strix-moe)**: Context compression, summarization, long docs - Agent profiles MUST route auxiliary tasks to the correct GPU: - `auxiliary.vision.model: gemma-4-12b` (RTX 5070) - `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070) - `auxiliary.compression.model: strix-moe` (Strix Halo) - Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing -- For 262K context window: `threshold: 0.65` (fires at ~170K tokens) +- For 128K context window: `threshold: 0.65` (fires at ~85K tokens) - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss -- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer -- `max_context_window: 262144` MUST match the model's actual capacity +- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin +- `max_context_window: 131072` MUST match the model's actual capacity (128K) - See `devops-hermes-compression` skill for full reference -### Rule 9: Compression Threshold for 256K Models -- For 262K context window: `threshold: 0.65` (fires at ~170K tokens) +### Rule 9: Compression Threshold for 128K Models +- For 128K context window: `threshold: 0.65` (fires at ~85K tokens) - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss -- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer -- `max_context_window: 262144` MUST match the model's actual capacity (Strix Halo = 256K) +- Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin +- `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K) - See `devops-hermes-compression` skill for full reference ### Rule 10: Default Model Must Be `syslog-auto` (All Agents) @@ -312,7 +311,7 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee verify ALL FOUR of these against the live config. They are the only root causes found in production: 1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window` - MUST be `262144` (all GPUs are 256K). A value of `131072` causes premature compression at + MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used. ~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml` 2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`, and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/` diff --git a/inference-optimization.prose.md b/inference-optimization.prose.md new file mode 100644 index 0000000..58c8760 --- /dev/null +++ b/inference-optimization.prose.md @@ -0,0 +1,100 @@ +--- +name: inference-optimization +kind: responsibility +id: 067NC6KP02RG60S50M40E30928 +--- + +### Goal + +Syslog inference response times reduced to sub-15s average by optimizing the full +stack: LiteLLM routing weights, GPU model assignments, Hermes agent context +management, and prompt caching — without sacrificing agent capability. + +### Requires + +- `inference-metrics`: current SpendLogs from CT116 LiteLLM Postgres — avg + request_duration_ms, prompt_tokens, completion_tokens, model_group breakdown, + cache_hit rate over the last 3 hours +- `agent-configs`: current config.yaml from each active Hermes agent (Mumuni + .123, any others on .129/.122) including compression, model, context_window, + prompt_caching, memory settings +- `gpu-health`: health check response from all 3 GPU backends (ornith .15:8080, + qwen .8:8080, gemma .110:8080) + +### Maintains + +The optimized inference stack configuration — every change is applied and +verified end-to-end. Postcondition: avg request_duration_ms ≤ 15000 for 90% of +non-ornith traffic; ≤ 30000 for ornith-bound agentic calls. + +#### liteLLM-routing +The syslog-auto routing weights, model-specific timeouts, RPM limits, and +model_list entries on CT116 `/opt/inference-harness/litellm_config.yaml`. + +#### agent-compression +Each Hermes agent's `~/.hermes/config.yaml` compression, context_window, +prompt_caching, and model sections. + +#### prompt-caching +LiteLLM cache configuration and llama.cpp `--cache-prompt` flag on GPU hosts. + +#### verification +End-to-end latency measurements after changes applied — at least 3 test +inference calls per model path measuring ttft (time-to-first-token) and total +duration. + +### Continuity + +- input-driven + +### Strategies + +**Context is the root cause.** Every ~46K prompt token costs ~87s of ornith +prefill time at 532 tok/s. Fix context first, routing second. + +- **Route by task**: ornith for multi-step reasoning only; qwen for code/standard + queries; gemma for compression/auxiliary. Never send simple completion to a + 35B MoE. +- **Compress aggressively**: threshold at 40% (not 65%) — a 256K window should + compact at 102K, not 166K. Target 15% tail (not 30%). +- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these + never change between turns. Single-digit cache hit rate is unacceptable. +- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations. + GPUs reduced from 256K to 128K (2026-07-17). For larger contexts, route to external providers. + +### Shape + +- `self`: analyze metrics, compute optimal configs, apply changes, verify +- `delegates`: + - `apply-liteLLM`: update litellm_config.yaml and reload + - `apply-agent-config`: update hermes config.yaml per agent + - `verify-latency`: run test inference calls and measure response + +### Execution + +```prose +-- Phase 1: Analyze current state (already complete) +-- Phase 2: Apply LiteLLM routing optimization + +call apply-liteLLM-routing + config_path: /opt/inference-harness/litellm_config.yaml + host: 192.168.68.116 + +-- Phase 3: Apply agent context compression optimization + +call apply-agent-compression + agent: mumuni + host: 192.168.68.123 + config_path: /root/.hermes/config.yaml + +-- Phase 4: Enable llama.cpp prompt caching on GPU hosts + +call enable-prompt-caching + hosts: [192.168.68.15, 192.168.68.8, 192.168.68.110] + +-- Phase 5: Verify end-to-end latency + +call verify-latency + host: 192.168.68.116 + models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b, ornith-1.0-35b] +``` -- 2.54.0 From ba2c55c7e60f252c508a2f4d6be6cc9d16b849ac Mon Sep 17 00:00:00 2001 From: root Date: Fri, 17 Jul 2026 11:10:29 +0000 Subject: [PATCH 2/3] ci: re-trigger pipeline for PR #20 -- 2.54.0 From 4a1f47662368074e99705470b5c765d10b2c711a Mon Sep 17 00:00:00 2001 From: root Date: Fri, 17 Jul 2026 11:11:27 +0000 Subject: [PATCH 3/3] fix: add missing description to inference-optimization frontmatter --- inference-optimization.prose.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/inference-optimization.prose.md b/inference-optimization.prose.md index 58c8760..91a10e6 100644 --- a/inference-optimization.prose.md +++ b/inference-optimization.prose.md @@ -1,6 +1,10 @@ --- name: inference-optimization kind: responsibility +description: > + Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model + assignments, agent context management, and prompt caching — to reduce response + times to sub-15s average. All GPUs now at 128K context (stable ceiling). id: 067NC6KP02RG60S50M40E30928 --- -- 2.54.0