feat: GPU context 256K→128K fleet-wide + Genesis Hermes V3 on Strix Halo
GPU Changes: - All 3 GPUs reduced from 256K (-c 262144) to 128K (-c 131072) for stability - Observed instability near 100K at 256K — 128K is the stable ceiling - VRAM improved: RTX 3090 ~70% (was 90%), RTX 5070 ~65% (was 88%) - Strix Halo swapped to LuffyTheFox/Genesis Hermes V3 APEX - Hermes agent fine-tune, tensor repair (3 SSM layers, 76% W1 improvement) - Uncensored (0/465 refusals), multimodal (mmproj F16) - Speed: 65 tok/s gen, 140 tok/s prompt - Alias strix-moe maintained Agent Updates: - Mumuni: max_context_window 262144→131072, already aligned on strix-moe/0.65 - Tanko: max_context_window 262144→131072 - Koonimo: max_context_window + context_length 262144→131072 - CT114 SSH access confirmed (was 'Zulip only') LiteLLM (CT116): - Updated backend model references qwen3.6-35B-udq4→strix-moe - Removed stale ornith-1.0-35b from model_cost - Fallback chains updated Contracts Updated: - gpu-fleet.prose.md: topology, VRAM, benchmarks, config lines, model assignments - gpu-self-heal.prose.md: Rule 9/10 context targets - hermes-config-template.prose.md: template values, Rules 7-9, compression thresholds - inference-optimization.prose.md: added to repo, 128K recommendation Compression: 0.65 fires at ~85K (~43K headroom before 128K ceiling) For >128K workloads: route to external providers (deepseek)
This commit is contained in:
@@ -118,9 +118,9 @@ depends_on:
|
||||
|
||||
### Rule 9: Context Window Optimization
|
||||
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context
|
||||
- RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
|
||||
- RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
|
||||
- Strix Halo (256K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
|
||||
- RTX 3090 (128K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
|
||||
- RTX 5070 (128K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
|
||||
- Strix Halo (128K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
|
||||
- **Fix**:
|
||||
- If tok/s > baseline → context has headroom, consider increasing
|
||||
- If tok/s < 90% baseline → reduce context by 25% and retest
|
||||
@@ -131,9 +131,9 @@ depends_on:
|
||||
### Rule 10: Workload Distribution Optimization
|
||||
- **Detect**: GPU roles misaligned with hardware capabilities
|
||||
- **Target distribution**:
|
||||
- RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations
|
||||
- RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
|
||||
- Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs
|
||||
- RTX 3090 (24GB, 128K, 75 tok/s) → Heavy reasoning, code gen, long conversations
|
||||
- RTX 5070 (12GB, 128K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
|
||||
- Strix Halo (64GB, 128K, 72 tok/s) → Context compression, summarization, long docs
|
||||
- **Fix**:
|
||||
- Alert if any GPU is handling workload outside its designated role
|
||||
- Recommend Hermes agent profile updates to match workload to GPU
|
||||
|
||||
Reference in New Issue
Block a user