gpu-fleet: July 8 optimization sweep — parallel 2 fleet-wide, 128K ctx on NVIDIA, LiteLLM timeout fixes
- All GPUs: --parallel 1 → 2 (6 concurrent slots, was 3) - .8 RTX 3090: ctx 256K→128K, VRAM 96%→83%, turbo4 KV cache - .110 RTX 5070: ctx 256K→128K, ubatch 4096→512 (was inverted), VRAM 90%→77%, q4_0 KV - .15 Strix Halo: parallel 1→2, 256K ctx (41GB free), q8_0 KV, AMD metrics via /sys/class/drm - LiteLLM: gemma timeout 25→120s, qwen timeout 40→90s, syslog-auto (qwen) 40→90s - Agent configs: context_length 262144 for syslog-auto, 131072 for direct qwen/gemma - Updated health-check operation, agent config implications, benchmark table (fixed model↔GPU mapping)
This commit is contained in:
@@ -5,7 +5,8 @@ description: >
|
||||
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
||||
Updated 2026-06-30: new LiteLLM keys, nginx routing, router back online.
|
||||
Updated 2026-07-08: GPU context reduced to 128K on NVIDIA (.8, .110),
|
||||
parallel 2 on all GPUs, LiteLLM timeouts tuned, context_length guidance added.
|
||||
---
|
||||
|
||||
## Maintains
|
||||
@@ -92,7 +93,8 @@ model:
|
||||
base_url: http://192.168.68.116/v1
|
||||
api_key_env: LITELLM_API_KEY # Set in /etc/environment AND ~/.hermes/.env
|
||||
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
||||
context_length: 262144 # Must match model's max context window
|
||||
context_length: 262144 # For syslog-auto (ornith route supports 256K).
|
||||
# Set 131072 if using qwen3.6-27B-code or gemma-4-12b directly.
|
||||
|
||||
fallback_providers:
|
||||
provider: deepseek
|
||||
@@ -123,8 +125,9 @@ compression:
|
||||
enabled: true
|
||||
model: gemma-4-12b # ⚠️ Must match auxiliary.compression.model
|
||||
provider: harness
|
||||
max_context_window: 262144 # Must match model's actual capacity
|
||||
threshold: 0.65 # Fires at ~170K for 262K window (not 0.25!)
|
||||
max_context_window: 262144 # For syslog-auto (ornith supports 256K).
|
||||
# Set 131072 if using qwen or gemma directly.
|
||||
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
|
||||
target_ratio: 0.30
|
||||
protect_last_n: 40
|
||||
hygiene_hard_message_limit: 350
|
||||
@@ -236,6 +239,28 @@ The following MUST be identical across ALL profiles:
|
||||
- `max_context_window: 262144` MUST match the model's actual capacity
|
||||
- See `devops-hermes-compression` skill for full reference
|
||||
|
||||
### Rule 9: Default Model Must Be `syslog-auto` (All Agents)
|
||||
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
|
||||
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
|
||||
- `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b
|
||||
and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against:
|
||||
- Model name typos that cause 403 errors and silent worker failures
|
||||
- Single GPU downtime (routing falls back automatically)
|
||||
- Key/model authorization mismatches
|
||||
- **Exception**: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models
|
||||
for specialized tasks, but MUST validate those models exist in the key's authorized list
|
||||
|
||||
### Rule 10: Validate Model IDs Before Deployment (pi Agents)
|
||||
- After configuring a pi agent's `models.json`, verify every model ID:
|
||||
```bash
|
||||
curl -s http://192.168.68.116:4000/v1/models \
|
||||
-H "Authorization: Bearer <AGENT_KEY>" | jq '.data[].id'
|
||||
```
|
||||
- All model IDs in `models.json` MUST appear in the LiteLLM response
|
||||
- The agent's API key may have a SUBSET of the full model catalog — check per-key
|
||||
- A non-existent model ID causes 403 errors that silently break the pi RPC worker
|
||||
(no `agent_end` emitted, worker stays "busy", Zulip messages pile up unprocessed)
|
||||
|
||||
## Execution
|
||||
|
||||
1. **Check current config** — Read the target agent's config.yaml
|
||||
|
||||
Reference in New Issue
Block a user