feat: GPU workload redistribution — compression → Strix Halo
- Move compression model from gemma-4-12b (RTX 5070) to ornith-1.0-35b (Strix Halo) - Add Rule 8: GPU Workload Distribution — per-GPU role assignment - Add Rule 9: Compression Threshold for 256K models - Update Rule 7: Auxiliary Model Consistency with new compression routing - Add gpu-self-heal.prose.md contract with 10 remediation rules - Strix Halo (64GB, 256K, 72.4 tok/s) → compression specialist - RTX 5070 (12GB) → vision/web search specialist - RTX 3090 (24GB, 256K) → heavy reasoning specialist - All rules grilled and confirmed with Kwame 2026-07-12
This commit is contained in:
@@ -13,27 +13,30 @@ description: >
|
||||
|
||||
- template_version: "2.1.0"
|
||||
- last_applied: timestamp
|
||||
- agents_configured: ["tanko", "mumuni", "abiba", "tdunna", "baggy", "kagenz0"]
|
||||
- agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
|
||||
- agent_keys: map (see Agent Keys section)
|
||||
- infra_endpoints_verified: array
|
||||
|
||||
## Agent Keys (LiteLLM — Current 2026-07-04)
|
||||
## Agent Keys (LiteLLM — Current 2026-07-11)
|
||||
|
||||
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
|
||||
PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB, not in
|
||||
config files. The env var `LITELLM_API_KEY` is set in `/etc/environment` on each agent
|
||||
host AND in `~/.hermes/.env` for gateway env propagation.
|
||||
PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB.
|
||||
The env var `LITELLM_API_KEY` is injected at runtime via `infisical run --` wrapper
|
||||
(project=agents, env=production). /etc/environment and ~/.hermes/.env are NO LONGER
|
||||
used for agent keys — stripped and tagged `# [INFISICAL]` post-migration.
|
||||
Sub-agent profiles inherit auth from the main config — no separate keys needed.
|
||||
|
||||
| Agent | Key Alias | Host | SSH | Sub-Agents |
|
||||
|-------|-----------|------|-----|-----------|
|
||||
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
|
||||
| Mumuni | `mumuni-jul2026` | 192.168.68.123 | root@.123 | 6 profiles ✱ |
|
||||
| Abiba | `abiba-*` | 192.168.68.24 | local | — |
|
||||
| Tdunna | `tdunna-*` | ? | Zulip | — |
|
||||
| Baggy | `baggy-*` | ? | Zulip | — |
|
||||
| Abiba | `abiba-pi` | 192.168.68.24 | local | — |
|
||||
| Koby | `koby` | CT 111 (tdunna) | Zulip | — |
|
||||
| Koonimo | `koonimo` | CT 113 (baggy) | Zulip | — |
|
||||
| Kagenz0 | `kagenz0-*` | ? | Zulip | — |
|
||||
|
||||
> CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
|
||||
|
||||
✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research,
|
||||
syslog-review, syslog-writer — all at `/root/.hermes/profiles/<name>/config.yaml`
|
||||
|
||||
@@ -51,11 +54,12 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
|
||||
## API Key Rules
|
||||
|
||||
- `api_key_env: LITELLM_API_KEY` — Use env var for main model auth (preferred)
|
||||
Key is injected at runtime via `infisical run --` wrapper — never in /etc/environment
|
||||
- `api_key: ''` — Sub-agents leave empty to inherit from main config's custom_provider
|
||||
- `api_key: sk-...` — Hardcoded key only as fallback when env var not possible
|
||||
- Set `LITELLM_API_KEY` in `/etc/environment` on each host
|
||||
- Store `LITELLM_API_KEY` in Infisical vault (project=agents, env=production)
|
||||
- Sub-agents NEVER get their own key — they share the host agent's key
|
||||
- Restart Hermes after updating `/etc/environment`
|
||||
- Restart Hermes gateway after updating vault secret (key auto-injected via wrapper)
|
||||
|
||||
### Sub-Agent Profiles (Mumuni pattern)
|
||||
|
||||
@@ -79,7 +83,7 @@ Sub-agent profile rules:
|
||||
5. **Auxiliary tasks** (vision, compression, etc.) also leave `api_key` empty
|
||||
6. **Never hardcode a key** in sub-agent profiles
|
||||
|
||||
This ensures all 6 sub-agents use the same LiteLLM key set in `/etc/environment`.
|
||||
This ensures all 6 sub-agents use the same LiteLLM key injected via `infisical run --` wrapper.
|
||||
When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs)
|
||||
work immediately after restart.
|
||||
|
||||
@@ -91,7 +95,7 @@ model:
|
||||
default: <agent_model> # e.g., ornith-1.0-35b, qwen3.6-27B-code
|
||||
provider: harness
|
||||
base_url: http://192.168.68.116/v1
|
||||
api_key_env: LITELLM_API_KEY # Set in /etc/environment AND ~/.hermes/.env
|
||||
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
|
||||
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
||||
context_length: 262144 # For syslog-auto (ornith route supports 256K).
|
||||
# Set 131072 if using qwen3.6-27B-code or gemma-4-12b directly.
|
||||
@@ -176,7 +180,7 @@ custom_providers:
|
||||
|
||||
When LiteLLM keys are regenerated (e.g., after infrastructure changes):
|
||||
|
||||
1. **If SSH available**: `ssh <host> "sudo sed -i 's/LITELLM_API_KEY=.*/LITELLM_API_KEY=sk-<NEW>/' /etc/environment"`
|
||||
1. **If SSH available**: Update Infisical vault: `infisical secrets set LITELLM_API_KEY=sk-<NEW> --project=agents --env=production`, then `ssh <host> "systemctl restart hermes-gateway"`
|
||||
2. **If SSH unavailable**: Send Zulip DM via abiba-bot with update command
|
||||
3. **After update**: Restart Hermes on the agent host
|
||||
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
||||
@@ -198,7 +202,7 @@ The following MUST be identical across ALL profiles:
|
||||
### Rule 3: API Keys via Environment
|
||||
- Prefer `api_key_env: LITELLM_API_KEY` over hardcoded keys
|
||||
- Hardcoded keys in config.yaml become stale after key rotation
|
||||
- `/etc/environment` persists across config updates
|
||||
- Infisical vault secrets persist across config updates / reinstalls
|
||||
- Restart Hermes after env var updates
|
||||
|
||||
### Rule 4: Sub-Agent Profiles Inherit Auth
|
||||
@@ -208,7 +212,7 @@ The following MUST be identical across ALL profiles:
|
||||
- Auxiliary tasks: `api_key: ''`, `provider: harness`
|
||||
- Never hardcode a key in sub-agent profiles
|
||||
- When main config uses `api_key_env`, sub-agents automatically use it
|
||||
- This means key rotation only touches ONE file (`/etc/environment`)
|
||||
- This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`)
|
||||
|
||||
### Rule 5: Main Config Base URL
|
||||
|- Use direct IP: `http://192.168.68.116/v1`
|
||||
@@ -223,23 +227,42 @@ The following MUST be identical across ALL profiles:
|
||||
- Apply to BOTH main config AND all sub-agent profiles
|
||||
- For agents needing longer outputs: raise to 8192, but never omit
|
||||
|
||||
### Rule 7: Auxiliary Model Consistency
|
||||
- All auxiliary services (vision, web_extract, compression) MUST use the same model:
|
||||
- `model: gemma-4-12b`
|
||||
### Rule 7: Auxiliary Model Consistency (UPDATED July 2026)
|
||||
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
|
||||
- Compression uses `ornith-1.0-35b` (Strix Halo — 64GB, 256K ctx, compression-optimized)
|
||||
- All auxiliary services MUST use identical routing:
|
||||
- `base_url: http://192.168.68.116/v1`
|
||||
- `api_key_env: LITELLM_API_KEY`
|
||||
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes to the primary 35B reasoning GPU
|
||||
- gemma-4-12b is a lightweight 12B model on the RTX 5070, keeping the Strix Halo free for reasoning
|
||||
- The `compression:` block's `model` MUST match `auxiliary: compression: model` — they are two different configs for the same service
|
||||
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
|
||||
- **Compression moved to Strix Halo (July 2026)**: The ornith-1.0-35b model on Strix Halo
|
||||
(64GB UMA, 256K context, 72.4 tok/s) is the designated compression GPU. This frees the
|
||||
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
|
||||
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
|
||||
- The `compression: max_context_window: 262144` MUST match Strix Halo's actual capacity
|
||||
|
||||
### Rule 8: Compression Threshold for 256K Models
|
||||
### Rule 8: GPU Workload Distribution (July 2026)
|
||||
- **RTX 3090 (24GB, 256K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
|
||||
- **RTX 5070 (12GB, 131K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract
|
||||
- **Strix Halo (64GB, 256K ctx, ornith-1.0-35b)**: Context compression, summarization, long docs
|
||||
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
||||
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
|
||||
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
|
||||
- `auxiliary.compression.model: ornith-1.0-35b` (Strix Halo)
|
||||
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
|
||||
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
|
||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
|
||||
- `max_context_window: 262144` MUST match the model's actual capacity
|
||||
- See `devops-hermes-compression` skill for full reference
|
||||
|
||||
### Rule 9: Default Model Must Be `syslog-auto` (All Agents)
|
||||
### Rule 9: Compression Threshold for 256K Models
|
||||
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
|
||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
|
||||
- `max_context_window: 262144` MUST match the model's actual capacity (Strix Halo = 256K)
|
||||
- See `devops-hermes-compression` skill for full reference
|
||||
|
||||
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
|
||||
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
|
||||
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
|
||||
- `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b
|
||||
@@ -250,7 +273,7 @@ The following MUST be identical across ALL profiles:
|
||||
- **Exception**: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models
|
||||
for specialized tasks, but MUST validate those models exist in the key's authorized list
|
||||
|
||||
### Rule 10: Validate Model IDs Before Deployment (pi Agents)
|
||||
### Rule 11: Validate Model IDs Before Deployment (pi Agents)
|
||||
- After configuring a pi agent's `models.json`, verify every model ID:
|
||||
```bash
|
||||
curl -s http://192.168.68.116:4000/v1/models \
|
||||
|
||||
Reference in New Issue
Block a user