feat: GPU workload redistribution — compression → Strix Halo
- Move compression model from gemma-4-12b (RTX 5070) to ornith-1.0-35b (Strix Halo) - Add Rule 8: GPU Workload Distribution — per-GPU role assignment - Add Rule 9: Compression Threshold for 256K models - Update Rule 7: Auxiliary Model Consistency with new compression routing - Add gpu-self-heal.prose.md contract with 10 remediation rules - Strix Halo (64GB, 256K, 72.4 tok/s) → compression specialist - RTX 5070 (12GB) → vision/web search specialist - RTX 3090 (24GB, 256K) → heavy reasoning specialist - All rules grilled and confirmed with Kwame 2026-07-12
This commit is contained in:
@@ -0,0 +1,258 @@
|
|||||||
|
---
|
||||||
|
kind: responsibility
|
||||||
|
name: gpu-self-heal
|
||||||
|
description: >
|
||||||
|
GPU fleet self-healing — detects anomalies, applies remediation, tracks
|
||||||
|
benchmarks, and predicts failures before they happen. Extends gpu-monitor
|
||||||
|
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
|
||||||
|
VRAM trend analysis, and predictive alerting.
|
||||||
|
agent: abiba
|
||||||
|
depends_on:
|
||||||
|
- gpu-monitor.prose.md (live data source on .24:9100)
|
||||||
|
- gpu-fleet.prose.md (source of truth for topology)
|
||||||
|
---
|
||||||
|
|
||||||
|
## Maintains
|
||||||
|
|
||||||
|
- gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array }
|
||||||
|
- gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail
|
||||||
|
- benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts }
|
||||||
|
- vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours }
|
||||||
|
- circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset }
|
||||||
|
|
||||||
|
## Requires
|
||||||
|
|
||||||
|
- gpu-monitor:function — Live fleet data from .24:9100/gpu-data
|
||||||
|
- Prometheus exporters on all 3 GPUs (:9400/metrics)
|
||||||
|
- SSH access to GPU hosts for restart operations
|
||||||
|
|
||||||
|
## Continuity
|
||||||
|
|
||||||
|
- Self-driven: check every 60 seconds against GPU monitor data
|
||||||
|
- Also wakes on gpu-fleet health degradation
|
||||||
|
- On fix: verify with benchmark inference test before declaring resolved
|
||||||
|
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Remediation Rules
|
||||||
|
|
||||||
|
### Rule 1: GPU Temperature Critical (>85°C for >2 min)
|
||||||
|
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
|
||||||
|
- **Fix**:
|
||||||
|
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
|
||||||
|
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains
|
||||||
|
3. If all GPUs hot, alert about cooling infrastructure
|
||||||
|
- **Verify**: Temp drops below 80°C within 5 minutes
|
||||||
|
- **Escalate after**: 3 verification failures → Zulip alert
|
||||||
|
|
||||||
|
### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
|
||||||
|
- **Detect**: VRAM growing at sustained rate over 6+ hour window
|
||||||
|
- RTX 3090 (24GB): ≥100MB/hour
|
||||||
|
- RTX 5070 (12GB): ≥50MB/hour
|
||||||
|
- Strix Halo (64GB UMA): ≥200MB/hour
|
||||||
|
- **Fix**:
|
||||||
|
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
|
||||||
|
2. If llama-server is the growth source → restart with memory cap flag
|
||||||
|
3. If unknown process → kill and alert
|
||||||
|
- **Verify**: VRAM growth rate drops below threshold
|
||||||
|
- **Escalate after**: persistent leak after restart → hardware investigation
|
||||||
|
|
||||||
|
### Rule 3: Model Inference Timeout / GPU Stuck
|
||||||
|
- **Detect**: >50% failure rate over 60s window + 30s grace period (not just single stuck request)
|
||||||
|
- **Fix**:
|
||||||
|
1. Restart llama-server on affected GPU host
|
||||||
|
2. Wait 15s for model to reload
|
||||||
|
3. Run benchmark inference test
|
||||||
|
- **Verify**: Model returns 200 with <30s response, failure rate drops to 0%
|
||||||
|
- **Escalate after**: 3 restarts in 1 hour → GPU hardware check
|
||||||
|
|
||||||
|
### Rule 4: Benchmark Regression (>20% drop)
|
||||||
|
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
|
||||||
|
- **Fix**:
|
||||||
|
1. Check GPU utilization — if >90%, other process is competing
|
||||||
|
2. Check power limit — if throttled, restore to max
|
||||||
|
3. Check thermal — if hot, apply Rule 1
|
||||||
|
- **Verify**: Benchmark returns to within 10% of baseline
|
||||||
|
- **Escalate after**: persistent regression → possible hardware degradation
|
||||||
|
|
||||||
|
### Rule 5: Circuit Breaker Stuck Open
|
||||||
|
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
|
||||||
|
- **Fix**:
|
||||||
|
1. Verify GPU /health returns 200
|
||||||
|
2. If GPU healthy, send 1 test inference
|
||||||
|
3. If test succeeds → reset circuit breaker via router API
|
||||||
|
4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset
|
||||||
|
5. Max 1 auto-reset per GPU per hour
|
||||||
|
- **Verify**: CB closes, inference succeeds, CB stays closed for 60s+
|
||||||
|
- **Escalate after**: CB won't close after reset → router issue
|
||||||
|
|
||||||
|
### Rule 6: Strix Halo Unreachable
|
||||||
|
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
|
||||||
|
- **Fix**:
|
||||||
|
1. SSH to .15 → check llama-server process
|
||||||
|
2. Restart llama-server if not running
|
||||||
|
3. Verify through both direct probe AND router
|
||||||
|
- **Verify**: Direct health probe returns 200, router reports Strix healthy
|
||||||
|
- **Escalate**: If host .15 itself is unreachable → infrastructure alert
|
||||||
|
|
||||||
|
### Rule 7: Prometheus Exporter Down
|
||||||
|
- **Detect**: Any GPU :9400/metrics unreachable for >2 polls
|
||||||
|
- **Fix**:
|
||||||
|
1. SSH to GPU host → check prometheus-exporter process
|
||||||
|
2. Restart exporter if dead
|
||||||
|
3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes
|
||||||
|
4. If exporter is running but unreachable → check firewall/host networking
|
||||||
|
- **Verify**: :9400/metrics returns 200
|
||||||
|
- **Escalate after**: 3 failed restarts → networking issue
|
||||||
|
|
||||||
|
### Rule 8: Predictive Thermal Warning (two-tier)
|
||||||
|
- **Detect**:
|
||||||
|
- Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert
|
||||||
|
- Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding
|
||||||
|
- **Fix**:
|
||||||
|
- Tier 1: silently reduce parallel requests to that GPU by 50%
|
||||||
|
- Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic
|
||||||
|
- **Verify**: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2)
|
||||||
|
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
|
||||||
|
|
||||||
|
### Rule 9: Context Window Optimization
|
||||||
|
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context
|
||||||
|
- RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
|
||||||
|
- RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
|
||||||
|
- Strix Halo (256K ctx, ornith-1.0-35b): target 70+ tok/s — currently above baseline
|
||||||
|
- **Fix**:
|
||||||
|
- If tok/s > baseline → context has headroom, consider increasing
|
||||||
|
- If tok/s < 90% baseline → reduce context by 25% and retest
|
||||||
|
- If tok/s within 10% of baseline → optimal, no change
|
||||||
|
- **Verify**: Re-benchmark after context change, confirm within 10% of target
|
||||||
|
- **Escalate**: If context can't be adjusted without significant perf loss
|
||||||
|
|
||||||
|
### Rule 10: Workload Distribution Optimization
|
||||||
|
- **Detect**: GPU roles misaligned with hardware capabilities
|
||||||
|
- **Target distribution**:
|
||||||
|
- RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations
|
||||||
|
- RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
|
||||||
|
- Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs
|
||||||
|
- **Fix**:
|
||||||
|
- Alert if any GPU is handling workload outside its designated role
|
||||||
|
- Recommend Hermes agent profile updates to match workload to GPU
|
||||||
|
- Track per-GPU request distribution via LiteLLM spend logs
|
||||||
|
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
||||||
|
- **Escalate**: If role mismatch persists >48h → agent profile audit needed
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Execution
|
||||||
|
|
||||||
|
```prose
|
||||||
|
-- Phase 1: Fetch live GPU data
|
||||||
|
let fleet = call gpu-monitor
|
||||||
|
endpoint: "http://192.168.68.24:9100/gpu-data"
|
||||||
|
|
||||||
|
-- Phase 2: Evaluate each GPU against remediation rules
|
||||||
|
let actions = []
|
||||||
|
for gpu in fleet.gpus:
|
||||||
|
-- Rule 1: Thermal critical
|
||||||
|
if gpu.temp_c > 85 and sustained_for(gpu, 120):
|
||||||
|
push actions apply-thermal-fix(gpu)
|
||||||
|
|
||||||
|
-- Rule 2: VRAM leak
|
||||||
|
let vram_rate = calculate-vram-trend(gpu, hours=6)
|
||||||
|
if vram_rate > 50:
|
||||||
|
push actions apply-vram-fix(gpu, vram_rate)
|
||||||
|
|
||||||
|
-- Rule 4: Benchmark regression
|
||||||
|
let bench = fleet.benchmarks[gpu.hostname]
|
||||||
|
if bench.current_tok_sec < bench.baseline_tok_sec * 0.8:
|
||||||
|
push actions apply-benchmark-fix(gpu, bench)
|
||||||
|
|
||||||
|
-- Rule 3: Model stuck
|
||||||
|
for model in fleet.router.available_models:
|
||||||
|
if model.consecutive_timeouts >= 3:
|
||||||
|
push actions apply-model-restart(model)
|
||||||
|
|
||||||
|
-- Rule 5: Circuit breaker
|
||||||
|
for cb in fleet.router.circuit_breaker:
|
||||||
|
if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu):
|
||||||
|
push actions apply-cb-reset(cb)
|
||||||
|
|
||||||
|
-- Rule 6: Strix Halo
|
||||||
|
if not fleet.strix.running and pingable("192.168.68.15"):
|
||||||
|
push actions apply-strix-restart()
|
||||||
|
|
||||||
|
-- Rule 7: Prometheus exporters
|
||||||
|
for gpu in fleet.gpus:
|
||||||
|
if not prometheus_reachable(gpu.hostname, 9400):
|
||||||
|
push actions apply-exporter-restart(gpu)
|
||||||
|
|
||||||
|
-- Rule 8: Predictive thermal
|
||||||
|
for gpu in fleet.gpus:
|
||||||
|
let rise_rate = calculate-temp-rise(gpu, minutes=5)
|
||||||
|
if rise_rate > 2.0 and gpu.temp_c < 80:
|
||||||
|
push actions apply-proactive-cooling(gpu)
|
||||||
|
|
||||||
|
-- Phase 3: Execute actions, verify, log
|
||||||
|
for action in actions:
|
||||||
|
let result = execute-with-verify(action)
|
||||||
|
log-to-kg(action, result)
|
||||||
|
if result.failed:
|
||||||
|
escalate-if-needed(action)
|
||||||
|
|
||||||
|
-- Phase 4: Update health state
|
||||||
|
call update-gpu-health
|
||||||
|
gpus: fleet.gpus
|
||||||
|
actions: actions
|
||||||
|
status: derive-overall-status(fleet, actions)
|
||||||
|
|
||||||
|
-- Wait 60s and repeat
|
||||||
|
```
|
||||||
|
|
||||||
|
## Audit Trail Format
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"run_id": "gpu-self-heal-20260712-001",
|
||||||
|
"timestamp": "2026-07-12T16:00:00Z",
|
||||||
|
"gpu": "ct8-rtx3090",
|
||||||
|
"issue": "thermal-critical",
|
||||||
|
"detected": { "temp_c": 87, "duration_s": 180 },
|
||||||
|
"action": "set-fan-100pct",
|
||||||
|
"result": "resolved",
|
||||||
|
"verification": { "temp_c": 76, "after_s": 300 },
|
||||||
|
"escalated": false
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Reporting
|
||||||
|
|
||||||
|
### 1. Knowledge Graph
|
||||||
|
Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail.
|
||||||
|
|
||||||
|
### 2. Zulip Alerts (#agent-hub → alerts-gpu)
|
||||||
|
- `issues_fixed > 0` → "🛠 GPU Self-Heal — <gpu> <issue> resolved"
|
||||||
|
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
|
||||||
|
- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
|
||||||
|
|
||||||
|
### 3. Prometheus/Grafana Integration
|
||||||
|
- GPU self-heal actions exposed as Prometheus counter metrics
|
||||||
|
- Dashboard panel: "GPU Interventions (24h)" showing count/type/result
|
||||||
|
|
||||||
|
### 4. Weekly Benchmark Report
|
||||||
|
- Per-GPU tok/s trend over 7 days
|
||||||
|
- Regression alerts if any GPU degrades >10% week-over-week
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Design Decisions (Grilled & Confirmed — 2026-07-12)
|
||||||
|
|
||||||
|
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
|
||||||
|
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
|
||||||
|
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
|
||||||
|
4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix).
|
||||||
|
5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU.
|
||||||
|
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana.
|
||||||
|
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
|
||||||
|
8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.
|
||||||
@@ -13,27 +13,30 @@ description: >
|
|||||||
|
|
||||||
- template_version: "2.1.0"
|
- template_version: "2.1.0"
|
||||||
- last_applied: timestamp
|
- last_applied: timestamp
|
||||||
- agents_configured: ["tanko", "mumuni", "abiba", "tdunna", "baggy", "kagenz0"]
|
- agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
|
||||||
- agent_keys: map (see Agent Keys section)
|
- agent_keys: map (see Agent Keys section)
|
||||||
- infra_endpoints_verified: array
|
- infra_endpoints_verified: array
|
||||||
|
|
||||||
## Agent Keys (LiteLLM — Current 2026-07-04)
|
## Agent Keys (LiteLLM — Current 2026-07-11)
|
||||||
|
|
||||||
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
|
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
|
||||||
PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB, not in
|
PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB.
|
||||||
config files. The env var `LITELLM_API_KEY` is set in `/etc/environment` on each agent
|
The env var `LITELLM_API_KEY` is injected at runtime via `infisical run --` wrapper
|
||||||
host AND in `~/.hermes/.env` for gateway env propagation.
|
(project=agents, env=production). /etc/environment and ~/.hermes/.env are NO LONGER
|
||||||
|
used for agent keys — stripped and tagged `# [INFISICAL]` post-migration.
|
||||||
Sub-agent profiles inherit auth from the main config — no separate keys needed.
|
Sub-agent profiles inherit auth from the main config — no separate keys needed.
|
||||||
|
|
||||||
| Agent | Key Alias | Host | SSH | Sub-Agents |
|
| Agent | Key Alias | Host | SSH | Sub-Agents |
|
||||||
|-------|-----------|------|-----|-----------|
|
|-------|-----------|------|-----|-----------|
|
||||||
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
|
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
|
||||||
| Mumuni | `mumuni-jul2026` | 192.168.68.123 | root@.123 | 6 profiles ✱ |
|
| Mumuni | `mumuni-jul2026` | 192.168.68.123 | root@.123 | 6 profiles ✱ |
|
||||||
| Abiba | `abiba-*` | 192.168.68.24 | local | — |
|
| Abiba | `abiba-pi` | 192.168.68.24 | local | — |
|
||||||
| Tdunna | `tdunna-*` | ? | Zulip | — |
|
| Koby | `koby` | CT 111 (tdunna) | Zulip | — |
|
||||||
| Baggy | `baggy-*` | ? | Zulip | — |
|
| Koonimo | `koonimo` | CT 113 (baggy) | Zulip | — |
|
||||||
| Kagenz0 | `kagenz0-*` | ? | Zulip | — |
|
| Kagenz0 | `kagenz0-*` | ? | Zulip | — |
|
||||||
|
|
||||||
|
> CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
|
||||||
|
|
||||||
✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research,
|
✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research,
|
||||||
syslog-review, syslog-writer — all at `/root/.hermes/profiles/<name>/config.yaml`
|
syslog-review, syslog-writer — all at `/root/.hermes/profiles/<name>/config.yaml`
|
||||||
|
|
||||||
@@ -51,11 +54,12 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
|
|||||||
## API Key Rules
|
## API Key Rules
|
||||||
|
|
||||||
- `api_key_env: LITELLM_API_KEY` — Use env var for main model auth (preferred)
|
- `api_key_env: LITELLM_API_KEY` — Use env var for main model auth (preferred)
|
||||||
|
Key is injected at runtime via `infisical run --` wrapper — never in /etc/environment
|
||||||
- `api_key: ''` — Sub-agents leave empty to inherit from main config's custom_provider
|
- `api_key: ''` — Sub-agents leave empty to inherit from main config's custom_provider
|
||||||
- `api_key: sk-...` — Hardcoded key only as fallback when env var not possible
|
- `api_key: sk-...` — Hardcoded key only as fallback when env var not possible
|
||||||
- Set `LITELLM_API_KEY` in `/etc/environment` on each host
|
- Store `LITELLM_API_KEY` in Infisical vault (project=agents, env=production)
|
||||||
- Sub-agents NEVER get their own key — they share the host agent's key
|
- Sub-agents NEVER get their own key — they share the host agent's key
|
||||||
- Restart Hermes after updating `/etc/environment`
|
- Restart Hermes gateway after updating vault secret (key auto-injected via wrapper)
|
||||||
|
|
||||||
### Sub-Agent Profiles (Mumuni pattern)
|
### Sub-Agent Profiles (Mumuni pattern)
|
||||||
|
|
||||||
@@ -79,7 +83,7 @@ Sub-agent profile rules:
|
|||||||
5. **Auxiliary tasks** (vision, compression, etc.) also leave `api_key` empty
|
5. **Auxiliary tasks** (vision, compression, etc.) also leave `api_key` empty
|
||||||
6. **Never hardcode a key** in sub-agent profiles
|
6. **Never hardcode a key** in sub-agent profiles
|
||||||
|
|
||||||
This ensures all 6 sub-agents use the same LiteLLM key set in `/etc/environment`.
|
This ensures all 6 sub-agents use the same LiteLLM key injected via `infisical run --` wrapper.
|
||||||
When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs)
|
When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs)
|
||||||
work immediately after restart.
|
work immediately after restart.
|
||||||
|
|
||||||
@@ -91,7 +95,7 @@ model:
|
|||||||
default: <agent_model> # e.g., ornith-1.0-35b, qwen3.6-27B-code
|
default: <agent_model> # e.g., ornith-1.0-35b, qwen3.6-27B-code
|
||||||
provider: harness
|
provider: harness
|
||||||
base_url: http://192.168.68.116/v1
|
base_url: http://192.168.68.116/v1
|
||||||
api_key_env: LITELLM_API_KEY # Set in /etc/environment AND ~/.hermes/.env
|
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
|
||||||
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
|
||||||
context_length: 262144 # For syslog-auto (ornith route supports 256K).
|
context_length: 262144 # For syslog-auto (ornith route supports 256K).
|
||||||
# Set 131072 if using qwen3.6-27B-code or gemma-4-12b directly.
|
# Set 131072 if using qwen3.6-27B-code or gemma-4-12b directly.
|
||||||
@@ -176,7 +180,7 @@ custom_providers:
|
|||||||
|
|
||||||
When LiteLLM keys are regenerated (e.g., after infrastructure changes):
|
When LiteLLM keys are regenerated (e.g., after infrastructure changes):
|
||||||
|
|
||||||
1. **If SSH available**: `ssh <host> "sudo sed -i 's/LITELLM_API_KEY=.*/LITELLM_API_KEY=sk-<NEW>/' /etc/environment"`
|
1. **If SSH available**: Update Infisical vault: `infisical secrets set LITELLM_API_KEY=sk-<NEW> --project=agents --env=production`, then `ssh <host> "systemctl restart hermes-gateway"`
|
||||||
2. **If SSH unavailable**: Send Zulip DM via abiba-bot with update command
|
2. **If SSH unavailable**: Send Zulip DM via abiba-bot with update command
|
||||||
3. **After update**: Restart Hermes on the agent host
|
3. **After update**: Restart Hermes on the agent host
|
||||||
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
|
||||||
@@ -198,7 +202,7 @@ The following MUST be identical across ALL profiles:
|
|||||||
### Rule 3: API Keys via Environment
|
### Rule 3: API Keys via Environment
|
||||||
- Prefer `api_key_env: LITELLM_API_KEY` over hardcoded keys
|
- Prefer `api_key_env: LITELLM_API_KEY` over hardcoded keys
|
||||||
- Hardcoded keys in config.yaml become stale after key rotation
|
- Hardcoded keys in config.yaml become stale after key rotation
|
||||||
- `/etc/environment` persists across config updates
|
- Infisical vault secrets persist across config updates / reinstalls
|
||||||
- Restart Hermes after env var updates
|
- Restart Hermes after env var updates
|
||||||
|
|
||||||
### Rule 4: Sub-Agent Profiles Inherit Auth
|
### Rule 4: Sub-Agent Profiles Inherit Auth
|
||||||
@@ -208,7 +212,7 @@ The following MUST be identical across ALL profiles:
|
|||||||
- Auxiliary tasks: `api_key: ''`, `provider: harness`
|
- Auxiliary tasks: `api_key: ''`, `provider: harness`
|
||||||
- Never hardcode a key in sub-agent profiles
|
- Never hardcode a key in sub-agent profiles
|
||||||
- When main config uses `api_key_env`, sub-agents automatically use it
|
- When main config uses `api_key_env`, sub-agents automatically use it
|
||||||
- This means key rotation only touches ONE file (`/etc/environment`)
|
- This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`)
|
||||||
|
|
||||||
### Rule 5: Main Config Base URL
|
### Rule 5: Main Config Base URL
|
||||||
|- Use direct IP: `http://192.168.68.116/v1`
|
|- Use direct IP: `http://192.168.68.116/v1`
|
||||||
@@ -223,23 +227,42 @@ The following MUST be identical across ALL profiles:
|
|||||||
- Apply to BOTH main config AND all sub-agent profiles
|
- Apply to BOTH main config AND all sub-agent profiles
|
||||||
- For agents needing longer outputs: raise to 8192, but never omit
|
- For agents needing longer outputs: raise to 8192, but never omit
|
||||||
|
|
||||||
### Rule 7: Auxiliary Model Consistency
|
### Rule 7: Auxiliary Model Consistency (UPDATED July 2026)
|
||||||
- All auxiliary services (vision, web_extract, compression) MUST use the same model:
|
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
|
||||||
- `model: gemma-4-12b`
|
- Compression uses `ornith-1.0-35b` (Strix Halo — 64GB, 256K ctx, compression-optimized)
|
||||||
|
- All auxiliary services MUST use identical routing:
|
||||||
- `base_url: http://192.168.68.116/v1`
|
- `base_url: http://192.168.68.116/v1`
|
||||||
- `api_key_env: LITELLM_API_KEY`
|
- `api_key_env: LITELLM_API_KEY`
|
||||||
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes to the primary 35B reasoning GPU
|
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
|
||||||
- gemma-4-12b is a lightweight 12B model on the RTX 5070, keeping the Strix Halo free for reasoning
|
- **Compression moved to Strix Halo (July 2026)**: The ornith-1.0-35b model on Strix Halo
|
||||||
- The `compression:` block's `model` MUST match `auxiliary: compression: model` — they are two different configs for the same service
|
(64GB UMA, 256K context, 72.4 tok/s) is the designated compression GPU. This frees the
|
||||||
|
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
|
||||||
|
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
|
||||||
|
- The `compression: max_context_window: 262144` MUST match Strix Halo's actual capacity
|
||||||
|
|
||||||
### Rule 8: Compression Threshold for 256K Models
|
### Rule 8: GPU Workload Distribution (July 2026)
|
||||||
|
- **RTX 3090 (24GB, 256K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
|
||||||
|
- **RTX 5070 (12GB, 131K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract
|
||||||
|
- **Strix Halo (64GB, 256K ctx, ornith-1.0-35b)**: Context compression, summarization, long docs
|
||||||
|
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
||||||
|
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
|
||||||
|
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
|
||||||
|
- `auxiliary.compression.model: ornith-1.0-35b` (Strix Halo)
|
||||||
|
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
|
||||||
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
|
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
|
||||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||||
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
|
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
|
||||||
- `max_context_window: 262144` MUST match the model's actual capacity
|
- `max_context_window: 262144` MUST match the model's actual capacity
|
||||||
- See `devops-hermes-compression` skill for full reference
|
- See `devops-hermes-compression` skill for full reference
|
||||||
|
|
||||||
### Rule 9: Default Model Must Be `syslog-auto` (All Agents)
|
### Rule 9: Compression Threshold for 256K Models
|
||||||
|
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
|
||||||
|
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||||
|
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
|
||||||
|
- `max_context_window: 262144` MUST match the model's actual capacity (Strix Halo = 256K)
|
||||||
|
- See `devops-hermes-compression` skill for full reference
|
||||||
|
|
||||||
|
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
|
||||||
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
|
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
|
||||||
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
|
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
|
||||||
- `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b
|
- `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b
|
||||||
@@ -250,7 +273,7 @@ The following MUST be identical across ALL profiles:
|
|||||||
- **Exception**: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models
|
- **Exception**: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models
|
||||||
for specialized tasks, but MUST validate those models exist in the key's authorized list
|
for specialized tasks, but MUST validate those models exist in the key's authorized list
|
||||||
|
|
||||||
### Rule 10: Validate Model IDs Before Deployment (pi Agents)
|
### Rule 11: Validate Model IDs Before Deployment (pi Agents)
|
||||||
- After configuring a pi agent's `models.json`, verify every model ID:
|
- After configuring a pi agent's `models.json`, verify every model ID:
|
||||||
```bash
|
```bash
|
||||||
curl -s http://192.168.68.116:4000/v1/models \
|
curl -s http://192.168.68.116:4000/v1/models \
|
||||||
|
|||||||
Reference in New Issue
Block a user