feat: GPU workload redistribution — compression → Strix Halo

- Move compression model from gemma-4-12b (RTX 5070) to ornith-1.0-35b (Strix Halo)
- Add Rule 8: GPU Workload Distribution — per-GPU role assignment
- Add Rule 9: Compression Threshold for 256K models
- Update Rule 7: Auxiliary Model Consistency with new compression routing
- Add gpu-self-heal.prose.md contract with 10 remediation rules
- Strix Halo (64GB, 256K, 72.4 tok/s) → compression specialist
- RTX 5070 (12GB) → vision/web search specialist
- RTX 3090 (24GB, 256K) → heavy reasoning specialist
- All rules grilled and confirmed with Kwame 2026-07-12
This commit is contained in:
root
2026-07-12 22:27:15 +00:00
parent 0134ad8be6
commit 19b6db9891
2 changed files with 305 additions and 24 deletions
+258
View File
@@ -0,0 +1,258 @@
---
kind: responsibility
name: gpu-self-heal
description: >
GPU fleet self-healing — detects anomalies, applies remediation, tracks
benchmarks, and predicts failures before they happen. Extends gpu-monitor
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
VRAM trend analysis, and predictive alerting.
agent: abiba
depends_on:
- gpu-monitor.prose.md (live data source on .24:9100)
- gpu-fleet.prose.md (source of truth for topology)
---
## Maintains
- gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array }
- gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail
- benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts }
- vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours }
- circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset }
## Requires
- gpu-monitor:function — Live fleet data from .24:9100/gpu-data
- Prometheus exporters on all 3 GPUs (:9400/metrics)
- SSH access to GPU hosts for restart operations
## Continuity
- Self-driven: check every 60 seconds against GPU monitor data
- Also wakes on gpu-fleet health degradation
- On fix: verify with benchmark inference test before declaring resolved
- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert
---
## Remediation Rules
### Rule 1: GPU Temperature Critical (>85°C for >2 min)
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
- **Fix**:
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains
3. If all GPUs hot, alert about cooling infrastructure
- **Verify**: Temp drops below 80°C within 5 minutes
- **Escalate after**: 3 verification failures → Zulip alert
### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
- **Detect**: VRAM growing at sustained rate over 6+ hour window
- RTX 3090 (24GB): ≥100MB/hour
- RTX 5070 (12GB): ≥50MB/hour
- Strix Halo (64GB UMA): ≥200MB/hour
- **Fix**:
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
2. If llama-server is the growth source → restart with memory cap flag
3. If unknown process → kill and alert
- **Verify**: VRAM growth rate drops below threshold
- **Escalate after**: persistent leak after restart → hardware investigation
### Rule 3: Model Inference Timeout / GPU Stuck
- **Detect**: >50% failure rate over 60s window + 30s grace period (not just single stuck request)
- **Fix**:
1. Restart llama-server on affected GPU host
2. Wait 15s for model to reload
3. Run benchmark inference test
- **Verify**: Model returns 200 with <30s response, failure rate drops to 0%
- **Escalate after**: 3 restarts in 1 hour → GPU hardware check
### Rule 4: Benchmark Regression (>20% drop)
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
- **Fix**:
1. Check GPU utilization — if >90%, other process is competing
2. Check power limit — if throttled, restore to max
3. Check thermal — if hot, apply Rule 1
- **Verify**: Benchmark returns to within 10% of baseline
- **Escalate after**: persistent regression → possible hardware degradation
### Rule 5: Circuit Breaker Stuck Open
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
- **Fix**:
1. Verify GPU /health returns 200
2. If GPU healthy, send 1 test inference
3. If test succeeds → reset circuit breaker via router API
4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset
5. Max 1 auto-reset per GPU per hour
- **Verify**: CB closes, inference succeeds, CB stays closed for 60s+
- **Escalate after**: CB won't close after reset → router issue
### Rule 6: Strix Halo Unreachable
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
- **Fix**:
1. SSH to .15 → check llama-server process
2. Restart llama-server if not running
3. Verify through both direct probe AND router
- **Verify**: Direct health probe returns 200, router reports Strix healthy
- **Escalate**: If host .15 itself is unreachable → infrastructure alert
### Rule 7: Prometheus Exporter Down
- **Detect**: Any GPU :9400/metrics unreachable for >2 polls
- **Fix**:
1. SSH to GPU host → check prometheus-exporter process
2. Restart exporter if dead
3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes
4. If exporter is running but unreachable → check firewall/host networking
- **Verify**: :9400/metrics returns 200
- **Escalate after**: 3 failed restarts → networking issue
### Rule 8: Predictive Thermal Warning (two-tier)
- **Detect**:
- Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert
- Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding
- **Fix**:
- Tier 1: silently reduce parallel requests to that GPU by 50%
- Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic
- **Verify**: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2)
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
### Rule 9: Context Window Optimization
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context
- RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
- RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
- Strix Halo (256K ctx, ornith-1.0-35b): target 70+ tok/s — currently above baseline
- **Fix**:
- If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest
- If tok/s within 10% of baseline → optimal, no change
- **Verify**: Re-benchmark after context change, confirm within 10% of target
- **Escalate**: If context can't be adjusted without significant perf loss
### Rule 10: Workload Distribution Optimization
- **Detect**: GPU roles misaligned with hardware capabilities
- **Target distribution**:
- RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations
- RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
- Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs
- **Fix**:
- Alert if any GPU is handling workload outside its designated role
- Recommend Hermes agent profile updates to match workload to GPU
- Track per-GPU request distribution via LiteLLM spend logs
- **Verify**: Each GPU's request pattern matches its designated role within 24h
- **Escalate**: If role mismatch persists >48h → agent profile audit needed
---
## Execution
```prose
-- Phase 1: Fetch live GPU data
let fleet = call gpu-monitor
endpoint: "http://192.168.68.24:9100/gpu-data"
-- Phase 2: Evaluate each GPU against remediation rules
let actions = []
for gpu in fleet.gpus:
-- Rule 1: Thermal critical
if gpu.temp_c > 85 and sustained_for(gpu, 120):
push actions apply-thermal-fix(gpu)
-- Rule 2: VRAM leak
let vram_rate = calculate-vram-trend(gpu, hours=6)
if vram_rate > 50:
push actions apply-vram-fix(gpu, vram_rate)
-- Rule 4: Benchmark regression
let bench = fleet.benchmarks[gpu.hostname]
if bench.current_tok_sec < bench.baseline_tok_sec * 0.8:
push actions apply-benchmark-fix(gpu, bench)
-- Rule 3: Model stuck
for model in fleet.router.available_models:
if model.consecutive_timeouts >= 3:
push actions apply-model-restart(model)
-- Rule 5: Circuit breaker
for cb in fleet.router.circuit_breaker:
if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu):
push actions apply-cb-reset(cb)
-- Rule 6: Strix Halo
if not fleet.strix.running and pingable("192.168.68.15"):
push actions apply-strix-restart()
-- Rule 7: Prometheus exporters
for gpu in fleet.gpus:
if not prometheus_reachable(gpu.hostname, 9400):
push actions apply-exporter-restart(gpu)
-- Rule 8: Predictive thermal
for gpu in fleet.gpus:
let rise_rate = calculate-temp-rise(gpu, minutes=5)
if rise_rate > 2.0 and gpu.temp_c < 80:
push actions apply-proactive-cooling(gpu)
-- Phase 3: Execute actions, verify, log
for action in actions:
let result = execute-with-verify(action)
log-to-kg(action, result)
if result.failed:
escalate-if-needed(action)
-- Phase 4: Update health state
call update-gpu-health
gpus: fleet.gpus
actions: actions
status: derive-overall-status(fleet, actions)
-- Wait 60s and repeat
```
## Audit Trail Format
```json
{
"run_id": "gpu-self-heal-20260712-001",
"timestamp": "2026-07-12T16:00:00Z",
"gpu": "ct8-rtx3090",
"issue": "thermal-critical",
"detected": { "temp_c": 87, "duration_s": 180 },
"action": "set-fan-100pct",
"result": "resolved",
"verification": { "temp_c": 76, "after_s": 300 },
"escalated": false
}
```
---
## Reporting
### 1. Knowledge Graph
Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail.
### 2. Zulip Alerts (#agent-hub → alerts-gpu)
- `issues_fixed > 0` → "🛠 GPU Self-Heal — <gpu> <issue> resolved"
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
### 3. Prometheus/Grafana Integration
- GPU self-heal actions exposed as Prometheus counter metrics
- Dashboard panel: "GPU Interventions (24h)" showing count/type/result
### 4. Weekly Benchmark Report
- Per-GPU tok/s trend over 7 days
- Regression alerts if any GPU degrades >10% week-over-week
---
## Design Decisions (Grilled & Confirmed — 2026-07-12)
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix).
5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU.
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana.
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.
+47 -24
View File
@@ -13,27 +13,30 @@ description: >
- template_version: "2.1.0" - template_version: "2.1.0"
- last_applied: timestamp - last_applied: timestamp
- agents_configured: ["tanko", "mumuni", "abiba", "tdunna", "baggy", "kagenz0"] - agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
- agent_keys: map (see Agent Keys section) - agent_keys: map (see Agent Keys section)
- infra_endpoints_verified: array - infra_endpoints_verified: array
## Agent Keys (LiteLLM — Current 2026-07-04) ## Agent Keys (LiteLLM — Current 2026-07-11)
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB, not in PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB.
config files. The env var `LITELLM_API_KEY` is set in `/etc/environment` on each agent The env var `LITELLM_API_KEY` is injected at runtime via `infisical run --` wrapper
host AND in `~/.hermes/.env` for gateway env propagation. (project=agents, env=production). /etc/environment and ~/.hermes/.env are NO LONGER
used for agent keys — stripped and tagged `# [INFISICAL]` post-migration.
Sub-agent profiles inherit auth from the main config — no separate keys needed. Sub-agent profiles inherit auth from the main config — no separate keys needed.
| Agent | Key Alias | Host | SSH | Sub-Agents | | Agent | Key Alias | Host | SSH | Sub-Agents |
|-------|-----------|------|-----|-----------| |-------|-----------|------|-----|-----------|
| Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — | | Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — |
| Mumuni | `mumuni-jul2026` | 192.168.68.123 | root@.123 | 6 profiles ✱ | | Mumuni | `mumuni-jul2026` | 192.168.68.123 | root@.123 | 6 profiles ✱ |
| Abiba | `abiba-*` | 192.168.68.24 | local | — | | Abiba | `abiba-pi` | 192.168.68.24 | local | — |
| Tdunna | `tdunna-*` | ? | Zulip | — | | Koby | `koby` | CT 111 (tdunna) | Zulip | — |
| Baggy | `baggy-*` | ? | Zulip | — | | Koonimo | `koonimo` | CT 113 (baggy) | Zulip | — |
| Kagenz0 | `kagenz0-*` | ? | Zulip | — | | Kagenz0 | `kagenz0-*` | ? | Zulip | — |
> CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research, ✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research,
syslog-review, syslog-writer — all at `/root/.hermes/profiles/<name>/config.yaml` syslog-review, syslog-writer — all at `/root/.hermes/profiles/<name>/config.yaml`
@@ -51,11 +54,12 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
## API Key Rules ## API Key Rules
- `api_key_env: LITELLM_API_KEY` — Use env var for main model auth (preferred) - `api_key_env: LITELLM_API_KEY` — Use env var for main model auth (preferred)
Key is injected at runtime via `infisical run --` wrapper — never in /etc/environment
- `api_key: ''` — Sub-agents leave empty to inherit from main config's custom_provider - `api_key: ''` — Sub-agents leave empty to inherit from main config's custom_provider
- `api_key: sk-...` — Hardcoded key only as fallback when env var not possible - `api_key: sk-...` — Hardcoded key only as fallback when env var not possible
- Set `LITELLM_API_KEY` in `/etc/environment` on each host - Store `LITELLM_API_KEY` in Infisical vault (project=agents, env=production)
- Sub-agents NEVER get their own key — they share the host agent's key - Sub-agents NEVER get their own key — they share the host agent's key
- Restart Hermes after updating `/etc/environment` - Restart Hermes gateway after updating vault secret (key auto-injected via wrapper)
### Sub-Agent Profiles (Mumuni pattern) ### Sub-Agent Profiles (Mumuni pattern)
@@ -79,7 +83,7 @@ Sub-agent profile rules:
5. **Auxiliary tasks** (vision, compression, etc.) also leave `api_key` empty 5. **Auxiliary tasks** (vision, compression, etc.) also leave `api_key` empty
6. **Never hardcode a key** in sub-agent profiles 6. **Never hardcode a key** in sub-agent profiles
This ensures all 6 sub-agents use the same LiteLLM key set in `/etc/environment`. This ensures all 6 sub-agents use the same LiteLLM key injected via `infisical run --` wrapper.
When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs) When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs)
work immediately after restart. work immediately after restart.
@@ -91,7 +95,7 @@ model:
default: <agent_model> # e.g., ornith-1.0-35b, qwen3.6-27B-code default: <agent_model> # e.g., ornith-1.0-35b, qwen3.6-27B-code
provider: harness provider: harness
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY # Set in /etc/environment AND ~/.hermes/.env api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 262144 # For syslog-auto (ornith route supports 256K). context_length: 262144 # For syslog-auto (ornith route supports 256K).
# Set 131072 if using qwen3.6-27B-code or gemma-4-12b directly. # Set 131072 if using qwen3.6-27B-code or gemma-4-12b directly.
@@ -176,7 +180,7 @@ custom_providers:
When LiteLLM keys are regenerated (e.g., after infrastructure changes): When LiteLLM keys are regenerated (e.g., after infrastructure changes):
1. **If SSH available**: `ssh <host> "sudo sed -i 's/LITELLM_API_KEY=.*/LITELLM_API_KEY=sk-<NEW>/' /etc/environment"` 1. **If SSH available**: Update Infisical vault: `infisical secrets set LITELLM_API_KEY=sk-<NEW> --project=agents --env=production`, then `ssh <host> "systemctl restart hermes-gateway"`
2. **If SSH unavailable**: Send Zulip DM via abiba-bot with update command 2. **If SSH unavailable**: Send Zulip DM via abiba-bot with update command
3. **After update**: Restart Hermes on the agent host 3. **After update**: Restart Hermes on the agent host
4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models` 4. **Verify**: `curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models`
@@ -198,7 +202,7 @@ The following MUST be identical across ALL profiles:
### Rule 3: API Keys via Environment ### Rule 3: API Keys via Environment
- Prefer `api_key_env: LITELLM_API_KEY` over hardcoded keys - Prefer `api_key_env: LITELLM_API_KEY` over hardcoded keys
- Hardcoded keys in config.yaml become stale after key rotation - Hardcoded keys in config.yaml become stale after key rotation
- `/etc/environment` persists across config updates - Infisical vault secrets persist across config updates / reinstalls
- Restart Hermes after env var updates - Restart Hermes after env var updates
### Rule 4: Sub-Agent Profiles Inherit Auth ### Rule 4: Sub-Agent Profiles Inherit Auth
@@ -208,7 +212,7 @@ The following MUST be identical across ALL profiles:
- Auxiliary tasks: `api_key: ''`, `provider: harness` - Auxiliary tasks: `api_key: ''`, `provider: harness`
- Never hardcode a key in sub-agent profiles - Never hardcode a key in sub-agent profiles
- When main config uses `api_key_env`, sub-agents automatically use it - When main config uses `api_key_env`, sub-agents automatically use it
- This means key rotation only touches ONE file (`/etc/environment`) - This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`)
### Rule 5: Main Config Base URL ### Rule 5: Main Config Base URL
|- Use direct IP: `http://192.168.68.116/v1` |- Use direct IP: `http://192.168.68.116/v1`
@@ -223,23 +227,42 @@ The following MUST be identical across ALL profiles:
- Apply to BOTH main config AND all sub-agent profiles - Apply to BOTH main config AND all sub-agent profiles
- For agents needing longer outputs: raise to 8192, but never omit - For agents needing longer outputs: raise to 8192, but never omit
### Rule 7: Auxiliary Model Consistency ### Rule 7: Auxiliary Model Consistency (UPDATED July 2026)
- All auxiliary services (vision, web_extract, compression) MUST use the same model: - Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
- `model: gemma-4-12b` - Compression uses `ornith-1.0-35b` (Strix Halo — 64GB, 256K ctx, compression-optimized)
- All auxiliary services MUST use identical routing:
- `base_url: http://192.168.68.116/v1` - `base_url: http://192.168.68.116/v1`
- `api_key_env: LITELLM_API_KEY` - `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes to the primary 35B reasoning GPU - **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- gemma-4-12b is a lightweight 12B model on the RTX 5070, keeping the Strix Halo free for reasoning - **Compression moved to Strix Halo (July 2026)**: The ornith-1.0-35b model on Strix Halo
- The `compression:` block's `model` MUST match `auxiliary: compression: model` — they are two different configs for the same service (64GB UMA, 256K context, 72.4 tok/s) is the designated compression GPU. This frees the
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
- The `compression: max_context_window: 262144` MUST match Strix Halo's actual capacity
### Rule 8: Compression Threshold for 256K Models ### Rule 8: GPU Workload Distribution (July 2026)
- **RTX 3090 (24GB, 256K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 131K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract
- **Strix Halo (64GB, 256K ctx, ornith-1.0-35b)**: Context compression, summarization, long docs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
- `auxiliary.compression.model: ornith-1.0-35b` (Strix Halo)
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens) - For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer - Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
- `max_context_window: 262144` MUST match the model's actual capacity - `max_context_window: 262144` MUST match the model's actual capacity
- See `devops-hermes-compression` skill for full reference - See `devops-hermes-compression` skill for full reference
### Rule 9: Default Model Must Be `syslog-auto` (All Agents) ### Rule 9: Compression Threshold for 256K Models
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer
- `max_context_window: 262144` MUST match the model's actual capacity (Strix Halo = 256K)
- See `devops-hermes-compression` skill for full reference
### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto` - **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto`
- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json` - **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json`
- `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b - `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b
@@ -250,7 +273,7 @@ The following MUST be identical across ALL profiles:
- **Exception**: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models - **Exception**: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models
for specialized tasks, but MUST validate those models exist in the key's authorized list for specialized tasks, but MUST validate those models exist in the key's authorized list
### Rule 10: Validate Model IDs Before Deployment (pi Agents) ### Rule 11: Validate Model IDs Before Deployment (pi Agents)
- After configuring a pi agent's `models.json`, verify every model ID: - After configuring a pi agent's `models.json`, verify every model ID:
```bash ```bash
curl -s http://192.168.68.116:4000/v1/models \ curl -s http://192.168.68.116:4000/v1/models \