diff --git a/gpu-self-heal.prose.md b/gpu-self-heal.prose.md new file mode 100644 index 0000000..c093178 --- /dev/null +++ b/gpu-self-heal.prose.md @@ -0,0 +1,258 @@ +--- +kind: responsibility +name: gpu-self-heal +description: > + GPU fleet self-healing — detects anomalies, applies remediation, tracks + benchmarks, and predicts failures before they happen. Extends gpu-monitor + (v2.1.0) with active remediation rules, Prometheus metrics consumption, + VRAM trend analysis, and predictive alerting. +agent: abiba +depends_on: + - gpu-monitor.prose.md (live data source on .24:9100) + - gpu-fleet.prose.md (source of truth for topology) +--- + +## Maintains + +- gpu-health: { status: "healthy"|"degraded"|"down", issues: array, actions: array } +- gpu-self-heal-log: array of { timestamp, gpu, issue, action, result } — audit trail +- benchmark-regression: { gpu, baseline_tok_sec, current_tok_sec, trend, alerts } +- vram-trend: { gpu, current_mb, rate_mb_per_hour, projected_full_in_hours } +- circuit-breaker-status: { gpu, open, auto_reset_attempted, last_reset } + +## Requires + +- gpu-monitor:function — Live fleet data from .24:9100/gpu-data +- Prometheus exporters on all 3 GPUs (:9400/metrics) +- SSH access to GPU hosts for restart operations + +## Continuity + +- Self-driven: check every 60 seconds against GPU monitor data +- Also wakes on gpu-fleet health degradation +- On fix: verify with benchmark inference test before declaring resolved +- Escalate: after 3 failed remediation attempts → Zulip #agent-hub alert + +--- + +## Remediation Rules + +### Rule 1: GPU Temperature Critical (>85°C for >2 min) +- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls +- **Fix**: + 1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control) + 2. Redirect new requests to cooler GPUs via LiteLLM fallback chains + 3. If all GPUs hot, alert about cooling infrastructure +- **Verify**: Temp drops below 80°C within 5 minutes +- **Escalate after**: 3 verification failures → Zulip alert + +### Rule 2: VRAM Leak Detection (tiered by GPU capacity) +- **Detect**: VRAM growing at sustained rate over 6+ hour window + - RTX 3090 (24GB): ≥100MB/hour + - RTX 5070 (12GB): ≥50MB/hour + - Strix Halo (64GB UMA): ≥200MB/hour +- **Fix**: + 1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux) + 2. If llama-server is the growth source → restart with memory cap flag + 3. If unknown process → kill and alert +- **Verify**: VRAM growth rate drops below threshold +- **Escalate after**: persistent leak after restart → hardware investigation + +### Rule 3: Model Inference Timeout / GPU Stuck +- **Detect**: >50% failure rate over 60s window + 30s grace period (not just single stuck request) +- **Fix**: + 1. Restart llama-server on affected GPU host + 2. Wait 15s for model to reload + 3. Run benchmark inference test +- **Verify**: Model returns 200 with <30s response, failure rate drops to 0% +- **Escalate after**: 3 restarts in 1 hour → GPU hardware check + +### Rule 4: Benchmark Regression (>20% drop) +- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks +- **Fix**: + 1. Check GPU utilization — if >90%, other process is competing + 2. Check power limit — if throttled, restore to max + 3. Check thermal — if hot, apply Rule 1 +- **Verify**: Benchmark returns to within 10% of baseline +- **Escalate after**: persistent regression → possible hardware degradation + +### Rule 5: Circuit Breaker Stuck Open +- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy +- **Fix**: + 1. Verify GPU /health returns 200 + 2. If GPU healthy, send 1 test inference + 3. If test succeeds → reset circuit breaker via router API + 4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset + 5. Max 1 auto-reset per GPU per hour +- **Verify**: CB closes, inference succeeds, CB stays closed for 60s+ +- **Escalate after**: CB won't close after reset → router issue + +### Rule 6: Strix Halo Unreachable +- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15) +- **Fix**: + 1. SSH to .15 → check llama-server process + 2. Restart llama-server if not running + 3. Verify through both direct probe AND router +- **Verify**: Direct health probe returns 200, router reports Strix healthy +- **Escalate**: If host .15 itself is unreachable → infrastructure alert + +### Rule 7: Prometheus Exporter Down +- **Detect**: Any GPU :9400/metrics unreachable for >2 polls +- **Fix**: + 1. SSH to GPU host → check prometheus-exporter process + 2. Restart exporter if dead + 3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes + 4. If exporter is running but unreachable → check firewall/host networking +- **Verify**: :9400/metrics returns 200 +- **Escalate after**: 3 failed restarts → networking issue + +### Rule 8: Predictive Thermal Warning (two-tier) +- **Detect**: + - Tier 1 (warning): temp >70°C AND rising >2°C/min → reduce concurrency, no alert + - Tier 2 (critical): temp >80°C AND still rising → full alert + aggressive load shedding +- **Fix**: + - Tier 1: silently reduce parallel requests to that GPU by 50% + - Tier 2: redirect all new requests away, alert #agent-hub, apply Rule 1 logic +- **Verify**: Temp rise rate drops below 1°C/min (Tier 1) or temp drops below 80°C (Tier 2) +- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure + +### Rule 9: Context Window Optimization +- **Detect**: Benchmark tok/s vs baseline for each GPU at current context + - RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline + - RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role + - Strix Halo (256K ctx, ornith-1.0-35b): target 70+ tok/s — currently above baseline +- **Fix**: + - If tok/s > baseline → context has headroom, consider increasing + - If tok/s < 90% baseline → reduce context by 25% and retest + - If tok/s within 10% of baseline → optimal, no change +- **Verify**: Re-benchmark after context change, confirm within 10% of target +- **Escalate**: If context can't be adjusted without significant perf loss + +### Rule 10: Workload Distribution Optimization +- **Detect**: GPU roles misaligned with hardware capabilities +- **Target distribution**: + - RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations + - RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks + - Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs +- **Fix**: + - Alert if any GPU is handling workload outside its designated role + - Recommend Hermes agent profile updates to match workload to GPU + - Track per-GPU request distribution via LiteLLM spend logs +- **Verify**: Each GPU's request pattern matches its designated role within 24h +- **Escalate**: If role mismatch persists >48h → agent profile audit needed + +--- + +## Execution + +```prose +-- Phase 1: Fetch live GPU data +let fleet = call gpu-monitor + endpoint: "http://192.168.68.24:9100/gpu-data" + +-- Phase 2: Evaluate each GPU against remediation rules +let actions = [] +for gpu in fleet.gpus: + -- Rule 1: Thermal critical + if gpu.temp_c > 85 and sustained_for(gpu, 120): + push actions apply-thermal-fix(gpu) + + -- Rule 2: VRAM leak + let vram_rate = calculate-vram-trend(gpu, hours=6) + if vram_rate > 50: + push actions apply-vram-fix(gpu, vram_rate) + + -- Rule 4: Benchmark regression + let bench = fleet.benchmarks[gpu.hostname] + if bench.current_tok_sec < bench.baseline_tok_sec * 0.8: + push actions apply-benchmark-fix(gpu, bench) + +-- Rule 3: Model stuck +for model in fleet.router.available_models: + if model.consecutive_timeouts >= 3: + push actions apply-model-restart(model) + +-- Rule 5: Circuit breaker +for cb in fleet.router.circuit_breaker: + if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu): + push actions apply-cb-reset(cb) + +-- Rule 6: Strix Halo +if not fleet.strix.running and pingable("192.168.68.15"): + push actions apply-strix-restart() + +-- Rule 7: Prometheus exporters +for gpu in fleet.gpus: + if not prometheus_reachable(gpu.hostname, 9400): + push actions apply-exporter-restart(gpu) + +-- Rule 8: Predictive thermal +for gpu in fleet.gpus: + let rise_rate = calculate-temp-rise(gpu, minutes=5) + if rise_rate > 2.0 and gpu.temp_c < 80: + push actions apply-proactive-cooling(gpu) + +-- Phase 3: Execute actions, verify, log +for action in actions: + let result = execute-with-verify(action) + log-to-kg(action, result) + if result.failed: + escalate-if-needed(action) + +-- Phase 4: Update health state +call update-gpu-health + gpus: fleet.gpus + actions: actions + status: derive-overall-status(fleet, actions) + +-- Wait 60s and repeat +``` + +## Audit Trail Format + +```json +{ + "run_id": "gpu-self-heal-20260712-001", + "timestamp": "2026-07-12T16:00:00Z", + "gpu": "ct8-rtx3090", + "issue": "thermal-critical", + "detected": { "temp_c": 87, "duration_s": 180 }, + "action": "set-fan-100pct", + "result": "resolved", + "verification": { "temp_c": 76, "after_s": 300 }, + "escalated": false +} +``` + +--- + +## Reporting + +### 1. Knowledge Graph +Every action logged as `[GPU-SELF-HEAL] ` node with full audit trail. + +### 2. Zulip Alerts (#agent-hub → alerts-gpu) +- `issues_fixed > 0` → "🛠 GPU Self-Heal — resolved" +- `issues_escalated > 0` → "⚠ GPU Self-Heal — needs attention" +- Every 100th clean cycle → "✅ GPU Fleet: All Clear" + +### 3. Prometheus/Grafana Integration +- GPU self-heal actions exposed as Prometheus counter metrics +- Dashboard panel: "GPU Interventions (24h)" showing count/type/result + +### 4. Weekly Benchmark Report +- Per-GPU tok/s trend over 7 days +- Regression alerts if any GPU degrades >10% week-over-week + +--- + +## Design Decisions (Grilled & Confirmed — 2026-07-12) + +1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect). +2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request. +3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart. +4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix). +5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU. +6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana. +7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising. +8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down. diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index 2edf9ce..0578f67 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -13,27 +13,30 @@ description: > - template_version: "2.1.0" - last_applied: timestamp -- agents_configured: ["tanko", "mumuni", "abiba", "tdunna", "baggy", "kagenz0"] +- agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"] - agent_keys: map (see Agent Keys section) - infra_endpoints_verified: array -## Agent Keys (LiteLLM — Current 2026-07-04) +## Agent Keys (LiteLLM — Current 2026-07-11) Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM -PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB, not in -config files. The env var `LITELLM_API_KEY` is set in `/etc/environment` on each agent -host AND in `~/.hermes/.env` for gateway env propagation. +PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB. +The env var `LITELLM_API_KEY` is injected at runtime via `infisical run --` wrapper +(project=agents, env=production). /etc/environment and ~/.hermes/.env are NO LONGER +used for agent keys — stripped and tagged `# [INFISICAL]` post-migration. Sub-agent profiles inherit auth from the main config — no separate keys needed. | Agent | Key Alias | Host | SSH | Sub-Agents | |-------|-----------|------|-----|-----------| | Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — | | Mumuni | `mumuni-jul2026` | 192.168.68.123 | root@.123 | 6 profiles ✱ | -| Abiba | `abiba-*` | 192.168.68.24 | local | — | -| Tdunna | `tdunna-*` | ? | Zulip | — | -| Baggy | `baggy-*` | ? | Zulip | — | +| Abiba | `abiba-pi` | 192.168.68.24 | local | — | +| Koby | `koby` | CT 111 (tdunna) | Zulip | — | +| Koonimo | `koonimo` | CT 113 (baggy) | Zulip | — | | Kagenz0 | `kagenz0-*` | ? | Zulip | — | +> CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo). + ✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research, syslog-review, syslog-writer — all at `/root/.hermes/profiles//config.yaml` @@ -51,11 +54,12 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed ## API Key Rules - `api_key_env: LITELLM_API_KEY` — Use env var for main model auth (preferred) + Key is injected at runtime via `infisical run --` wrapper — never in /etc/environment - `api_key: ''` — Sub-agents leave empty to inherit from main config's custom_provider - `api_key: sk-...` — Hardcoded key only as fallback when env var not possible -- Set `LITELLM_API_KEY` in `/etc/environment` on each host +- Store `LITELLM_API_KEY` in Infisical vault (project=agents, env=production) - Sub-agents NEVER get their own key — they share the host agent's key -- Restart Hermes after updating `/etc/environment` +- Restart Hermes gateway after updating vault secret (key auto-injected via wrapper) ### Sub-Agent Profiles (Mumuni pattern) @@ -79,7 +83,7 @@ Sub-agent profile rules: 5. **Auxiliary tasks** (vision, compression, etc.) also leave `api_key` empty 6. **Never hardcode a key** in sub-agent profiles -This ensures all 6 sub-agents use the same LiteLLM key set in `/etc/environment`. +This ensures all 6 sub-agents use the same LiteLLM key injected via `infisical run --` wrapper. When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs) work immediately after restart. @@ -91,7 +95,7 @@ model: default: # e.g., ornith-1.0-35b, qwen3.6-27B-code provider: harness base_url: http://192.168.68.116/v1 - api_key_env: LITELLM_API_KEY # Set in /etc/environment AND ~/.hermes/.env + api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation context_length: 262144 # For syslog-auto (ornith route supports 256K). # Set 131072 if using qwen3.6-27B-code or gemma-4-12b directly. @@ -176,7 +180,7 @@ custom_providers: When LiteLLM keys are regenerated (e.g., after infrastructure changes): -1. **If SSH available**: `ssh "sudo sed -i 's/LITELLM_API_KEY=.*/LITELLM_API_KEY=sk-/' /etc/environment"` +1. **If SSH available**: Update Infisical vault: `infisical secrets set LITELLM_API_KEY=sk- --project=agents --env=production`, then `ssh "systemctl restart hermes-gateway"` 2. **If SSH unavailable**: Send Zulip DM via abiba-bot with update command 3. **After update**: Restart Hermes on the agent host 4. **Verify**: `curl -H "Authorization: Bearer sk-" http://192.168.68.116/v1/models` @@ -198,7 +202,7 @@ The following MUST be identical across ALL profiles: ### Rule 3: API Keys via Environment - Prefer `api_key_env: LITELLM_API_KEY` over hardcoded keys - Hardcoded keys in config.yaml become stale after key rotation -- `/etc/environment` persists across config updates +- Infisical vault secrets persist across config updates / reinstalls - Restart Hermes after env var updates ### Rule 4: Sub-Agent Profiles Inherit Auth @@ -208,7 +212,7 @@ The following MUST be identical across ALL profiles: - Auxiliary tasks: `api_key: ''`, `provider: harness` - Never hardcode a key in sub-agent profiles - When main config uses `api_key_env`, sub-agents automatically use it -- This means key rotation only touches ONE file (`/etc/environment`) +- This means key rotation only touches ONE vault secret (`LITELLM_API_KEY`) ### Rule 5: Main Config Base URL |- Use direct IP: `http://192.168.68.116/v1` @@ -223,23 +227,42 @@ The following MUST be identical across ALL profiles: - Apply to BOTH main config AND all sub-agent profiles - For agents needing longer outputs: raise to 8192, but never omit -### Rule 7: Auxiliary Model Consistency -- All auxiliary services (vision, web_extract, compression) MUST use the same model: - - `model: gemma-4-12b` +### Rule 7: Auxiliary Model Consistency (UPDATED July 2026) +- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized) +- Compression uses `ornith-1.0-35b` (Strix Halo — 64GB, 256K ctx, compression-optimized) +- All auxiliary services MUST use identical routing: - `base_url: http://192.168.68.116/v1` - `api_key_env: LITELLM_API_KEY` -- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes to the primary 35B reasoning GPU -- gemma-4-12b is a lightweight 12B model on the RTX 5070, keeping the Strix Halo free for reasoning -- The `compression:` block's `model` MUST match `auxiliary: compression: model` — they are two different configs for the same service +- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably +- **Compression moved to Strix Halo (July 2026)**: The ornith-1.0-35b model on Strix Halo + (64GB UMA, 256K context, 72.4 tok/s) is the designated compression GPU. This frees the + RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning. +- The `compression:` block's `model` MUST match `auxiliary: compression: model` +- The `compression: max_context_window: 262144` MUST match Strix Halo's actual capacity -### Rule 8: Compression Threshold for 256K Models +### Rule 8: GPU Workload Distribution (July 2026) +- **RTX 3090 (24GB, 256K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations +- **RTX 5070 (12GB, 131K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract +- **Strix Halo (64GB, 256K ctx, ornith-1.0-35b)**: Context compression, summarization, long docs +- Agent profiles MUST route auxiliary tasks to the correct GPU: + - `auxiliary.vision.model: gemma-4-12b` (RTX 5070) + - `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070) + - `auxiliary.compression.model: ornith-1.0-35b` (Strix Halo) +- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing - For 262K context window: `threshold: 0.65` (fires at ~170K tokens) - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss - Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer - `max_context_window: 262144` MUST match the model's actual capacity - See `devops-hermes-compression` skill for full reference -### Rule 9: Default Model Must Be `syslog-auto` (All Agents) +### Rule 9: Compression Threshold for 256K Models +- For 262K context window: `threshold: 0.65` (fires at ~170K tokens) +- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss +- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer +- `max_context_window: 262144` MUST match the model's actual capacity (Strix Halo = 256K) +- See `devops-hermes-compression` skill for full reference + +### Rule 10: Default Model Must Be `syslog-auto` (All Agents) - **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto` - **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json` - `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b @@ -250,7 +273,7 @@ The following MUST be identical across ALL profiles: - **Exception**: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models for specialized tasks, but MUST validate those models exist in the key's authorized list -### Rule 10: Validate Model IDs Before Deployment (pi Agents) +### Rule 11: Validate Model IDs Before Deployment (pi Agents) - After configuring a pi agent's `models.json`, verify every model ID: ```bash curl -s http://192.168.68.116:4000/v1/models \