Compare commits
12
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6d975360d8 | ||
|
|
3e246835d5 | ||
|
|
06d2bcbc9e | ||
|
|
c4a8c45835 | ||
|
|
a74229ee74 | ||
|
|
aebc98ead6 | ||
|
|
17a77e6b3f | ||
|
|
23f3f378c5 | ||
|
|
14d27a09b5 | ||
|
|
bddbb22f03 | ||
|
|
31ec70ae36 | ||
|
|
5eb6d3bfbd |
+7
-7
@@ -125,7 +125,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
|||||||
|
|
||||||
| Alias | RPM Cap | Routes To | Purpose |
|
| Alias | RPM Cap | Routes To | Purpose |
|
||||||
|-------|---------|-----------|---------|
|
|-------|---------|-----------|---------|
|
||||||
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
|
| `strix-moe` | 40 | Strix Halo | Agent reasoning, compression (~30% via syslog-auto pool) (MoE models) |
|
||||||
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
|
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
|
||||||
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
|
| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks |
|
||||||
|
|
||||||
@@ -136,7 +136,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
|
|||||||
- syslog-auto → qwen → gemma → qwen3.6-35B-udq4
|
- syslog-auto → qwen → gemma → qwen3.6-35B-udq4
|
||||||
|
|
||||||
### Why Strix Halo RPM Is Capped
|
### Why Strix Halo RPM Is Capped
|
||||||
- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks
|
- Direct (strix-moe): 40 RPM (tight) — Strix Halo handles agent reasoning + compression via syslog-auto pool
|
||||||
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
|
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
|
||||||
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
|
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
|
||||||
|
|
||||||
@@ -284,7 +284,7 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
|
|||||||
### Stable Aliases — CRITICAL
|
### Stable Aliases — CRITICAL
|
||||||
|
|
||||||
All agent configs MUST use stable role-based aliases, never model-specific names:
|
All agent configs MUST use stable role-based aliases, never model-specific names:
|
||||||
- `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`)
|
- `compression.model: syslog-auto` — switched from strix-moe 2026-07-18 to distribute across weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070); relieves Strix Halo thermal pressure
|
||||||
- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`)
|
- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`)
|
||||||
- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`)
|
- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`)
|
||||||
- `auxiliary.web_extract.model: gpu-light`
|
- `auxiliary.web_extract.model: gpu-light`
|
||||||
@@ -295,7 +295,7 @@ When the underlying model is swapped, only the LiteLLM config changes — agent
|
|||||||
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
- RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
|
||||||
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
- **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
|
||||||
- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling)
|
- Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling)
|
||||||
- Mumuni compression model alias: `strix-moe` with 300s timeout
|
- Mumuni compression model alias: `syslog-auto` (switched from strix-moe 2026-07-18) with 300s timeout
|
||||||
|
|
||||||
### Mumuni Agent Profile
|
### Mumuni Agent Profile
|
||||||
|
|
||||||
@@ -305,8 +305,8 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
|
|||||||
|---------|-------|-------|
|
|---------|-------|-------|
|
||||||
| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) |
|
| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) |
|
||||||
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
|
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
|
||||||
| `compression.model` | `strix-moe` | Stable alias — survives model swaps |
|
| `compression.model` | `syslog-auto` | Switched from strix-moe 2026-07-18 — distributes across weighted pool |
|
||||||
| `aux.compression.model` | `strix-moe` | Compression auxiliary model |
|
| `aux.compression.model` | `syslog-auto` | Compression auxiliary — switched from strix-moe 2026-07-18 |
|
||||||
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) |
|
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) |
|
||||||
| `aux.web_extract.model` | `gpu-light` | Web extraction |
|
| `aux.web_extract.model` | `gpu-light` | Web extraction |
|
||||||
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
|
||||||
@@ -318,7 +318,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
|
|||||||
| `personalities` | `creative` | Creative assistant personality |
|
| `personalities` | `creative` | Creative assistant personality |
|
||||||
| Platforms | cli, discord, homeassistant, signal, telegram, zulip | All Hermes platforms |
|
| Platforms | cli, discord, homeassistant, signal, telegram, zulip | All Hermes platforms |
|
||||||
| Main model timeout | 300s | LiteLLM global timeout |
|
| Main model timeout | 300s | LiteLLM global timeout |
|
||||||
| Compression model timeout | 300s | strix-moe timeout increased from 120s |
|
| Compression model timeout | 300s | syslog-auto timeout (switched from strix-moe 2026-07-18) |
|
||||||
|
|
||||||
### Agent Update Status (2026-07-15)
|
### Agent Update Status (2026-07-15)
|
||||||
|
|
||||||
|
|||||||
+93
-69
@@ -6,10 +6,15 @@ description: >
|
|||||||
benchmarks, and predicts failures before they happen. Extends gpu-monitor
|
benchmarks, and predicts failures before they happen. Extends gpu-monitor
|
||||||
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
|
(v2.1.0) with active remediation rules, Prometheus metrics consumption,
|
||||||
VRAM trend analysis, and predictive alerting.
|
VRAM trend analysis, and predictive alerting.
|
||||||
|
UPDATED 2026-07-18: Model assignments synced to 2026-07-17 swaps.
|
||||||
|
Router (port 9000) references replaced with direct GPU routing.
|
||||||
|
Benchmark baselines refreshed to live values.
|
||||||
|
Prometheus exporters removed — not deployed; fall back to direct sidecar probes.
|
||||||
|
Stable role-based aliases (strix-moe, gpu-dense, gpu-light) from gpu-fleet.
|
||||||
agent: abiba
|
agent: abiba
|
||||||
depends_on:
|
depends_on:
|
||||||
- gpu-monitor.prose.md (live data source on .24:9100)
|
- gpu-monitor.prose.md (live data source on .24:9100)
|
||||||
- gpu-fleet.prose.md (source of truth for topology)
|
- gpu-fleet.prose.md (source of truth for topology, aliases, model assignments)
|
||||||
---
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -22,8 +27,8 @@ depends_on:
|
|||||||
|
|
||||||
## Requires
|
## Requires
|
||||||
|
|
||||||
- gpu-monitor:function — Live fleet data from .24:9100/gpu-data
|
- gpu-monitor:function — Live fleet data from localhost:9100/gpu-data
|
||||||
- Prometheus exporters on all 3 GPUs (:9400/metrics)
|
- Direct sidecar probe access to all GPU hosts (:8080/health)
|
||||||
- SSH access to GPU hosts for restart operations
|
- SSH access to GPU hosts for restart operations
|
||||||
|
|
||||||
## Continuity
|
## Continuity
|
||||||
@@ -35,21 +40,37 @@ depends_on:
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## Current Fleet Baseline (2026-07-18)
|
||||||
|
|
||||||
|
| Alias | GPU | Host | Model | VRAM | Ctx | tok/s | Role |
|
||||||
|
|-------|-----|------|-------|------|-----|-------|------|
|
||||||
|
| `gpu-dense` | RTX 3090 24GB | ct8 (.8:8080) | ThinkingCap Qwen3.6-27B Q4_K_M + MTP + vision | 21.6/24.6GB (88%) | 128K | 74.9 | Heavy reasoning, code gen |
|
||||||
|
| `gpu-light` | RTX 5070 12GB | ct110 (.110:8080) | HauhauCS Gemma4-12B QAT Q4_K_M + MTP draft | 10.1/12.2GB (83%) | 128K | 169.6 | Vision, web extract, light tasks |
|
||||||
|
| `strix-moe` | Strix Halo 64GB | ct15 (.15:8080) | Genesis Hermes V3 APEX (LuffyTheFox, 24GB) | ~10/64GB (16%) | 128K | 62.9 | Agent reasoning, compression (~30% via syslog-auto pool), summarization, long docs |
|
||||||
|
|
||||||
|
Key notes:
|
||||||
|
- All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path.
|
||||||
|
- Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated.
|
||||||
|
- RTX 5070 tok/s is 2.3x faster than RTX 3090 for its model — gpu-light is the fastest endpoint. Route vision/web/light work there first.
|
||||||
|
- Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads.
|
||||||
|
- RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom).
|
||||||
|
- RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes).
|
||||||
|
|
||||||
## Remediation Rules
|
## Remediation Rules
|
||||||
|
|
||||||
### Rule 1: GPU Temperature Critical (>85°C for >2 min)
|
### Rule 1: GPU Temperature Critical (>85°C for >2 min)
|
||||||
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
|
- **Detect**: Any GPU temp >85°C sustained for 2+ consecutive polls
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
|
1. Reduce inference concurrency on that GPU (load-side cooling only — NO fan control)
|
||||||
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains
|
2. Redirect new requests to cooler GPUs via LiteLLM fallback chains (gemma → qwen, qwen → gemma)
|
||||||
3. If all GPUs hot, alert about cooling infrastructure
|
3. If all GPUs hot, alert about cooling infrastructure
|
||||||
- **Verify**: Temp drops below 80°C within 5 minutes
|
- **Verify**: Temp drops below 80°C within 5 minutes
|
||||||
- **Escalate after**: 3 verification failures → Zulip alert
|
- **Escalate after**: 3 verification failures → Zulip alert
|
||||||
|
|
||||||
### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
|
### Rule 2: VRAM Leak Detection (tiered by GPU capacity)
|
||||||
- **Detect**: VRAM growing at sustained rate over 6+ hour window
|
- **Detect**: VRAM growing at sustained rate over 6+ hour window
|
||||||
- RTX 3090 (24GB): ≥100MB/hour
|
- RTX 3090 (24GB): ≥300MB/hour
|
||||||
- RTX 5070 (12GB): ≥50MB/hour
|
- RTX 5070 (12GB): ≥300MB/hour
|
||||||
- Strix Halo (64GB UMA): ≥200MB/hour
|
- Strix Halo (64GB UMA): ≥200MB/hour
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
|
1. Log VRAM snapshot with process list (nvidia-smi/rocm-smi + ps aux)
|
||||||
@@ -69,6 +90,9 @@ depends_on:
|
|||||||
|
|
||||||
### Rule 4: Benchmark Regression (>20% drop)
|
### Rule 4: Benchmark Regression (>20% drop)
|
||||||
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
|
- **Detect**: gen_tok_per_sec drops >20% below baseline over 3+ benchmarks
|
||||||
|
- RTX 3090 baseline: 74.8 tok/s → alert at <59.8 tok/s
|
||||||
|
- RTX 5070 baseline: 165.2 tok/s → alert at <132.2 tok/s
|
||||||
|
- Strix Halo baseline: 70.5 tok/s → alert at <56.4 tok/s
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. Check GPU utilization — if >90%, other process is competing
|
1. Check GPU utilization — if >90%, other process is competing
|
||||||
2. Check power limit — if throttled, restore to max
|
2. Check power limit — if throttled, restore to max
|
||||||
@@ -78,32 +102,31 @@ depends_on:
|
|||||||
|
|
||||||
### Rule 5: Circuit Breaker Stuck Open
|
### Rule 5: Circuit Breaker Stuck Open
|
||||||
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
|
- **Detect**: Circuit breaker open >10 minutes with GPU reporting healthy
|
||||||
|
- **Note**: Router (port 9000) is deprecated. If circuit breakers are reported by gpu-monitor, they come from LiteLLM's internal tracking, not the old router.
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. Verify GPU /health returns 200
|
1. Verify GPU /health returns 200 on direct port (:8080)
|
||||||
2. If GPU healthy, send 1 test inference
|
2. If GPU healthy, alert but do NOT reset via router API (deprecated)
|
||||||
3. If test succeeds → reset circuit breaker via router API
|
3. Check LiteLLM health directly: http://192.168.68.116/litellm/health/liveliness
|
||||||
4. 60s cooldown — if CB re-opens immediately, it was legitimate, do NOT re-reset
|
4. Restart LiteLLM container on CT 116 if circuit breakers are stuck
|
||||||
5. Max 1 auto-reset per GPU per hour
|
- **Verify**: LiteLLM returns healthy, circuit breaker clears within 60s
|
||||||
- **Verify**: CB closes, inference succeeds, CB stays closed for 60s+
|
- **Escalate after**: LiteLLM restart doesn't clear → human investigation
|
||||||
- **Escalate after**: CB won't close after reset → router issue
|
|
||||||
|
|
||||||
### Rule 6: Strix Halo Unreachable
|
### Rule 6: Strix Halo Unreachable
|
||||||
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
|
- **Detect**: Strix not responding — probe .15:8080 directly (firewall opened .24→.15)
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. SSH to .15 → check llama-server process
|
1. SSH to .15 → check llama-server process
|
||||||
2. Restart llama-server if not running
|
2. Restart llama-server if not running
|
||||||
3. Verify through both direct probe AND router
|
3. Verify through both direct probe AND LiteLLM health
|
||||||
- **Verify**: Direct health probe returns 200, router reports Strix healthy
|
- **Verify**: Direct health probe returns 200, LiteLLM reports model healthy
|
||||||
- **Escalate**: If host .15 itself is unreachable → infrastructure alert
|
- **Escalate**: If host .15 itself is unreachable → infrastructure alert
|
||||||
|
|
||||||
### Rule 7: Prometheus Exporter Down
|
### Rule 7: GPU Data Source Unreachable (replaces old Prometheus rule)
|
||||||
- **Detect**: Any GPU :9400/metrics unreachable for >2 polls
|
- **Detect**: gpu-monitor endpoint (localhost:9100/gpu-data) or sidecar port (:8080) on any GPU unreachable for >2 polls
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
1. SSH to GPU host → check prometheus-exporter process
|
1. If gpu-monitor is down: restart systemd service `gpu-monitor.service` on this host
|
||||||
2. Restart exporter if dead
|
2. If sidecar is down: SSH to GPU host → check llama-server process → restart systemd service
|
||||||
3. While exporter is down, fall back to nvidia-smi/rocm-smi direct probes
|
3. Fall back to direct nvidia-smi/rocm-smi probe via SSH if all API paths fail
|
||||||
4. If exporter is running but unreachable → check firewall/host networking
|
- **Verify**: gpu-monitor returns healthy + all sidecars reachable
|
||||||
- **Verify**: :9400/metrics returns 200
|
|
||||||
- **Escalate after**: 3 failed restarts → networking issue
|
- **Escalate after**: 3 failed restarts → networking issue
|
||||||
|
|
||||||
### Rule 8: Predictive Thermal Warning (two-tier)
|
### Rule 8: Predictive Thermal Warning (two-tier)
|
||||||
@@ -117,29 +140,31 @@ depends_on:
|
|||||||
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
|
- **Escalate**: If Tier 2 triggers and temp still rising after 5 min → possible hardware failure
|
||||||
|
|
||||||
### Rule 9: Context Window Optimization
|
### Rule 9: Context Window Optimization
|
||||||
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context
|
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context (all 128K)
|
||||||
- RTX 3090 (128K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
|
- RTX 3090 (128K ctx, ThinkingCap): baseline 74.8 tok/s — currently at 74.9 (100%)
|
||||||
- RTX 5070 (128K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
|
- RTX 5070 (128K ctx, HauhauCS QAT): baseline 165.2 tok/s — currently at 169.6 (103%)
|
||||||
- Strix Halo (128K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
|
- Strix Halo (128K ctx, Genesis Hermes V3): baseline 70.5 tok/s — currently at 62.9 (89%)
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
- If tok/s > baseline → context has headroom, consider increasing
|
- If tok/s > baseline → context has headroom, consider increasing
|
||||||
- If tok/s < 90% baseline → reduce context by 25% and retest
|
- If tok/s < 90% baseline → reduce context by 25% and retest
|
||||||
- If tok/s within 10% of baseline → optimal, no change
|
- If tok/s within 10% of baseline → optimal, no change
|
||||||
|
- Strix Halo at 89% of baseline → MONITOR but do not reduce yet (recent model swap may still be settling)
|
||||||
- **Verify**: Re-benchmark after context change, confirm within 10% of target
|
- **Verify**: Re-benchmark after context change, confirm within 10% of target
|
||||||
- **Escalate**: If context can't be adjusted without significant perf loss
|
- **Escalate**: If context can't be adjusted without significant perf loss
|
||||||
|
|
||||||
### Rule 10: Workload Distribution Optimization
|
### Rule 10: Workload Distribution Optimization (updated 2026-07-18)
|
||||||
- **Detect**: GPU roles misaligned with hardware capabilities
|
- **Detect**: GPU roles misaligned with hardware capabilities
|
||||||
- **Target distribution**:
|
- **Target distribution**:
|
||||||
- RTX 3090 (24GB, 128K, 75 tok/s) → Heavy reasoning, code gen, long conversations
|
- RTX 3090 (gpu-dense, 24GB, 74.9 tok/s) → Heavy reasoning, code gen, long conversations (slowest per-token but largest context capacity). Weight: 0.55 (LiteLLM).
|
||||||
- RTX 5070 (12GB, 128K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
|
- RTX 5070 (gpu-light, 12GB, 169.6 tok/s) → Vision/image, web search, lightweight tasks (2.3x faster than 3090 per token). Weight: 0.15 (LiteLLM).
|
||||||
- Strix Halo (64GB, 128K, 72 tok/s) → Context compression, summarization, long docs
|
- Strix Halo (strix-moe, 64GB, 62.9 tok/s) → Agent reasoning, compression (~30% via syslog-auto pool), summarization, long docs (MoE model). Weight: 0.30 (LiteLLM).
|
||||||
|
- **Note**: RTX 5070 is the fastest endpoint per token. Route high-volume, low-complexity work there first.
|
||||||
- **Fix**:
|
- **Fix**:
|
||||||
- Alert if any GPU is handling workload outside its designated role
|
- Alert if any GPU is handling workload outside its designated role
|
||||||
- Recommend Hermes agent profile updates to match workload to GPU
|
- Recommend agent alias updates to match workload to GPU role (use stable aliases: gpu-dense, gpu-light, strix-moe)
|
||||||
- Track per-GPU request distribution via LiteLLM spend logs
|
- Track per-GPU request distribution via LiteLLM spend logs
|
||||||
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
- **Verify**: Each GPU's request pattern matches its designated role within 24h
|
||||||
- **Escalate**: If role mismatch persists >48h → agent profile audit needed
|
- **Escalate**: If role mismatch persists >48h → agent alias audit needed
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -148,7 +173,7 @@ depends_on:
|
|||||||
```prose
|
```prose
|
||||||
-- Phase 1: Fetch live GPU data
|
-- Phase 1: Fetch live GPU data
|
||||||
let fleet = call gpu-monitor
|
let fleet = call gpu-monitor
|
||||||
endpoint: "http://192.168.68.24:9100/gpu-data"
|
endpoint: "http://localhost:9100/gpu-data"
|
||||||
|
|
||||||
-- Phase 2: Evaluate each GPU against remediation rules
|
-- Phase 2: Evaluate each GPU against remediation rules
|
||||||
let actions = []
|
let actions = []
|
||||||
@@ -164,27 +189,25 @@ for gpu in fleet.gpus:
|
|||||||
|
|
||||||
-- Rule 4: Benchmark regression
|
-- Rule 4: Benchmark regression
|
||||||
let bench = fleet.benchmarks[gpu.hostname]
|
let bench = fleet.benchmarks[gpu.hostname]
|
||||||
if bench.current_tok_sec < bench.baseline_tok_sec * 0.8:
|
if bench.current_tok_s < bench.baseline_tok_s * 0.8:
|
||||||
push actions apply-benchmark-fix(gpu, bench)
|
push actions apply-benchmark-fix(gpu, bench)
|
||||||
|
|
||||||
-- Rule 3: Model stuck
|
-- Rule 3: Model stuck
|
||||||
for model in fleet.router.available_models:
|
for model in fleet.summary.available_models:
|
||||||
if model.consecutive_timeouts >= 3:
|
if model.consecutive_timeouts >= 3:
|
||||||
push actions apply-model-restart(model)
|
push actions apply-model-restart(model)
|
||||||
|
|
||||||
-- Rule 5: Circuit breaker
|
-- Rule 5: Circuit breaker check via LiteLLM (router deprecated)
|
||||||
for cb in fleet.router.circuit_breaker:
|
if fleet.summary.circuit_breakers_open > 0:
|
||||||
if cb.open and cb.open_duration > 600 and gpu_is_healthy(cb.gpu):
|
push actions check-litellm-circuit-breakers()
|
||||||
push actions apply-cb-reset(cb)
|
|
||||||
|
|
||||||
-- Rule 6: Strix Halo
|
-- Rule 6: Strix Halo
|
||||||
if not fleet.strix.running and pingable("192.168.68.15"):
|
if not fleet.strix.running and pingable("192.168.68.15"):
|
||||||
push actions apply-strix-restart()
|
push actions apply-strix-restart()
|
||||||
|
|
||||||
-- Rule 7: Prometheus exporters
|
-- Rule 7: GPU data source
|
||||||
for gpu in fleet.gpus:
|
if not fleet.gpus or len(fleet.gpus) < 2:
|
||||||
if not prometheus_reachable(gpu.hostname, 9400):
|
push actions check-gpu-monitor-service()
|
||||||
push actions apply-exporter-restart(gpu)
|
|
||||||
|
|
||||||
-- Rule 8: Predictive thermal
|
-- Rule 8: Predictive thermal
|
||||||
for gpu in fleet.gpus:
|
for gpu in fleet.gpus:
|
||||||
@@ -212,12 +235,12 @@ call update-gpu-health
|
|||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
"run_id": "gpu-self-heal-20260712-001",
|
"run_id": "gpu-self-heal-20260718-001",
|
||||||
"timestamp": "2026-07-12T16:00:00Z",
|
"timestamp": "2026-07-18T08:00:00Z",
|
||||||
"gpu": "ct8-rtx3090",
|
"gpu": "ct8-rtx3090",
|
||||||
"issue": "thermal-critical",
|
"issue": "thermal-critical",
|
||||||
"detected": { "temp_c": 87, "duration_s": 180 },
|
"detected": { "temp_c": 87, "duration_s": 180 },
|
||||||
"action": "set-fan-100pct",
|
"action": "load-shedding",
|
||||||
"result": "resolved",
|
"result": "resolved",
|
||||||
"verification": { "temp_c": 76, "after_s": 300 },
|
"verification": { "temp_c": 76, "after_s": 300 },
|
||||||
"escalated": false
|
"escalated": false
|
||||||
@@ -236,52 +259,53 @@ Every action logged as `[GPU-SELF-HEAL] <run_id>` node with full audit trail.
|
|||||||
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
|
- `issues_escalated > 0` → "⚠ GPU Self-Heal — <gpu> needs attention"
|
||||||
- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
|
- Every 100th clean cycle → "✅ GPU Fleet: All Clear"
|
||||||
|
|
||||||
### 3. Prometheus/Grafana Integration
|
### 3. Weekly Benchmark Report
|
||||||
- GPU self-heal actions exposed as Prometheus counter metrics
|
|
||||||
- Dashboard panel: "GPU Interventions (24h)" showing count/type/result
|
|
||||||
|
|
||||||
### 4. Weekly Benchmark Report
|
|
||||||
- Per-GPU tok/s trend over 7 days
|
- Per-GPU tok/s trend over 7 days
|
||||||
- Regression alerts if any GPU degrades >10% week-over-week
|
- Regression alerts if any GPU degrades >10% week-over-week
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Design Decisions (Grilled & Confirmed — 2026-07-12)
|
## Design Decisions (Verified 2026-07-12, Reaffirmed 2026-07-18)
|
||||||
|
|
||||||
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
|
1. **Fan control**: ❌ NO auto fan control. Load-side cooling only (reduce concurrency, redirect).
|
||||||
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
|
2. **Model restart**: ✅ Only if >50% failure rate over 60s + 30s grace period. Not on single stuck request.
|
||||||
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
|
3. **Strix direct access**: ✅ Open firewall .15:8080 → .24 for direct health probe + restart.
|
||||||
4. **VRAM thresholds**: Tiered — 100MB/h (RTX 3090), 50MB/h (RTX 5070), 200MB/h (Strix).
|
4. **VRAM thresholds**: Tiered — **300MB/h** (RTX 3090), **300MB/h** (RTX 5070), 200MB/h (Strix). Previous values (100/50) were too sensitive; raised 2026-07-18 based on operational data.
|
||||||
5. **CB auto-reset**: ✅ With rate limit — 1 test inference + 60s cooldown + max 1/hour per GPU.
|
5. **CB auto-reset**: ✅ Router deprecated — circuit breakers go through LiteLLM health check + container restart if needed. No per-GPU auto-reset.
|
||||||
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Original baseline kept in Grafana.
|
6. **Benchmark baseline**: Rolling 30-day average, recalculated weekly. Current baselines live in gpu-monitor.
|
||||||
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
|
7. **Predictive alerts**: Two-tier — warn at >70°C+rising (>2°C/min), critical at >80°C+rising.
|
||||||
8. **Prometheus**: Primary source. Fall back to nvidia-smi/rocm-smi direct probes if exporter down.
|
8. **Prometheus**: ❌ Not deployed. Use direct sidecar probes (:8080/health) and gpu-monitor API. Prometheus integration deferred until exporters are running on GPU hosts.
|
||||||
|
|
||||||
## Lessons Learned (2026-07-12)
|
## Lessons Learned (2026-07-12, Updated 2026-07-18)
|
||||||
|
|
||||||
### L1: API Key Standardization Is Critical
|
### L1: API Key Standardization Is Critical
|
||||||
- All GPU llama-servers MUST use the same api-key as the LiteLLM config.
|
- All GPU llama-servers MUST use the same api-key as the LiteLLM config.
|
||||||
- RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`.
|
- RTX 5070 had `--api-key sk-loc...5678` while LiteLLM sent `not-needed`.
|
||||||
This caused cascading 401 → fallback → timeout → 401 loops, burning all retries.
|
This caused cascading 401 → fallback → timeout → 401 loops.
|
||||||
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config.
|
- **Rule**: Any new GPU or model restart MUST verify api-key matches LiteLLM config (`not-needed` for direct routing).
|
||||||
|
|
||||||
### L2: Fallback Chain Cascading Failures
|
### L2: Fallback Chain Cascading Failures
|
||||||
- When one model returns 401 (auth) and another is slow (timeout), the fallback
|
- When one model returns 401 (auth) and another is slow (timeout), the fallback
|
||||||
chain creates an infinite loop: gemma 401 → qwen timeout → gemma 401 → ...
|
chain creates an infinite loop.
|
||||||
- **Rule**: If a model returns 401 (auth error), do NOT fall back to it again.
|
- **Rule**: If a model returns 401 (auth error), do NOT fall back to it again.
|
||||||
Mark it as permanently failed for this request.
|
Mark it as permanently failed for this request.
|
||||||
|
|
||||||
### L3: Verify Running State, Not Docs
|
### L3: Verify Running State, Not Docs
|
||||||
- RTX 3090 was documented at 128K context. Actually running at 256K.
|
- RTX 3090 was documented at 128K context. Running at 128K (verified 2026-07-18).
|
||||||
- Parallel count wrong (docs said 2, actual is 1 on RTX 3090).
|
- Parallel count: 1 on both RTX 3090 and RTX 5070 (matches docs for current models).
|
||||||
- **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts.
|
- **Rule**: Before making decisions, check `/proc/PID/cmdline` on GPU hosts.
|
||||||
|
|
||||||
### L4: Infisical Is Not Always Available
|
### L4: Infisical Is Not Always Available
|
||||||
- Tanko's Infisical service token was 404 — gateway ran without API key for hours.
|
- Keep a local `.env` fallback for `LITELLM_API_KEY`.
|
||||||
- **Rule**: Always keep a local `.env` fallback for `LITELLM_API_KEY`.
|
- **Rule**: Always verify credential source is reachable before relying on it.
|
||||||
- Contract hermes-config-template Rule 3 updated.
|
|
||||||
|
|
||||||
### L5: Zulip Event Queue Can Silently Die
|
### L5: GPU Monitor Response Size Can Cause Self-Heal Crash
|
||||||
- Mumuni's queue accumulated 41 errors/reconnects then stopped polling.
|
- gpu-self-heal crashed with KeyboardInterrupt during json.loads() of 20MB response.
|
||||||
Gateway was running but ignoring all messages.
|
- Root cause: router poll returns accumulated data → cache balloons.
|
||||||
- **Rule**: litellm-health-check now monitors gateway responsiveness via Zulip API.
|
- **Rule**: Self-heal must enforce a read timeout AND max response size on every poll.
|
||||||
|
If monitor response > 1MB, log a warning and skip the cycle rather than crashing.
|
||||||
|
|
||||||
|
### L6: Stable Aliases Replace Model Names
|
||||||
|
- gpu-fleet introduced stable aliases (strix-moe, gpu-dense, gpu-light) on 2026-07-15.
|
||||||
|
- Self-heal must use aliases for reporting and alerting, not model-specific names.
|
||||||
|
- **Rule**: All alert messages and KG nodes use the stable alias as the GPU identifier.
|
||||||
|
|||||||
@@ -5,7 +5,7 @@ version: 1.0.0
|
|||||||
description: >
|
description: >
|
||||||
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
|
Canonical known-good baseline for all Syslog Hermes agents. Captures the exact
|
||||||
configuration state, keys, workarounds, and audit procedure. When an agent's
|
configuration state, keys, workarounds, and audit procedure. When an agent's
|
||||||
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 256K context (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (Strix Halo handles compression solo).
|
configuration goes sideways, restore from this baseline. Last verified 2026-07-16. All GPUs 128K context (RTX 3090 .8, RTX 5070 .110, Strix Halo .15). Parallel 1 fleet-wide (compression via syslog-auto pool — switched from strix-moe 2026-07-18).
|
||||||
author: Abiba (pi agent)
|
author: Abiba (pi agent)
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -193,6 +193,8 @@ Key is injected via `infisical run --` wrapper at PM2 startup:
|
|||||||
"models": [
|
"models": [
|
||||||
{ "id": "syslog-auto" },
|
{ "id": "syslog-auto" },
|
||||||
{ "id": "strix-moe" },
|
{ "id": "strix-moe" },
|
||||||
|
{ "id": "gpu-dense" },
|
||||||
|
{ "id": "gpu-light" },
|
||||||
{ "id": "qwen3.6-27B-code" },
|
{ "id": "qwen3.6-27B-code" },
|
||||||
{ "id": "gemma-4-12b" }
|
{ "id": "gemma-4-12b" }
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -5,11 +5,14 @@ description: >
|
|||||||
Standard Hermes configuration template for Syslog Solution LLC agents.
|
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||||
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||||
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
||||||
UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
|
UPDATED 2026-07-18: Compression model switched to `syslog-auto` (was `strix-moe`)
|
||||||
|
to relieve Strix Halo pressure. syslog-auto distributes compression across the
|
||||||
|
weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).
|
||||||
|
UPDATED 2026-07-16: Compression model was the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
|
||||||
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
|
which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
|
||||||
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
||||||
2026-07-16 Mumuni root-cause investigation (WAL #1300).
|
2026-07-16 Mumuni root-cause investigation (WAL #1300).
|
||||||
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
|
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo (later switched to syslog-auto 2026-07-18). RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
|
||||||
---
|
---
|
||||||
|
|
||||||
## Maintains
|
## Maintains
|
||||||
@@ -130,7 +133,9 @@ mcp_servers:
|
|||||||
# ─── Compression ───
|
# ─── Compression ───
|
||||||
compression:
|
compression:
|
||||||
enabled: true
|
enabled: true
|
||||||
model: strix-moe # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
|
model: syslog-auto # ⚠️ Switched from strix-moe 2026-07-18 to relieve Strix Halo.
|
||||||
|
# syslog-auto distributes across weighted pool (55% RTX 3090,
|
||||||
|
# 30% Strix Halo, 15% RTX 5070). All GPUs at 128K.
|
||||||
provider: harness
|
provider: harness
|
||||||
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
|
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
|
||||||
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
|
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
|
||||||
@@ -145,8 +150,9 @@ compression:
|
|||||||
# model: gpu-light # stable alias (NOT raw "gemma-4-12b")
|
# model: gpu-light # stable alias (NOT raw "gemma-4-12b")
|
||||||
# base_url: http://192.168.68.116/v1
|
# base_url: http://192.168.68.116/v1
|
||||||
# api_key_env: LITELLM_API_KEY
|
# api_key_env: LITELLM_API_KEY
|
||||||
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
|
# Compression uses syslog-auto (switched from strix-moe 2026-07-18) to distribute
|
||||||
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
|
# load across the weighted pool and relieve Strix Halo pressure.
|
||||||
|
# Vision and web_extract use gpu-light = RTX 5070 (12B).
|
||||||
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
|
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
|
||||||
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
|
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
|
||||||
# in agent configs — use the stable aliases so model swaps don't break agents.
|
# in agent configs — use the stable aliases so model swaps don't break agents.
|
||||||
@@ -166,7 +172,7 @@ auxiliary:
|
|||||||
timeout: 30
|
timeout: 30
|
||||||
compression:
|
compression:
|
||||||
provider: harness
|
provider: harness
|
||||||
model: strix-moe # MUST match compression.model above. Stable alias for Strix Halo.
|
model: syslog-auto # Switched from strix-moe 2026-07-18. Relieves Strix Halo pressure.
|
||||||
base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1
|
base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1
|
||||||
api_key_env: LITELLM_API_KEY
|
api_key_env: LITELLM_API_KEY
|
||||||
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
|
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
|
||||||
@@ -247,29 +253,30 @@ The following MUST be identical across ALL profiles:
|
|||||||
- Apply to BOTH main config AND all sub-agent profiles
|
- Apply to BOTH main config AND all sub-agent profiles
|
||||||
- For agents needing longer outputs: raise to 8192, but never omit
|
- For agents needing longer outputs: raise to 8192, but never omit
|
||||||
|
|
||||||
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
|
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-18)
|
||||||
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
|
- Vision and web_extract use `gpu-light` (stable alias, RTX 5070 — 12GB, vision-optimized)
|
||||||
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
|
- Compression now uses `syslog-auto` (switched from `strix-moe` 2026-07-18) to distribute
|
||||||
- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
|
compression load across the weighted pool (55% RTX 3090, 30% Strix Halo, 15% RTX 5070).
|
||||||
(it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
|
This relieves Strix Halo pressure while keeping compression functional on all GPUs.
|
||||||
|
- **`syslog-auto` is the valid compression model** — LiteLLM serves it as the weighted pool.
|
||||||
|
Old configs with `strix-moe` for compression should be updated to `syslog-auto`.
|
||||||
- All auxiliary services MUST use identical routing:
|
- All auxiliary services MUST use identical routing:
|
||||||
- `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`)
|
- `base_url: http://192.168.68.116/v1` (Rule 5: `/v1`, NOT `/litellm/v1`)
|
||||||
- `api_key_env: LITELLM_API_KEY`
|
- `api_key_env: LITELLM_API_KEY`
|
||||||
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
|
- **Compression via syslog-auto**: Routes through the weighted pool. Strix Halo still handles
|
||||||
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
|
~30% of compression calls (at 60 RPM via pool vs 40 RPM direct), but the bulk (55%)
|
||||||
(64GB UMA, 128K context) — the designated compression GPU. This frees the
|
goes to RTX 3090 which has ample spare capacity.
|
||||||
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
|
|
||||||
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
|
- The `compression:` block's `model` MUST match `auxiliary: compression: model`
|
||||||
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
|
- The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
|
||||||
|
|
||||||
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
|
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-18)
|
||||||
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
|
- **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations — also handles ~55% of compression via syslog-auto pool
|
||||||
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
|
- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract — handles ~15% of compression via syslog-auto pool
|
||||||
- **Strix Halo (64GB, 128K ctx, strix-moe)**: Context compression, summarization, long docs
|
- **Strix Halo (64GB, 128K ctx, Genesis Hermes V3 APEX)**: Agent reasoning, compression (~30% via syslog-auto pool), fallback for other GPUs
|
||||||
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
- Agent profiles MUST route auxiliary tasks to the correct GPU:
|
||||||
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
|
- `auxiliary.vision.model: gpu-light` (RTX 5070)
|
||||||
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
|
- `auxiliary.web_extract.model: gpu-light` (RTX 5070)
|
||||||
- `auxiliary.compression.model: strix-moe` (Strix Halo)
|
- `auxiliary.compression.model: syslog-auto` (distributed pool, switched from strix-moe 2026-07-18)
|
||||||
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
|
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
|
||||||
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
|
- For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
|
||||||
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
|
||||||
@@ -325,7 +332,8 @@ verify ALL FOUR of these against the live config. They are the only root causes
|
|||||||
|
|
||||||
One-line agent health check (run on the agent host):
|
One-line agent health check (run on the agent host):
|
||||||
```bash
|
```bash
|
||||||
PID=$(pgrep -f "python -m hermes_cli.main gateway run" | head -1)
|
# Use grep -v infisical to avoid matching the bash wrapper that contains the same string
|
||||||
|
PID=$(pgrep -f "python -m hermes_cli.main gateway run" | grep -v infisical | head -1)
|
||||||
cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | sed 's/=.*/<set>/'
|
cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | sed 's/=.*/<set>/'
|
||||||
curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $(cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | cut -d= -f2)" http://192.168.68.116/v1/models
|
curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $(cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | cut -d= -f2)" http://192.168.68.116/v1/models
|
||||||
```
|
```
|
||||||
@@ -351,9 +359,19 @@ directly (no infisical). Apply with `systemctl daemon-reload && systemctl restar
|
|||||||
The wrapper sources `~/.hermes/.env` then exports `LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY"`.
|
The wrapper sources `~/.hermes/.env` then exports `LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY"`.
|
||||||
See litellm-api-keys.prose.md § Machine Identity for Vault Writes for vault sync.
|
See litellm-api-keys.prose.md § Machine Identity for Vault Writes for vault sync.
|
||||||
|
|
||||||
|
**⚠️ Vault empty-key guard:** If the vault stores the secret as an empty string,
|
||||||
|
the wrapper will inject an empty key and the gateway will silently get 401 errors
|
||||||
|
on all LiteLLM requests (triggering silent DeepSeek fallback). The `.env` fallback
|
||||||
|
is present but the vault takes precedence when the secret key exists (even if empty).
|
||||||
|
|
||||||
|
**Fix:** The wrapper MUST validate the key length after injection. If LITELLM_API_KEY
|
||||||
|
is empty or shorter than 20 chars, log a warning and either fail with a clear error
|
||||||
|
message or fall back to the `.env` value before starting the gateway.
|
||||||
|
|
||||||
**Verification (all agents):**
|
**Verification (all agents):**
|
||||||
```bash
|
```bash
|
||||||
GP=$(pgrep -f "python -m hermes_cli.main gateway run" | head -1)
|
# Use grep -v infisical to avoid matching the bash wrapper that contains the same string
|
||||||
|
GP=$(pgrep -f "python -m hermes_cli.main gateway run" | grep -v infisical | head -1)
|
||||||
K=$(cat /proc/$GP/environ | tr '\0' '\n' | grep '^LITELLM_API_KEY=' | cut -d= -f2)
|
K=$(cat /proc/$GP/environ | tr '\0' '\n' | grep '^LITELLM_API_KEY=' | cut -d= -f2)
|
||||||
curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.168.68.116/v1/models # must be 200
|
curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.168.68.116/v1/models # must be 200
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -57,7 +57,7 @@ duration.
|
|||||||
prefill time at 532 tok/s. Fix context first, routing second.
|
prefill time at 532 tok/s. Fix context first, routing second.
|
||||||
|
|
||||||
- **Route by task**: ornith for multi-step reasoning only; qwen for code/standard
|
- **Route by task**: ornith for multi-step reasoning only; qwen for code/standard
|
||||||
queries; gemma for compression/auxiliary. Never send simple completion to a
|
queries; gemma for vision/web extraction; syslog-auto for compression (via weighted pool). Never send simple completion to a
|
||||||
35B MoE.
|
35B MoE.
|
||||||
- **Compress aggressively**: threshold at 40% (not 65%) — a 256K window should
|
- **Compress aggressively**: threshold at 40% (not 65%) — a 256K window should
|
||||||
compact at 102K, not 166K. Target 15% tail (not 30%).
|
compact at 102K, not 166K. Target 15% tail (not 30%).
|
||||||
|
|||||||
@@ -8,18 +8,18 @@ description: >
|
|||||||
Ensures agents never use the master key directly. Rotation is event-driven,
|
Ensures agents never use the master key directly. Rotation is event-driven,
|
||||||
not calendar-driven — rotate only on compromise, personnel change, or
|
not calendar-driven — rotate only on compromise, personnel change, or
|
||||||
periodic security hygiene (quarterly/annually).
|
periodic security hygiene (quarterly/annually).
|
||||||
|
|
||||||
UPDATED 2026-07-12: Keys are stored in Infisical vault (project=agents, env=production)
|
UPDATED 2026-07-12: Keys are stored in Infisical vault (project=agents, env=production)
|
||||||
BUT each agent host MUST keep a local .env fallback. Infisical service tokens can
|
BUT each agent host MUST keep a local .env fallback. Infisical service tokens can
|
||||||
expire/404. The .env fallback prevents agents from running without keys.
|
expire/404. The .env fallback prevents agents from running without keys.
|
||||||
Tanko incident: token 404 → gateway had no LITELLM_API_KEY for hours.
|
Tanko incident: token 404 → gateway had no LITELLM_API_KEY for hours.
|
||||||
|
|
||||||
UPDATED 2026-07-16: Vault is SYNCED (session-13 keys written to vault via abiba service
|
UPDATED 2026-07-16: Vault is SYNCED (session-13 keys written to vault via abiba service
|
||||||
token, all validate 200). Koby/Koonimo migrated from hardcoded drop-ins to the
|
token, all validate 200). Koby/Koonimo migrated from hardcoded drop-ins to the
|
||||||
infisical-gateway.sh wrapper (live vault injection). 4/5 agents now vault-backed.
|
infisical-gateway.sh wrapper (live vault injection). 4/5 agents now vault-backed.
|
||||||
Canonical process: see § Production Vault Access Process. Tanko (user jerome) pending.
|
Canonical process: see § Production Vault Access Process. Tanko (user jerome) pending.
|
||||||
Abiba's key is now a proper agent key (NOT the master key — stale note removed).
|
Abiba's key is now a proper agent key (NOT the master key — stale note removed).
|
||||||
|
|
||||||
UPDATED 2026-07-17: FLEET-WIDE STANDARDIZATION. All 4 agents (Mumuni, Tanko, Koby, Koonimo)
|
UPDATED 2026-07-17: FLEET-WIDE STANDARDIZATION. All 4 agents (Mumuni, Tanko, Koby, Koonimo)
|
||||||
standardized on a single pattern: systemd drop-in (ExecStart= reset + wrapper path) →
|
standardized on a single pattern: systemd drop-in (ExecStart= reset + wrapper path) →
|
||||||
infisical-gateway.sh while-true loop → /usr/bin/infisical run --token → bash -c key
|
infisical-gateway.sh while-true loop → /usr/bin/infisical run --token → bash -c key
|
||||||
@@ -30,7 +30,7 @@ description: >
|
|||||||
Critical lessons: (1) NEVER use shell variables inside single-quoted bash -c in wrappers
|
Critical lessons: (1) NEVER use shell variables inside single-quoted bash -c in wrappers
|
||||||
— hardcode absolute paths. (2) Drop-ins override unit file ExecStart permanently.
|
— hardcode absolute paths. (2) Drop-ins override unit file ExecStart permanently.
|
||||||
(3) Capture /proc/<pid>/environ before gateway restarts to preserve running env set.
|
(3) Capture /proc/<pid>/environ before gateway restarts to preserve running env set.
|
||||||
|
|
||||||
Current key inventory and agent list: see gpu-fleet.prose.md § Agent Keys.
|
Current key inventory and agent list: see gpu-fleet.prose.md § Agent Keys.
|
||||||
Source of truth for LiteLLM config: /opt/inference-harness/litellm_config.yaml
|
Source of truth for LiteLLM config: /opt/inference-harness/litellm_config.yaml
|
||||||
on CT 116. Last verified: 2026-07-17.
|
on CT 116. Last verified: 2026-07-17.
|
||||||
@@ -154,7 +154,7 @@ through its agent wrapper.
|
|||||||
The `ExecStart=` (empty reset) clears any ExecStart from the main unit file,
|
The `ExecStart=` (empty reset) clears any ExecStart from the main unit file,
|
||||||
then the second `ExecStart=` sets the wrapper. This drop-in **survives unit file
|
then the second `ExecStart=` sets the wrapper. This drop-in **survives unit file
|
||||||
regeneration** by `hermes gateway install` — the drop-in always wins.
|
regeneration** by `hermes gateway install` — the drop-in always wins.
|
||||||
|
|
||||||
**Why a drop-in instead of editing the unit file:** `hermes gateway install`
|
**Why a drop-in instead of editing the unit file:** `hermes gateway install`
|
||||||
(called during Hermes updates and some self-heal operations) regenerates the
|
(called during Hermes updates and some self-heal operations) regenerates the
|
||||||
systemd unit file with `ExecStart=/path/to/python -m hermes_cli.main gateway run`.
|
systemd unit file with `ExecStart=/path/to/python -m hermes_cli.main gateway run`.
|
||||||
@@ -171,7 +171,7 @@ through its agent wrapper.
|
|||||||
- **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=always` revive the gateway. Two-layer defense.
|
- **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=always` revive the gateway. Two-layer defense.
|
||||||
- **Survives Hermes updates**: systemd drop-in overrides unit file ExecStart — `hermes gateway install` cannot break the vault injection.
|
- **Survives Hermes updates**: systemd drop-in overrides unit file ExecStart — `hermes gateway install` cannot break the vault injection.
|
||||||
- **Survives reboot**: systemd user service + `loginctl enable-linger` ensures gateway starts at boot without a login session.
|
- **Survives reboot**: systemd user service + `loginctl enable-linger` ensures gateway starts at boot without a login session.
|
||||||
- **Auditable**: `cat /proc/$(pgrep hermes_cli)/environ` shows all injected keys; `infisical secrets` shows the vault source.
|
- **Auditable**: `cat /proc/$(pgrep -f 'python.*hermes_cli.main.gateway.run' | grep -v infisical | head -1)/environ` shows all injected keys (note: pipe through grep -v infisical to avoid matching the bash wrapper); `infisical secrets` shows the vault source.
|
||||||
|
|
||||||
### Migration status (2026-07-17)
|
### Migration status (2026-07-17)
|
||||||
|
|
||||||
|
|||||||
@@ -184,8 +184,8 @@ def check_agents():
|
|||||||
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
|
print(f" ⬜ {name} (CT {ct}): cannot SSH — skip liveness check")
|
||||||
continue
|
continue
|
||||||
|
|
||||||
# Gateway process
|
# Gateway process (exclude the infisical bash wrapper that contains the same string)
|
||||||
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | head -1", user=user)
|
pid = ssh(host, "pgrep -f 'hermes_cli.main gateway run' | grep -v infisical | head -1", user=user)
|
||||||
if not pid:
|
if not pid:
|
||||||
print(f" ❌ {name}: GATEWAY NOT RUNNING")
|
print(f" ❌ {name}: GATEWAY NOT RUNNING")
|
||||||
FAIL.append(f"gateway-down:{name}")
|
FAIL.append(f"gateway-down:{name}")
|
||||||
|
|||||||
@@ -409,3 +409,68 @@ Backup v2 before starting: `cp index.js index.js.v2-backup-$(date +%Y%m%d-%H%M%S
|
|||||||
| Queue expiry handling | Crash | Auto re-register |
|
| Queue expiry handling | Crash | Auto re-register |
|
||||||
| Busy worker deadlock | Router death | Worker SIGKILL + error DM |
|
| Busy worker deadlock | Router death | Worker SIGKILL + error DM |
|
||||||
| PM2 restart exhaustion | Yes (max_restarts=10) | No (max_restarts=100 + watchdog) |
|
| PM2 restart exhaustion | Yes (max_restarts=10) | No (max_restarts=100 + watchdog) |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Incident Log — 2026-07-18 Fleet-Wide Audit
|
||||||
|
|
||||||
|
### Fleet State After Audit
|
||||||
|
|
||||||
|
| Agent | Platform | Zulip State | Issues Found | Fix Applied |
|
||||||
|
|-------|----------|-------------|--------------|-------------|
|
||||||
|
| **Abiba** | pi (CT 100) | ✅ Connected | API key missing from Infisical injection; poll timeout noise | Added .env fallback; AbortError treated as empty poll (no retry); poll timeout 65s→90s |
|
||||||
|
| **Tanko** | Hermes (CT 112) | ✅ Connected | Gateway disconnected since Jul 11; watchdog restart didn't re-establish Zulip | Full gateway restart (kill wrapper, let infisical-gateway.sh respawn) |
|
||||||
|
| **Mumuni** | Hermes (CT 114) | ✅ Connected | No issues found | None needed |
|
||||||
|
|
||||||
|
### Key Fixes Applied
|
||||||
|
|
||||||
|
**1. Abiba — Credential Fallback (L4 Pattern)**
|
||||||
|
- Root cause: `zulip.api_key` in config.yaml is `""` (expected from Infisical). Infisical vault `ABIBA_ZULIP_API_KEY` wasn't being injected into the process environment.
|
||||||
|
- Fix: Added `.env` file fallback at `/root/.pi/agent/extensions/zulip/.env` with known-working key, sourced before the Infisical `exec`.
|
||||||
|
- Lesson: Per L4 from gpu-self-heal, Infisical is not always available — always keep a local `.env` fallback.
|
||||||
|
|
||||||
|
**2. Abiba — Poll Timeout Handling**
|
||||||
|
- Root cause: Zulip long-poll uses `AbortSignal.timeout(65000)`. Zulip's default `event_queue_longpoll_timeout_seconds` can exceed 65s. When the signal fires, an `AbortError` is thrown and caught by the circuit breaker as a failure.
|
||||||
|
- Fix: Caught `AbortError` inside `poll()` and return empty array (no events) instead of throwing. Extended timeout to 90s to match Zulip server default.
|
||||||
|
- Reference: [Zulip Events System — long-poll timeout](https://zulip.readthedocs.io/en/11.6/subsystems/events-system.html)
|
||||||
|
|
||||||
|
**3. Tanko — Gateway Restart**
|
||||||
|
- Root cause: Gateway process was running but Zulip platform stayed in "disconnected" state since Jul 11, 2026. The wrapper script (`infisical-gateway.sh`) restarts on crash but the gateway wasn't re-establishing Zulip on restart.
|
||||||
|
- Fix: Killed gateway PID to trigger wrapper restart. New gateway (PID 331991) established Zulip connection successfully.
|
||||||
|
|
||||||
|
### Fleet-Wide Zulip Health Metrics (as of 2026-07-18)
|
||||||
|
|
||||||
|
| Metric | Value |
|
||||||
|
|--------|-------|
|
||||||
|
| Zulip server | ✅ HTTP 200 |
|
||||||
|
| Agents connected | 3/3 (Abiba, Tanko, Mumuni) |
|
||||||
|
| Abiba circuit breaker | CLOSED (0 failures) |
|
||||||
|
| Abiba uptime | 2D (post-restart) |
|
||||||
|
| Tanko gateway uptime | Ongoing |
|
||||||
|
| Mumuni gateway uptime | Ongoing |
|
||||||
|
| Watchdog status | ✅ Online (2D uptime) |
|
||||||
|
|
||||||
|
### Hermes Agent Zulip Plugin Improvements
|
||||||
|
|
||||||
|
Based on the audit, improvements that should be ported to all Hermes Zulip adapters:
|
||||||
|
|
||||||
|
1. **Circuit breaker pattern** — Already in Abiba's pi extension. Hermes adapters should add the same CLOSED→OPEN→HALF_OPEN state machine with exponential backoff.
|
||||||
|
2. **Credential fallback** — All Hermes agents use Infisical for credentials. Add `.env` local fallback per L4 pattern for `ZULIP_API_KEY`.
|
||||||
|
3. **Queue re-registration** — Handle `BAD_EVENT_QUEUE_ID` with automatic re-registration instead of gateway restart.
|
||||||
|
4. **Supervisor watchdog** — Hermes uses PM2 which auto-restarts on crash, but has no health-check watchdog. Add lightweight external health checks.
|
||||||
|
5. **Streaming** — All agents have `streaming: true` in their zulip config. Verify `edit_message()` is implemented in each adapter.
|
||||||
|
|
||||||
|
### Abiba pi Zulip Extension v2 — Implemented Resilience Summary
|
||||||
|
|
||||||
|
| Feature | Status | Notes |
|
||||||
|
|---------|--------|-------|
|
||||||
|
| Circuit breaker | ✅ | CLOSED→OPEN→HALF_OPEN; 50% failure threshold; 30s reset timeout |
|
||||||
|
| Retry with jitter | ✅ | 2 attempts, 200ms base, 50-100% jitter |
|
||||||
|
| Queue lifecycle | ✅ | 10min idle_queue_timeout; BAD_EVENT_QUEUE_ID handling |
|
||||||
|
| Crash prevention | ✅ | uncaughtException + unhandledRejection recovery |
|
||||||
|
| Worker timeout | ✅ | 5min busy timeout → SIGKILL + error DM |
|
||||||
|
| Health endpoint | ✅ | :9200 with circuit breaker metrics |
|
||||||
|
| Echo prevention | ✅ | Dynamic bot user resolution |
|
||||||
|
| Poll timeout (AbortError) | ✅ v2.1 | Normal timeout returns [] instead of error |
|
||||||
|
| Credential fallback | ✅ v2.1 | .env file before Infisical exec |
|
||||||
|
| Provider auto-fix | ✅ | Detects reasoning_content models, switches to compatible |
|
||||||
|
|||||||
Reference in New Issue
Block a user