feat: GPU context 256K→128K fleet-wide + Genesis Hermes V3 on Strix Halo #20

Merged
jerome merged 5 commits from feat/gpu-128k-genesis-hermes-v3-20260717 into master 2026-07-17 11:21:32 +00:00
5 changed files with 304 additions and 155 deletions
+27 -23
View File
@@ -9,8 +9,12 @@ description: >
gpu-dense, gpu-light. These never change — only the underlying model does. gpu-dense, gpu-light. These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster). RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
RTX 5070 context: 131K → 256K. VRAM: 88% (10.8/12.2GB). UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Compression timeout: 300s (was 120s). Mumuni context: 128K (was 256K). Strix Halo model swapped to Genesis Hermes V3 APEX (LuffyTheFox, 24GB, uncensored,
Hermes agent fine-tune, tensor repair, multimodal with mmproj).
Instability observed near 100K at 256K. 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
agent: abiba agent: abiba
triggers: triggers:
- on model add/remove - on model add/remove
@@ -67,7 +71,7 @@ triggers:
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │ │ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │ │ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │ │ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
256K ctx │ │ 256K ctx │ │ 256K ctx │ │ Watchdog │ 128K ctx │ │ 128K ctx │ │ 128K ctx │ │ Watchdog │
│ qwen3.6 │ │ gemma-4-12b │ │ qwen3.6 │ │ Prometheus │ │ qwen3.6 │ │ gemma-4-12b │ │ qwen3.6 │ │ Prometheus │
│ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │ │ 27B-code │ │ :8080 │ │ -35B-udq4 │ │ exporter │
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │ │ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
@@ -93,9 +97,9 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | | Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------| |-------|-----|------|------|-----|----------|----------|-------------|--------|
| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | 22.2/24.6GB (90%) | **256K** 🚀 | turbo4 | 2 | default | ✅ 63 tok/s | | qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | ~17/24.6GB (70%) | **128K** | turbo4 | 2 | default | ✅ 63 tok/s |
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 10.0/12.2GB (82%) | 256K | q4_0 | 2 | 2048/1024 | ✅ healthy | | gemma-4-12b | RTX 5070 | .110 (ocu-llm) | ~7.8/12.2GB (65%) | 128K | q4_0 | 2 | 2048/1024 | ✅ healthy |
| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~9GB/64GB | 256K | q8_0 | 2 | 2048/512 | ✅ healthy | | Genesis Hermes V3 APEX | Strix Halo Vulkan | .15 (amdpve) | ~10GB/64GB | 128K | q4_0 | 1 | 4096/1024 | ✅ 65 tok/s |
## Routing Configuration (LiteLLM — July 2026) ## Routing Configuration (LiteLLM — July 2026)
@@ -104,7 +108,7 @@ but are deprecated for agent configs. Only the stable aliases survive model swap
| Model | GPU | Weight | RPM Cap | Timeout | | Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------| |-------|-----|--------|---------|---------|
| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | | qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** | | Genesis Hermes V3 APEX | Strix Halo (.15:8080) | **0.30** | 60 | **300s** |
| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** | | gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **120s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path. Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path.
@@ -113,7 +117,7 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
| Model | RPM Cap | Notes | | Model | RPM Cap | Notes |
|-------|---------|-------| |-------|---------|-------|
| qwen3.6-35B-udq4 | 40 | Tight cap — prevents Strix overload | | strix-moe (Hermes V3) | 40 | Tight cap — prevents Strix overload |
| qwen3.6-27B-code | 500 | High cap — primary workhorse | | qwen3.6-27B-code | 500 | High cap — primary workhorse |
| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s | | gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s |
@@ -128,11 +132,11 @@ Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`.
### Fallback Chains ### Fallback Chains
- gemma → qwen - gemma → qwen
- qwen → gemma - qwen → gemma
- qwen3.6-35B-udq4 → qwen → gemma - strix-moe → qwen → gemma
- syslog-auto → qwen → gemma → qwen3.6-35B-udq4 - syslog-auto → qwen → gemma → qwen3.6-35B-udq4
### Why Strix Halo RPM Is Capped ### Why Strix Halo RPM Is Capped
- Direct (qwen3.6-35B-udq4): 40 RPM (tight) — Strix Halo is shared with compression tasks - Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously - Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C - Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
@@ -243,12 +247,12 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first. - **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster(). Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
- **VRAM (2026-07-15)**: RTX 3090 at 22.2/24.6GB (90%) with **256K context** (corrected from 131K). RTX 5070 at 10.8/12.2GB (88%) with 256K context + MTP. Strix Halo at ~9GB/64GB. - **VRAM (2026-07-15)**: RTX 3090 at ~17/24.6GB (~70%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~7.8/12.2GB (~65%) with 128K context + MTP. Strix Halo at ~7GB/64GB.
- **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2). - **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2).
- **RTX 3090 config**: `-c 262144 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context corrected to 256K (2026-07-15). VRAM: 90%. Service: `/home/llmuser/llama-wrapper.sh`. - **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context reduced to 128K (2026-07-17, was 256K). VRAM: ~70%. Service: `/home/llmuser/llama-wrapper.sh`.
- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 256K context. Gen speed: 122 tok/s (was 70). VRAM: 10.8/12.2GB (88%). No draft model pre-upgrade due to VRAM constraints. Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 262144`. - **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 128K context. Gen speed: 122 tok/s. VRAM: ~7.8/12.2GB (~65%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 131072`.
- **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. - **LiteLLM timeout tuning (verified 2026-07-16 against `/opt/inference-harness/litellm_config.yaml` on CT 116)**: gemma-4-12b 120s, qwen3.6-27B-code 300s, qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s.
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080 (was `ornith-server.service`), model changed to `unsloth/Qwen3.6-35B-A3B-MTP-GGUF` (UD-Q4_K_M), alias `qwen3.6-35B-udq4`, 256K context, flash-attn + q8 KV. MTP support enabled for 1.4-2.2x faster inference. - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF` (APEX quant), alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). Hermes agent fine-tune, tensor repair (SSM layers fixed via SVD), uncensored (0/465 refusals).
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. - **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
- **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116. - **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116.
@@ -262,12 +266,12 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions.
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context | | GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context |
|-----|-------|-----------|--------------|----------|---------| |-----|-------|-----------|--------------|----------|---------|
| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **256K** | | RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | **128K** |
| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **256K** | | RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **128K** |
| Strix Halo (.15) | qwen3.6-35B-udq4 | **71** | | — | **256K** | | Strix Halo (.15) | Genesis Hermes V3 APEX | **65** | 140 | — | **128K** |
Benchmarks from 2026-07-15 verification run. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s. Benchmarks from 2026-07-17. Strix Halo swapped to Genesis Hermes V3 APEX (LuffyTheFox). RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s.
All 3 GPUs now at 256K context (2026-07-15). All 3 GPUs now at 128K context (2026-07-17, reduced from 256K for stability).
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes. Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
Degradation alerts fire at 30% (warning) and 50% (critical) below baseline. Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
@@ -288,9 +292,9 @@ All agent configs MUST use stable role-based aliases, never model-specific names
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched. When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
### Context Windows ### Context Windows
- RTX 3090: **256K** (was 131K, bumped 2026-07-15) | RTX 5070: **256K** (up from 131K) | Strix Halo: **256K** - RTX 3090: **128K** (reduced from 256K 2026-07-17) | RTX 5070: **128K** (reduced from 256K) | Strix Halo: **128K**
- **Mumuni compression context**: 128K (down from 256K) — ensures compression model doesn't timeout - **All agents**: 128K ceiling — stable margin. For >128K workloads, use external providers (deepseek)
- Compression threshold 0.65: fires at ~85K for 128K context window - Compression threshold 0.65: fires at ~85K (~43K headroom before 128K ceiling)
- Mumuni compression model alias: `strix-moe` with 300s timeout - Mumuni compression model alias: `strix-moe` with 300s timeout
### Mumuni Agent Profile ### Mumuni Agent Profile
@@ -306,7 +310,7 @@ Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile i
| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) | | `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) |
| `aux.web_extract.model` | `gpu-light` | Web extraction | | `aux.web_extract.model` | `gpu-light` | Web extraction |
| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | | `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) |
| `context.max_context_window` | 262144 (256K) | Fixed 2026-07-16 (was 131072 — caused premature compression, WAL #1300) | | `context.max_context_window` | 131072 (128K) | Reduced from 256K 2026-07-17 — stable 128K ceiling |
| `compression.threshold` | 0.65 | Triggers at ~85K | | `compression.threshold` | 0.65 | Triggers at ~85K |
| `compression.target_ratio` | 0.3 | Compresses to ~38K | | `compression.target_ratio` | 0.3 | Compresses to ~38K |
| `compression.protect_last_n` | 40 | Preserves last 40 messages | | `compression.protect_last_n` | 40 | Preserves last 40 messages |
+6 -6
View File
@@ -118,9 +118,9 @@ depends_on:
### Rule 9: Context Window Optimization ### Rule 9: Context Window Optimization
- **Detect**: Benchmark tok/s vs baseline for each GPU at current context - **Detect**: Benchmark tok/s vs baseline for each GPU at current context
- RTX 3090 (256K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline - RTX 3090 (128K ctx, qwen3.6-27B-code): target 75+ tok/s — currently at baseline
- RTX 5070 (131K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role - RTX 5070 (128K ctx, gemma-4-12b): target 76+ tok/s — optimal for vision/web role
- Strix Halo (256K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline - Strix Halo (128K ctx, strix-moe / qwen3.6-35B-udq4): target 70+ tok/s — currently above baseline
- **Fix**: - **Fix**:
- If tok/s > baseline → context has headroom, consider increasing - If tok/s > baseline → context has headroom, consider increasing
- If tok/s < 90% baseline → reduce context by 25% and retest - If tok/s < 90% baseline → reduce context by 25% and retest
@@ -131,9 +131,9 @@ depends_on:
### Rule 10: Workload Distribution Optimization ### Rule 10: Workload Distribution Optimization
- **Detect**: GPU roles misaligned with hardware capabilities - **Detect**: GPU roles misaligned with hardware capabilities
- **Target distribution**: - **Target distribution**:
- RTX 3090 (24GB, 256K, 75 tok/s) → Heavy reasoning, code gen, long conversations - RTX 3090 (24GB, 128K, 75 tok/s) → Heavy reasoning, code gen, long conversations
- RTX 5070 (12GB, 131K, 76 tok/s) → Vision/image, web search, quick lightweight tasks - RTX 5070 (12GB, 128K, 76 tok/s) → Vision/image, web search, quick lightweight tasks
- Strix Halo (64GB, 256K, 72 tok/s) → Context compression, summarization, long docs - Strix Halo (64GB, 128K, 72 tok/s) → Context compression, summarization, long docs
- **Fix**: - **Fix**:
- Alert if any GPU is handling workload outside its designated role - Alert if any GPU is handling workload outside its designated role
- Recommend Hermes agent profile updates to match workload to GPU - Recommend Hermes agent profile updates to match workload to GPU
+20 -21
View File
@@ -6,11 +6,10 @@ description: >
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
RA-H OS MCP) while keeping agent-specific API keys and model choices. RA-H OS MCP) while keeping agent-specific API keys and model choices.
UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`, UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`,
which LiteLLM does not serve). All 3 GPUs verified 256K (RTX 5070 bumped 131K→256K on Jul 15). which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability).
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
2026-07-16 Mumuni root-cause investigation (WAL #1300). 2026-07-16 Mumuni root-cause investigation (WAL #1300).
UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).
verified at 256K. Infisical .env fallback required (Rule 3/13).
--- ---
## Maintains ## Maintains
@@ -36,7 +35,7 @@ Sub-agent profiles inherit auth from the main config — no separate keys needed
| Mumuni | `mumuni` | 192.168.68.123 | root@.123 | 6 profiles ✱ | | Mumuni | `mumuni` | 192.168.68.123 | root@.123 | 6 profiles ✱ |
| Abiba | `abiba-pi` | 192.168.68.24 | local | — | | Abiba | `abiba-pi` | 192.168.68.24 | local | — |
| Koby | `koby` | CT 111 (tdunna) | Zulip | — | | Koby | `koby` | CT 111 (tdunna) | Zulip | — |
| Koonimo | `koonimo` | CT 113 (baggy) | Zulip | — | | Koonimo | `koonimo` | CT 114 (baggy) | SSH root | — |
| Kagenz0 | `kagenz0-*` | ? | Zulip | — | | Kagenz0 | `kagenz0-*` | ? | Zulip | — |
> CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo). > CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
@@ -101,8 +100,8 @@ model:
base_url: http://192.168.68.116/v1 base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 262144 # For syslog-auto (all GPUs support 256K). context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
# Set 131072 if using gemma-4-12b directly (12GB VRAM constraint). # Set 65536 if using gemma-4-12b directly (tight VRAM).
fallback_providers: fallback_providers:
provider: deepseek provider: deepseek
@@ -133,7 +132,7 @@ compression:
enabled: true enabled: true
model: strix-moe # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name). model: strix-moe # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
provider: harness provider: harness
max_context_window: 262144 # MUST match actual GPU capacity. All 3 GPUs are 256K (Jul 15). max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
target_ratio: 0.30 target_ratio: 0.30
protect_last_n: 40 protect_last_n: 40
@@ -250,7 +249,7 @@ The following MUST be identical across ALL profiles:
### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16) ### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized) - Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized)
- Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 256K ctx, compression-optimized) - Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
- **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b` - **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b`
(it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls. (it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls.
- All auxiliary services MUST use identical routing: - All auxiliary services MUST use identical routing:
@@ -258,31 +257,31 @@ The following MUST be identical across ALL profiles:
- `api_key_env: LITELLM_API_KEY` - `api_key_env: LITELLM_API_KEY`
- **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably - **Do NOT use `syslog-auto`** for auxiliary tasks — it routes unpredictably
- **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo - **Compression on Strix Halo**: The strix-moe alias routes to Strix Halo
(64GB UMA, 256K context) — the designated compression GPU. This frees the (64GB UMA, 128K context) — the designated compression GPU. This frees the
RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning. RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
- The `compression:` block's `model` MUST match `auxiliary: compression: model` - The `compression:` block's `model` MUST match `auxiliary: compression: model`
- The `compression: max_context_window: 262144` MUST match Strix Halo's actual capacity - The `compression: max_context_window: 131072` MUST match actual GPU capacity (128K)
### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16) ### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
- **RTX 3090 (24GB, 256K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations - **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations
- **RTX 5070 (12GB, 256K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (bumped 131K→256K Jul 15; IQ4_NL+MTP, 88% VRAM) - **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- **Strix Halo (64GB, 256K ctx, strix-moe)**: Context compression, summarization, long docs - **Strix Halo (64GB, 128K ctx, strix-moe)**: Context compression, summarization, long docs
- Agent profiles MUST route auxiliary tasks to the correct GPU: - Agent profiles MUST route auxiliary tasks to the correct GPU:
- `auxiliary.vision.model: gemma-4-12b` (RTX 5070) - `auxiliary.vision.model: gemma-4-12b` (RTX 5070)
- `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070) - `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070)
- `auxiliary.compression.model: strix-moe` (Strix Halo) - `auxiliary.compression.model: strix-moe` (Strix Halo)
- Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing - Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens) - For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer - Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 262144` MUST match the model's actual capacity - `max_context_window: 131072` MUST match the model's actual capacity (128K)
- See `devops-hermes-compression` skill for full reference - See `devops-hermes-compression` skill for full reference
### Rule 9: Compression Threshold for 256K Models ### Rule 9: Compression Threshold for 128K Models
- For 262K context window: `threshold: 0.65` (fires at ~170K tokens) - For 128K context window: `threshold: 0.65` (fires at ~85K tokens)
- Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss
- Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer - Do NOT use `threshold: 0.80` — this delays until ~105K, leaving only 23K margin
- `max_context_window: 262144` MUST match the model's actual capacity (Strix Halo = 256K) - `max_context_window: 131072` MUST match the model's actual capacity (all GPUs = 128K)
- See `devops-hermes-compression` skill for full reference - See `devops-hermes-compression` skill for full reference
### Rule 10: Default Model Must Be `syslog-auto` (All Agents) ### Rule 10: Default Model Must Be `syslog-auto` (All Agents)
@@ -312,7 +311,7 @@ When an agent shows "context issues" (premature compression, 401s, 504s, DeepSee
verify ALL FOUR of these against the live config. They are the only root causes found in production: verify ALL FOUR of these against the live config. They are the only root causes found in production:
1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window` 1. **max_context_window correct?** — BOTH `compression.max_context_window` AND `context.max_context_window`
MUST be `262144` (all GPUs are 256K). A value of `131072` causes premature compression at MUST be `131072` (all GPUs are 128K). A value of `262144` causes instability near 100K and must NOT be used.
~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml` ~83K instead of ~170K. Check: `grep -n max_context_window ~/.hermes/config.yaml`
2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`, 2. **base_url uses /v1 NOT /litellm/v1?** — `custom_providers[0].base_url`, `delegation.base_url`,
and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/` and ALL `auxiliary.*.base_url` MUST be `http://192.168.68.116/v1` (Rule 5). nginx `/litellm/`
+104
View File
@@ -0,0 +1,104 @@
---
name: inference-optimization
kind: responsibility
description: >
Optimizes the full Syslog inference stack — LiteLLM routing weights, GPU model
assignments, agent context management, and prompt caching — to reduce response
times to sub-15s average. All GPUs now at 128K context (stable ceiling).
id: 067NC6KP02RG60S50M40E30928
---
### Goal
Syslog inference response times reduced to sub-15s average by optimizing the full
stack: LiteLLM routing weights, GPU model assignments, Hermes agent context
management, and prompt caching — without sacrificing agent capability.
### Requires
- `inference-metrics`: current SpendLogs from CT116 LiteLLM Postgres — avg
request_duration_ms, prompt_tokens, completion_tokens, model_group breakdown,
cache_hit rate over the last 3 hours
- `agent-configs`: current config.yaml from each active Hermes agent (Mumuni
.123, any others on .129/.122) including compression, model, context_window,
prompt_caching, memory settings
- `gpu-health`: health check response from all 3 GPU backends (ornith .15:8080,
qwen .8:8080, gemma .110:8080)
### Maintains
The optimized inference stack configuration — every change is applied and
verified end-to-end. Postcondition: avg request_duration_ms ≤ 15000 for 90% of
non-ornith traffic; ≤ 30000 for ornith-bound agentic calls.
#### liteLLM-routing
The syslog-auto routing weights, model-specific timeouts, RPM limits, and
model_list entries on CT116 `/opt/inference-harness/litellm_config.yaml`.
#### agent-compression
Each Hermes agent's `~/.hermes/config.yaml` compression, context_window,
prompt_caching, and model sections.
#### prompt-caching
LiteLLM cache configuration and llama.cpp `--cache-prompt` flag on GPU hosts.
#### verification
End-to-end latency measurements after changes applied — at least 3 test
inference calls per model path measuring ttft (time-to-first-token) and total
duration.
### Continuity
- input-driven
### Strategies
**Context is the root cause.** Every ~46K prompt token costs ~87s of ornith
prefill time at 532 tok/s. Fix context first, routing second.
- **Route by task**: ornith for multi-step reasoning only; qwen for code/standard
queries; gemma for compression/auxiliary. Never send simple completion to a
35B MoE.
- **Compress aggressively**: threshold at 40% (not 65%) — a 256K window should
compact at 102K, not 166K. Target 15% tail (not 30%).
- **Cache everything repeated**: system prompts, skill docs, AGENTS.md — these
never change between turns. Single-digit cache hit rate is unacceptable.
- **Lower context ceiling**: 128K window is the stable ceiling for agent conversations.
GPUs reduced from 256K to 128K (2026-07-17). For larger contexts, route to external providers.
### Shape
- `self`: analyze metrics, compute optimal configs, apply changes, verify
- `delegates`:
- `apply-liteLLM`: update litellm_config.yaml and reload
- `apply-agent-config`: update hermes config.yaml per agent
- `verify-latency`: run test inference calls and measure response
### Execution
```prose
-- Phase 1: Analyze current state (already complete)
-- Phase 2: Apply LiteLLM routing optimization
call apply-liteLLM-routing
config_path: /opt/inference-harness/litellm_config.yaml
host: 192.168.68.116
-- Phase 3: Apply agent context compression optimization
call apply-agent-compression
agent: mumuni
host: 192.168.68.123
config_path: /root/.hermes/config.yaml
-- Phase 4: Enable llama.cpp prompt caching on GPU hosts
call enable-prompt-caching
hosts: [192.168.68.15, 192.168.68.8, 192.168.68.110]
-- Phase 5: Verify end-to-end latency
call verify-latency
host: 192.168.68.116
models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b, ornith-1.0-35b]
```
+147 -105
View File
@@ -20,9 +20,20 @@ description: >
Canonical process: see § Production Vault Access Process. Tanko (user jerome) pending. Canonical process: see § Production Vault Access Process. Tanko (user jerome) pending.
Abiba's key is now a proper agent key (NOT the master key — stale note removed). Abiba's key is now a proper agent key (NOT the master key — stale note removed).
UPDATED 2026-07-17: FLEET-WIDE STANDARDIZATION. All 4 agents (Mumuni, Tanko, Koby, Koonimo)
standardized on a single pattern: systemd drop-in (ExecStart= reset + wrapper path) →
infisical-gateway.sh while-true loop → /usr/bin/infisical run --token → bash -c key
injection → .env fallback → exec python. Systemd drop-ins are IMMUNE to hermes gateway
install which overwrites the unit file ExecStart. Infisical CLI updated to 0.43.109 on
all agents (was 0.38.0). Service token st.8e848433 shared across fleet (st.353699cd
for tanko was deleted). .env fallback on every agent protects against token loss.
Critical lessons: (1) NEVER use shell variables inside single-quoted bash -c in wrappers
— hardcode absolute paths. (2) Drop-ins override unit file ExecStart permanently.
(3) Capture /proc/<pid>/environ before gateway restarts to preserve running env set.
Current key inventory and agent list: see gpu-fleet.prose.md § Agent Keys. Current key inventory and agent list: see gpu-fleet.prose.md § Agent Keys.
Source of truth for LiteLLM config: /opt/inference-harness/litellm_config.yaml Source of truth for LiteLLM config: /opt/inference-harness/litellm_config.yaml
on CT 116. Last verified: 2026-07-16. on CT 116. Last verified: 2026-07-17.
--- ---
## Parameters ## Parameters
@@ -74,88 +85,147 @@ description: >
- Confirm key alias matches agent_name in LiteLLM key list - Confirm key alias matches agent_name in LiteLLM key list
- Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run` - Verify agent gateway uses vault wrapper: `cat /proc/<pid>/cmdline` shows `infisical run`
## Production Vault Access Process (canonical, 2026-07-16) ## Production Vault Access Process (canonical, 2026-07-17)
The non-fail approach to agentic vault access. Deployed on 4/5 agents (tanko pending — The non-fail approach to agentic vault access. Deployed on all 4 Hermes agents
runs as user `jerome`, not systemd root, needs user-scope adaptation). (Mumuni, Tanko, Koby, Koonimo) as of 2026-07-17. Abiba (pi) uses a similar pattern
through its agent wrapper.
### The canonical pattern ### The canonical pattern
1. **infisical CLI** installed on the host (`/usr/local/bin/infisical` or `/usr/bin/infisical`). 1. **infisical CLI** installed on the host at `/usr/bin/infisical` (v0.43.109+, from
2. **Service token** (Infisical Machine Identity, `st.…`) stored root-only at `/root/.infisical-token` (`chmod 600`). artifacts-cli.infisical.com apt repo). Update procedure:
- Interim: the shared `abiba` service token (`st.8e848433…`) has READ+WRITE on the `agents` project. ```bash
- Proper: one machine identity per agent (create in Infisical UI → Project Settings → Machine Identities). curl -1sLf 'https://artifacts-cli.infisical.com/setup.deb.sh' | sudo -E bash
3. **`infisical-gateway.sh` wrapper** at `/root/.hermes/infisical-gateway.sh` (`chmod 700`): sudo apt-get update && sudo apt-get install -y infisical
# Remove stale old binary if present
rm -f /usr/local/bin/infisical /bin/infisical
```
Wrappers use absolute path `/usr/bin/infisical run`. Never rely on PATH resolution.
2. **Service token** (Infisical Machine Identity, `st.…`) stored at `~/.infisical-token`
(`chmod 600`). Current: shared `st.8e848433…` (abiba, READ+WRITE on agents project).
Tanko's `st.353699cd…` (tanko-agent) was deleted — reverted to shared token.
Proper: one machine identity per agent (create in Infisical UI → Project Settings →
Machine Identities).
3. **`infisical-gateway.sh` wrapper** at `~/.hermes/infisical-gateway.sh` (`chmod 700`):
```bash ```bash
#!/bin/bash #!/bin/bash
export INFISICAL_API_URL="https://vault.sysloggh.net" export INFISICAL_API_URL="https://vault.sysloggh.net"
TOKEN=$(cat /root/.infisical-token) TOKEN=$(cat $HOME/.infisical-token)
LOG=/root/.hermes/logs/gateway.log; mkdir -p /root/.hermes/logs LOG=$HOME/.hermes/logs/gateway.log; mkdir -p $HOME/.hermes/logs
while true; do while true; do
infisical run --token="$TOKEN" --projectId=322fceab-39da-4854-a55a-568e76c0f13f \ echo "[$(date -Iseconds)] Starting gateway with Infisical injection..." >> $LOG
/usr/bin/infisical run --token="$TOKEN" \
--projectId=322fceab-39da-4854-a55a-568e76c0f13f \
--env=prod --domain=https://vault.sysloggh.net -- bash -c ' --env=prod --domain=https://vault.sysloggh.net -- bash -c '
. /root/.hermes/.env 2>/dev/null # [FALLBACK Rule 3] safety net only . $HOME/.hermes/.env 2>/dev/null # [FALLBACK Rule 3]
export LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY" export LITELLM_API_KEY="${<AGENT>_LITELLM_API_KEY}"
exec <HERMES_VENV>/bin/python -m hermes_cli.main gateway run export ZULIP_API_KEY="${<AGENT>_ZULIP_API_KEY}"
export ZULIP_SITE="https://chat.sysloggh.net"
export ZULIP_EMAIL="<agent>-bot@chat.sysloggh.net"
export SEARXNG_URL="http://192.168.68.7:8888"
# ⚠️ HARDCODE the full venv path. NEVER use $VENV inside single quotes.
exec /root/.hermes/hermes-agent/venv/bin/python -m hermes_cli.main gateway run
' >> $LOG 2>&1 ' >> $LOG 2>&1
sleep 5 # restart on exit EXIT_CODE=$?
echo "[$(date -Iseconds)] Gateway exited with code $EXIT_CODE — restarting in 5s..." >> $LOG
sleep 5
done done
``` ```
4. **Agent key in vault** as `<AGENT>_LITELLM_API_KEY` (e.g. `KOBY_LITELLM_API_KEY`). Vault = source of truth. **CRITICAL: VENV PATH.** The inner `bash -c '...'` uses single quotes. Shell
5. **`.env` fallback** at `/root/.hermes/.env` (`chmod 600`) with the same key — safety net ONLY for vault outage (Rule 3/13). Must be kept in sync on rotation. variables set in the outer wrapper are NOT expanded inside single quotes.
6. **systemd service** `hermes-gateway.service` with `ExecStart=/root/.hermes/infisical-gateway.sh`. NO `litellm-key.conf` drop-in (those hardcode keys and rot). `$VENV/bin/python` resolves to `/bin/python` (file not found). Always hardcode
7. **NEVER hardcode** LiteLLM keys in systemd drop-ins, config.yaml, or /etc/environment. The wrapper injects live from vault. the absolute path to the venv python binary.
4. **Agent key in vault** as `<AGENT>_LITELLM_API_KEY` and `<AGENT>_ZULIP_API_KEY`.
Vault = source of truth for ALL platform credentials.
5. **`.env` fallback** at `~/.hermes/.env` (`chmod 600`) with agent-specific keys —
safety net for vault outage or token revocation. Must be kept in sync on rotation.
Example:
```bash
MUMUNI_LITELLM_API_KEY=sk-OzuWsoX22Hmb3Ps3JY01gw
MUMUNI_ZULIP_API_KEY=H8dY6V7aHmWNcfgNtJaDBPZ1dGWn0Ttt
```
6. **systemd drop-in** at `~/.config/systemd/user/hermes-gateway.service.d/50-vault-wrapper.conf`:
```ini
[Service]
ExecStart=
ExecStart=/root/.hermes/infisical-gateway.sh
```
The `ExecStart=` (empty reset) clears any ExecStart from the main unit file,
then the second `ExecStart=` sets the wrapper. This drop-in **survives unit file
regeneration** by `hermes gateway install` — the drop-in always wins.
**Why a drop-in instead of editing the unit file:** `hermes gateway install`
(called during Hermes updates and some self-heal operations) regenerates the
systemd unit file with `ExecStart=/path/to/python -m hermes_cli.main gateway run`.
Editing the unit file directly is futile — it will be overwritten. The drop-in
approach explicitly resets ExecStart and sets the wrapper regardless of what the
main unit file says.
7. **NEVER hardcode** API keys in systemd drop-ins, config.yaml, or /etc/environment.
The wrapper injects live from vault at every start.
### Why this is non-fail ### Why this is non-fail
- **No rot**: keys pulled live from vault at every gateway start. Rotation = one `infisical secrets set` + `systemctl restart`. No per-host file edits. - **No rot**: keys pulled live from vault at every gateway start. Rotation = one `infisical secrets set` + `systemctl restart`. No per-host file edits.
- **Survives vault outage**: the `.env` fallback (Rule 3) keeps the gateway running if Infisical is unreachable. - **Survives vault outage**: the `.env` fallback (Rule 3) keeps the gateway running if Infisical is unreachable or the service token is revoked.
- **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=on-failure` revive the gateway. - **Survives gateway crash**: the wrapper's `while true` + systemd `Restart=always` revive the gateway. Two-layer defense.
- **Auditable**: `cat /proc/$(pgrep hermes_cli)/environ` shows the live key; `infisical secrets` shows the vault source. - **Survives Hermes updates**: systemd drop-in overrides unit file ExecStart — `hermes gateway install` cannot break the vault injection.
- **Survives reboot**: systemd user service + `loginctl enable-linger` ensures gateway starts at boot without a login session.
- **Auditable**: `cat /proc/$(pgrep hermes_cli)/environ` shows all injected keys; `infisical secrets` shows the vault source.
### Migration status (2026-07-16) ### Migration status (2026-07-17)
| Agent | Host | Pattern | Vault key | Status | | Agent | Host | Pattern | Keys | Status |
|-------|------|---------|-----------|--------| |-------|------|---------|------|--------|
| abiba | .24 | `infisical run` (pi agent wrapper, service token) | ABIBA_LITELLM_API_KEY | ✅ vault-backed | | abiba | .24 | pi agent wrapper | ABIBA_LITELLM_API_KEY + ABIBA_ZULIP_API_KEY | ✅ vault-backed |
| mumuni | .123 | infisical-gateway.sh + user-login machine identity | MUMUNI_LITELLM_API_KEY | ✅ vault-backed | | mumuni | .123 | systemd drop-in + while-true wrapper + st.8e848433 | MUMUNI_LITELLM_API_KEY + MUMUNI_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| koby | .129 | infisical-gateway.sh + service token (migrated 2026-07-16) | KOBY_LITELLM_API_KEY | ✅ vault-backed, Zulip (tanko-bot@) + Telegram | | tanko | .122 | systemd drop-in + while-true wrapper + st.8e848433 (user jerome) | TANKO_LITELLM_API_KEY + TANKO_ZULIP_API_KEY | ✅ vault-backed + .env fallback |
| koonimo | .114 | infisical-gateway.sh + service token (migrated 2026-07-16) | KOONIMO_LITELLM_API_KEY | ✅ vault-backed | | koby | .129 | systemd drop-in + while-true wrapper + st.8e848433 | KOBY_LITELLM_API_KEY, shares TANKO_ZULIP_API_KEY (tanko-bot) | ✅ vault-backed |
> **Baggy = Koonimo (CT 113).** Deleted `BAGGY_LITELLM_API_KEY` from vault 2026-07-16. Only `KOONIMO_LITELLM_API_KEY` exists — one secret per agent. | koonimo | .114 | systemd drop-in + while-true wrapper + st.8e848433 | KOONIMO_LITELLM_API_KEY + KOONIMO_ZULIP_API_KEY | ✅ vault-backed |
<<<<<<< HEAD
| tanko | .122 | infisical-gateway.sh + service token (migrated 2026-07-16) | TANKO_LITELLM_API_KEY | ✅ vault-backed |
=======
| tanko | .122 | **hardcoded in config.yaml** (runs as user jerome, not systemd) | TANKO_LITELLM_API_KEY | ⚠️ TODO: migrate to user-scope wrapper |
>>>>>>> origin/master
### Tanko migration (pending) > Tanko runs as user `jerome` — wrapper/token at `~/.hermes/infisical-gateway.sh` and
> `~/.infisical-token`. Linger enabled (`loginctl enable-linger jerome`) for boot startup.
Tanko runs the gateway as user `jerome` (not root/systemd), with the key hardcoded in ### Tanko migration (COMPLETED 2026-07-17)
`/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`, valid but not vault-sourced).
Migration: create a user-scope systemd service (`~/.config/systemd/user/hermes-gateway.service`)
with `infisical-gateway.sh` wrapper in jerome's home, token at `~/.infisical-token`, lingering
enabled (`loginctl enable-linger jerome`) so the user service runs without a login session.
### Koby migration lessons (2026-07-16) Tanko was the last agent migrated from hardcoded keys to vault wrapper.
Previously: key hardcoded in `/home/jerome/.hermes/config.yaml` (`api_key: sk-CggiHWlamQy…`)
and `zulip-env.conf` systemd drop-in. Now: user-scope systemd service with drop-in
`50-vault-wrapper.conf`, `infisical-gateway.sh` wrapper with while-true loop, token at
`~/.infisical-token`, `.env` fallback at `~/.hermes/.env`. Keys injected live from vault.
### Koby migration lessons (2026-07-16, updated 2026-07-17)
Migrated Koby from hardcoded systemd drop-in → `infisical-gateway.sh` wrapper. Migrated Koby from hardcoded systemd drop-in → `infisical-gateway.sh` wrapper.
**Two mistakes I made that broke the agent:** **Three mistakes made:**
1. **Overwrote `/root/.hermes/.env`** without backing it up. The Zulip API key only existed 1. **Overwrote `/root/.hermes/.env`** without backing it up. The Zulip API key only existed
in the running process memory — the old .env was minimal (just LiteLLM key). Zulip creds were in the running process memory — the old .env was minimal (just LiteLLM key). Zulip creds were
inherited from the pre-migration gateway env, not stored in any file. Lost on restart. inherited from the pre-migration gateway env, not stored in any file. Lost on restart.
2. **Only injected `LITELLM_API_KEY`** in the wrapper — forgot Zulip + Telegram credentials. 2. **Only injected `LITELLM_API_KEY`** in the wrapper — forgot Zulip + Telegram credentials.
Agents need ALL their platform env vars. Missing vars cause silent adapter failures. Agents need ALL their platform env vars. Missing vars cause silent adapter failures.
3. (2026-07-17 fix) **VENV variable in single-quoted bash -c**: `exec "$VENV/bin/python"`
inside single quotes resolved to `exec "/bin/python"` (file not found). Hardcoded full path.
**How Koby actually connects (2026-07-16):** **How Koby actually connects:**
- Zulip: shares **Tanko's bot** (`tanko-bot@chat.sysloggh.net`, `TANKO_ZULIP_API_KEY=5PeD6f3zo…`). - Zulip: shares **Tanko's bot** (`tanko-bot@chat.sysloggh.net`, `TANKO_ZULIP_API_KEY=5PeD6f3zo…`).
Koby doesn't have its own Zulip bot (koby-bot@ doesn't exist in the swarm config). - Telegram: token from `.env` fallback. Allowed users: 6679773481.
- Telegram: token `828640…` recovered from `.env.bak-20260603` (18KB backup from June 2026).
Allowed users: 6679773481. Home channel: 6679773481.
- Both platforms now connect through the wrapper's env injection. - Both platforms now connect through the wrapper's env injection.
**Golden rule for gateway restarts:** always `cat /proc/<pid>/environ` before killing the old **Golden rules for gateway restarts:**
process — captures the live env set. Especially important when migrating gateways between 1. Always `cat /proc/<pid>/environ` before killing the old process — captures the live env set.
injection mechanisms. 2. Hardcode venv python path in wrapper — never use variables inside single-quoted bash -c.
3. Use systemd drop-ins (not unit file edits) to override ExecStart — survives Hermes updates.
### Fleet-wide standardization lessons (2026-07-17)
After auditing all 4 agents, five systemic patterns caused repeated failures:
1. **Three incompatible startup patterns** coexisted (systemd drop-in, direct python, orphaned wrapper)
2. **Systemd unit files reverted** by `hermes gateway install` during updates
3. **VENV variable scoping** broke wrappers on Koby and Mumuni (single-quote bash -c)
4. **Service token expiry** — Tanko's `st.353699cd` was deleted from Infisical
5. **No ZULIP_API_KEY** in env on Tanko — wrapper bypassed by systemd direct python
All resolved by the canonical drop-in + while-true wrapper pattern documented above.
### Key rotation procedure (one vault operation with this standard) ### Key rotation procedure (one vault operation with this standard)
@@ -165,69 +235,41 @@ injection mechanisms.
4. Restart: `systemctl restart hermes-gateway`. The wrapper pulls the new key live. 4. Restart: `systemctl restart hermes-gateway`. The wrapper pulls the new key live.
5. Verify: `curl -H "Authorization: Bearer sk-NEW" http://192.168.68.116/v1/models` → 200. 5. Verify: `curl -H "Authorization: Bearer sk-NEW" http://192.168.68.116/v1/models` → 200.
## Machine Identity for Vault Writes (ADDED 2026-07-16, WAL #1300) ## Machine Identity for Vault Writes (UPDATED 2026-07-17)
**Problem:** The infisical CLI on agent hosts is logged in as a user session (jerome@sysloggh.com). **Current state:** Infisical CLI updated to v0.43.109 on all agents (from v0.38.0).
In CLI v0.38.0, `infisical secrets set` / `infisical export` fail with "project id missing" / "workspace The v0.38.0 bug (user-session auth fails for `secrets set`/`export`) is resolved.
key 404" — a known bug where user-session auth works for `run` but NOT for `secrets set`. The apt Service token `st.8e848433…` (abiba, READ+WRITE) can write to vault from CLI.
repo only ships 0.38.0, so `apt upgrade` does not help.
**Proper fix — Machine Identity (Infisical automation best practice):** **Proper fix — per-agent Machine Identities:**
Create a machine identity with READ+WRITE scope on the `agents` project (project_id= Create machine identities in Infisical UI → Project Settings → Machine Identities
`322fceab-39da-4854-a55a-568e76c0f13f`, env `prod`). Store client_id + client_secret securely. for each agent with READ-only scope on the `agents` project. Store client_id +
Then vault writes work from any host: client_secret per agent. Then vault writes use the shared abiba identity, and
```bash reads use per-agent identities. This eliminates the single shared token risk.
# Get a machine-identity access token
TOKEN=$(curl -fsSL -X POST https://vault.sysloggh.net/api/v1/auth/universal-auth/login \
-H 'Content-Type: application/json' \
-d '{"clientId":"<CLIENT_ID>","clientSecret":"<CLIENT_SECRET>"}' | jq -r .accessToken)
# Write a secret via REST API v3
curl -fsSL -X PATCH https://vault.sysloggh.net/api/v3/secrets/MUMUNI_LITELLM_API_KEY \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"environment":"prod","secretValue":"sk-<NEW_KEY>","workspaceId":"<WORKSPACE_ID>","type":"shared"}'
# OR via CLI: infisical secrets set --token=$TOKEN --projectId=322fceab... --env=prod ...
```
<<<<<<< HEAD
**2026-07-16 UPDATE — vault SYNCED + CLEANED.** The abiba service token (`st.8e848433…`, READ+WRITE)
can write to the vault. All session-13 rotated keys in vault and validate 200 against LiteLLM.
4 stale secrets deprecated, 5 personal creds flagged for separate project. Tanko migrated from
hardcoded keys to `infisical-gateway.sh` wrapper. Koonimo Zulip key restored.
**Machine Identity:** `8ddb9438-74fc-4ab6-bd74-929e7c47b53b` exists in Infisical UI. **Service Token Inventory (2026-07-17):**
Client secret needed to use universal auth for vault writes from automation. | Token ID | Name | Permissions | Used By | Status |
|----------|------|-------------|---------|--------|
| `st.8e848433…` | tanko-gateway | READ+WRITE | Mumuni, Tanko, Koby, Koonimo, Abiba | ✅ Active |
| `st.353699cd…` | tanko-agent | READ-only | — | ❌ Deleted from Infisical |
**Service Token Inventory (2026-07-16):** **Per-agent .env fallback inventory (2026-07-17):**
| Token ID | Name | Permissions | Used By | | Agent | .env Keys |
|----------|------|-------------|---------| |-------|-----------|
| `st.8e848433…` | tanko-gateway | READ+WRITE | Abiba, Koonimo | | Mumuni | MUMUNI_LITELLM_API_KEY, MUMUNI_ZULIP_API_KEY |
| `st.353699cd…` | tanko-agent | READ-only | Tanko | | Tanko | TANKO_LITELLM_API_KEY, TANKO_ZULIP_API_KEY |
======= | Koby | (wrapper injects from vault — .env has Telegram token) |
Creation requires the Infisical web UI (https://vault.sysloggh.net) under Project Settings → | Koonimo | KOONIMO_LITELLM_API_KEY, KOONIMO_ZULIP_API_KEY |
Machine Identities, or an admin API call. **TODO: create `abiba-automation` machine identity
and store its credentials in the vault itself (or a root-only file).**
**Interim (working now):** the `.env` fallback (hermes-config-template Rule 3/13). The
infisical-gateway.sh wrapper sources `~/.hermes/.env`, so its `<AGENT>_LITELLM_API_KEY`
overrides a stale vault value.
**2026-07-16 UPDATE — vault is now SYNCED.** The abiba service token (`st.8e848433…`, READ+WRITE)
can write to the vault, so the session-13 rotated keys (mumuni `sk-OzuWsoX2…`, koby `sk-BqRRMboTI…`,
koonimo `sk-OEK7z26n6E…`) are now in the vault as `MUMUNI_LITELLM_API_KEY` / `KOBY_LITELLM_API_KEY` /
`KOONIMO_LITELLM_API_KEY` and validate 200 against LiteLLM. The vault is the source of truth again.
Creating a dedicated `abiba-automation` machine identity (via UI) is still the proper long-term fix
so the shared service token isn't reused across hosts — but it is no longer blocking.
>>>>>>> origin/master
## Key Rotation Log ## Key Rotation Log
| Date | Agent | Action | Notes | | Date | Agent | Action | Notes |
|------|-------|--------|-------| |------|-------|--------|-------|
<<<<<<< HEAD | 2026-07-17 | fleet | standardize | All 4 agents standardized on systemd drop-in + while-true wrapper + infisical v0.43.109. Removed conflicting zulip-env.conf + litellm-key.conf drop-ins. Added .env fallbacks with ZULIP keys. WAL #1322. |
| 2026-07-16 | koonimo | add-zulip | Added KOONIMO_ZULIP_API_KEY to vault (Tt2bUL…). Updated wrapper to inject ZULIP_API_KEY + ZULIP_EMAIL. Restarted gateway → Zulip connected as koonimo-bot@ (bot_id=17). 3 platforms now. | | 2026-07-17 | tanko | fix-zulip | Added ZULIP_API_KEY to env (was missing — systemd bypassed vault). Updated wrapper from exec to while-true. Created .env fallback. Removed hardcoded zulip-env.conf drop-in. WAL #1321. |
| 2026-07-16 | tanko | migrate | Migrated from hardcoded config.yaml key to infisical-gateway.sh wrapper + service token st.353699cd… (tanko-agent). Systemd user service updated, hardcoded key drop-ins removed. Verified LITELLM_API_KEY from vault, 3 platforms connected. | | 2026-07-16 | vault | cleanup | 4 stale secrets deprecated. 5 personal creds flagged. |
| 2026-07-16 | vault | cleanup | 4 stale secrets deprecated: KAGENZ0_LITELLM_API_KEY, HERMES_OPENROUTER_KEY, LITELLM_API_KEY (generic duplicate), ZULIP_API_KEY (generic duplicate). 5 personal creds flagged for separate project. | | 2026-07-16 | koonimo | add-zulip | Added KOONIMO_ZULIP_API_KEY to vault. Wrapper injects ZULIP_API_KEY + ZULIP_EMAIL. 3 platforms. |
======= | 2026-07-16 | tanko | migrate | Migrated from hardcoded config.yaml to infisical-gateway.sh + st.353699cd. NOTE: st.353699cd later deleted — reverted to st.8e848433 on 2026-07-17. |
>>>>>>> origin/master
| 2026-07-16 | mumuni | rotate | Old key malformed (sk-_SWAl_Vu_, 47 chars, not LiteLLM format) → 401. Deleted old `mumuni` key (token 15cbca18…), generated fresh (alias `mumuni`, 7 models: syslog-auto, qwen3.6-27B-code, gemma-4-12b, strix-moe, gpu-dense, gpu-light, qwen3.6-35B-udq4). New key sk-OzuWsoX2… written to /root/.hermes/.env (Rule 3/13 fallback). Vault sync PENDING (needs machine identity). WAL #1300. | | 2026-07-16 | mumuni | rotate | Old key malformed (sk-_SWAl_Vu_, 47 chars, not LiteLLM format) → 401. Deleted old `mumuni` key (token 15cbca18…), generated fresh (alias `mumuni`, 7 models: syslog-auto, qwen3.6-27B-code, gemma-4-12b, strix-moe, gpu-dense, gpu-light, qwen3.6-35B-udq4). New key sk-OzuWsoX2… written to /root/.hermes/.env (Rule 3/13 fallback). Vault sync PENDING (needs machine identity). WAL #1300. |
| 2026-07-16 | koby | rotate | Old key sk-6sbCNjz (401, stale in /etc/environment). Deleted old `koby` key, generated fresh (alias `koby`). New key sk-BqRRMboTI… in systemd drop-in `hermes-gateway.service.d/litellm-key.conf` + /etc/environment. Created `hermes-gateway.service` unit (was missing — gateway wasn't persistent) with `--replace`. Verified HTTP 200, Telegram connected. | | 2026-07-16 | koby | rotate | Old key sk-6sbCNjz (401, stale in /etc/environment). Deleted old `koby` key, generated fresh (alias `koby`). New key sk-BqRRMboTI… in systemd drop-in `hermes-gateway.service.d/litellm-key.conf` + /etc/environment. Created `hermes-gateway.service` unit (was missing — gateway wasn't persistent) with `--replace`. Verified HTTP 200, Telegram connected. |
| 2026-07-16 | baggy (koonimo) | rotate | Old key sk-krnw_zGB (401, hardcoded in systemd drop-in). Deleted old `baggy` key, generated fresh (alias `baggy`, metadata agent=koonimo). New key sk-OEK7z26n6E… in drop-in `hermes-gateway.service.d/litellm-key.conf`. CT113 IP changed .113→.114. Verified HTTP 200, Zulip connected. | | 2026-07-16 | baggy (koonimo) | rotate | Old key sk-krnw_zGB (401, hardcoded in systemd drop-in). Deleted old `baggy` key, generated fresh (alias `baggy`, metadata agent=koonimo). New key sk-OEK7z26n6E… in drop-in `hermes-gateway.service.d/litellm-key.conf`. CT113 IP changed .113→.114. Verified HTTP 200, Zulip connected. |