GPU Changes: - All 3 GPUs reduced from 256K (-c 262144) to 128K (-c 131072) for stability - Observed instability near 100K at 256K — 128K is the stable ceiling - VRAM improved: RTX 3090 ~70% (was 90%), RTX 5070 ~65% (was 88%) - Strix Halo swapped to LuffyTheFox/Genesis Hermes V3 APEX - Hermes agent fine-tune, tensor repair (3 SSM layers, 76% W1 improvement) - Uncensored (0/465 refusals), multimodal (mmproj F16) - Speed: 65 tok/s gen, 140 tok/s prompt - Alias strix-moe maintained Agent Updates: - Mumuni: max_context_window 262144→131072, already aligned on strix-moe/0.65 - Tanko: max_context_window 262144→131072 - Koonimo: max_context_window + context_length 262144→131072 - CT114 SSH access confirmed (was 'Zulip only') LiteLLM (CT116): - Updated backend model references qwen3.6-35B-udq4→strix-moe - Removed stale ornith-1.0-35b from model_cost - Fallback chains updated Contracts Updated: - gpu-fleet.prose.md: topology, VRAM, benchmarks, config lines, model assignments - gpu-self-heal.prose.md: Rule 9/10 context targets - hermes-config-template.prose.md: template values, Rules 7-9, compression thresholds - inference-optimization.prose.md: added to repo, 128K recommendation Compression: 0.65 fires at ~85K (~43K headroom before 128K ceiling) For >128K workloads: route to external providers (deepseek)
19 KiB
kind, name, description
| kind | name | description |
|---|---|---|
| template | hermes-config-template | Standard Hermes configuration template for Syslog Solution LLC agents. Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, RA-H OS MCP) while keeping agent-specific API keys and model choices. UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`, which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability). Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the 2026-07-16 Mumuni root-cause investigation (WAL #1300). UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13). |
Maintains
- template_version: "2.1.0"
- last_applied: timestamp
- agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
- agent_keys: map (see Agent Keys section)
- infra_endpoints_verified: array
Agent Keys (LiteLLM — Current 2026-07-11)
Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM
PostgreSQL DB via POST /key/generate on CT 116. Keys are stored in the DB.
The env var LITELLM_API_KEY is injected at runtime via infisical run -- wrapper
(project=agents, env=production). /etc/environment and ~/.hermes/.env are NO LONGER
used for agent keys — stripped and tagged # [INFISICAL] post-migration.
Sub-agent profiles inherit auth from the main config — no separate keys needed.
| Agent | Key Alias | Host | SSH | Sub-Agents |
|---|---|---|---|---|
| Tanko | tanko-* |
192.168.68.122 | jerome@.122 | — |
| Mumuni | mumuni |
192.168.68.123 | root@.123 | 6 profiles ✱ |
| Abiba | abiba-pi |
192.168.68.24 | local | — |
| Koby | koby |
CT 111 (tdunna) | Zulip | — |
| Koonimo | koonimo |
CT 114 (baggy) | SSH root | — |
| Kagenz0 | kagenz0-* |
? | Zulip | — |
CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).
✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research,
syslog-review, syslog-writer — all at /root/.hermes/profiles/<name>/config.yaml
Infrastructure Stack
| Component | Endpoint | Purpose |
|---|---|---|
| Firecrawl | http://192.168.68.7:3002/ |
Web content extraction |
| SearXNG | http://storepve:8888 |
Privacy-respecting web search |
| LiteLLM | http://192.168.68.116/v1 |
Unified model gateway (via nginx) |
| LiteLLM (NetBird) | https://litellm.sysloggh.net/v1 |
Alternative (may have 502 issues) |
| RA-H OS MCP | http://192.168.68.65:3100/mcp |
Knowledge graph bridge |
| Context7 MCP | http://localhost:8079/mcp |
Documentation queries |
API Key Rules
api_key_env: LITELLM_API_KEY— Use env var for main model auth (preferred) Key is injected at runtime viainfisical run --wrapper — never in /etc/environmentapi_key: ''— Sub-agents leave empty to inherit from main config's custom_providerapi_key: sk-...— Hardcoded key only as fallback when env var not possible- Store
LITELLM_API_KEYin Infisical vault (project=agents, env=production) - Sub-agents NEVER get their own key — they share the host agent's key
- Restart Hermes gateway after updating vault secret (key auto-injected via wrapper)
Sub-Agent Profiles (Mumuni pattern)
Mumuni has 6 sub-agent profiles in /root/.hermes/profiles/<name>/config.yaml:
profiles/
├── syslog-code/config.yaml # Code generation
├── syslog-devops/config.yaml # DevOps/infrastructure
├── syslog-email/config.yaml # Email processing
├── syslog-research/config.yaml # Research & analysis
├── syslog-review/config.yaml # Code review
└── syslog-writer/config.yaml # Content writing
Sub-agent profile rules:
api_keymust be empty —api_key: ''or omitted entirelybase_urlmust be empty — inherits from main config's custom_providerproviderisautoorharness— routes through the shared LiteLLM gatewaymodelis agent-specific — each sub-agent can have its own default model- Auxiliary tasks (vision, compression, etc.) also leave
api_keyempty - Never hardcode a key in sub-agent profiles
This ensures all 6 sub-agents use the same LiteLLM key injected via infisical run -- wrapper.
When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs)
work immediately after restart.
Template — Required Sections
# ─── Model Selection ───
model:
default: <agent_model> # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
provider: harness
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper
max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation
context_length: 131072 # For syslog-auto (all GPUs at 128K for stability).
# Set 65536 if using gemma-4-12b directly (tight VRAM).
fallback_providers:
provider: deepseek
model: deepseek-chat
api_key_env: DEEPSEEK_API_KEY
# ─── Web Stack (SHARED INFRA — DO NOT CHANGE) ───
web:
backend: firecrawl
search_backend: searxng
extract_backend: firecrawl
firecrawl:
base_url: http://192.168.68.7:3002/
# ─── MCP Servers (SHARED INFRA) ───
mcp_servers:
context7:
connect_timeout: 60
timeout: 300
url: http://localhost:8079/mcp
ra-h-os:
url: http://192.168.68.65:3100/mcp
timeout: 120
connect_timeout: 60
# ─── Compression ───
compression:
enabled: true
model: strix-moe # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
provider: harness
max_context_window: 131072 # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K
target_ratio: 0.30
protect_last_n: 40
hygiene_hard_message_limit: 350
protect_first_n: 3
abort_on_summary_failure: false
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# All auxiliary services MUST use identical model, base_url, and api_key_env:
# model: gpu-light # stable alias (NOT raw "gemma-4-12b")
# base_url: http://192.168.68.116/v1
# api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
# in agent configs — use the stable aliases so model swaps don't break agents.
auxiliary:
vision:
provider: harness
model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b)
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY
timeout: 60
download_timeout: 30
web_extract:
provider: harness
model: gpu-light # stable alias for RTX 5070
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY
timeout: 30
compression:
provider: harness
model: strix-moe # MUST match compression.model above. Stable alias for Strix Halo.
base_url: http://192.168.68.116/v1 # Rule 5: /v1 NOT /litellm/v1
api_key_env: LITELLM_API_KEY
timeout: 300 # gpu-fleet: 300s for large-history summarization (was 60)
# ─── Delegation / Heavy Aux (use gpu-dense = RTX 3090) ───
# delegation.model and x_search.model use gpu-dense (NOT raw qwen3.6-27B-code).
delegation:
model: gpu-dense # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
provider: harness
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY
# ─── Custom Provider ───
custom_providers:
- name: harness
model: syslog-auto # weighted pool (default)
base_url: http://192.168.68.116/v1
api_key_env: LITELLM_API_KEY
api_mode: chat_completions
Key Update Procedure
When LiteLLM keys are regenerated (e.g., after infrastructure changes):
- If SSH available: Update Infisical vault:
infisical secrets set LITELLM_API_KEY=sk-<NEW> --project=agents --env=production, thenssh <host> "systemctl restart hermes-gateway" - If SSH unavailable: Send Zulip DM via abiba-bot with update command
- After update: Restart Hermes on the agent host
- Verify:
curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models
Configuration Rules
Rule 1: Shared Infra Is Locked
The following MUST be identical across ALL profiles:
web.backend,web.search_backend,web.extract_backendweb.firecrawl.base_urlmcp_servers.ra-h-os.urlcustom_providers[0].base_url
Rule 2: Model Choice Is Free
model.default— per agentfallback_providers.model— per agentcustom_providers[0].model— per agent
Rule 3: API Keys via Environment
- Prefer
api_key_env: LITELLM_API_KEYover hardcoded keys - Hardcoded keys in config.yaml become stale after key rotation
- Infisical vault secrets persist across config updates / reinstalls
- NEW (July 2026): Always keep a local
.envfallback. Infisical service tokens can expire/404 (tanko incident: token not found, gateway ran without key for hours). The.envfile should have the key uncommented as a fallback:LITELLM_API_KEY=sk-... # [INFISICAL] Also sourced from vault.sysloggh.net - Restart Hermes after env var updates
Rule 4: Sub-Agent Profiles Inherit Auth
- Sub-agent profiles (
/root/.hermes/profiles/*/config.yaml) must have:api_key: ''— inherit from main config's custom_providerbase_url: ''— inherit from main config- Auxiliary tasks:
api_key: '',provider: harness
- Never hardcode a key in sub-agent profiles
- When main config uses
api_key_env, sub-agents automatically use it - This means key rotation only touches ONE vault secret (
LITELLM_API_KEY)
Rule 5: Main Config Base URL
|- Use direct IP: http://192.168.68.116/v1
|- NOT the NetBird URL (litellm.sysloggh.net) — can cause 502 when NetBird is down
|- NOT the old path (/litellm/v1) — nginx now routes /v1 directly
Rule 6: max_tokens Is Required (Thermal Safety)
- Every Hermes config MUST set
model.max_tokens: 4096— this is non-negotiable - Prevents unbounded generation that caused the July 2 Strix Halo GPU thermal incident
- The value flows through:
config → agent.max_tokens → transport build_kwargs → API max_tokens - Even though the server now has
-n 8192hard cap (set by Abiba), the client cap is the first line of defense - Apply to BOTH main config AND all sub-agent profiles
- For agents needing longer outputs: raise to 8192, but never omit
Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)
- Vision and web_extract use
gemma-4-12b(RTX 5070 — 12GB, vision-optimized) - Compression uses
strix-moe(stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized) strix-moeis the only valid compression model name — LiteLLM does NOT serveornith-1.0-35b(it servesstrix-moe,qwen3.6-35B-udq4,gpu-dense,gpu-light,syslog-auto,gemma-4-12b,qwen3.6-27B-code). Old configs withornith-1.0-35bcause 403/model-not-found on compression calls.- All auxiliary services MUST use identical routing:
base_url: http://192.168.68.116/v1(Rule 5:/v1, NOT/litellm/v1)api_key_env: LITELLM_API_KEY
- Do NOT use
syslog-autofor auxiliary tasks — it routes unpredictably - Compression on Strix Halo: The strix-moe alias routes to Strix Halo (64GB UMA, 128K context) — the designated compression GPU. This frees the RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
- The
compression:block'smodelMUST matchauxiliary: compression: model - The
compression: max_context_window: 131072MUST match actual GPU capacity (128K)
Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)
- RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code): Heavy reasoning, code gen, long conversations
- RTX 5070 (12GB, 128K ctx, gemma-4-12b): Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
- Strix Halo (64GB, 128K ctx, strix-moe): Context compression, summarization, long docs
- Agent profiles MUST route auxiliary tasks to the correct GPU:
auxiliary.vision.model: gemma-4-12b(RTX 5070)auxiliary.web_extract.model: gemma-4-12b(RTX 5070)auxiliary.compression.model: strix-moe(Strix Halo)
- Default model (
model.default) and custom_provider remainsyslog-autofor auto-routing - For 128K context window:
threshold: 0.65(fires at ~85K tokens) - Do NOT use
threshold: 0.25— this fires at 65K, causing premature context loss - Do NOT use
threshold: 0.80— this delays until ~105K, leaving only 23K margin max_context_window: 131072MUST match the model's actual capacity (128K)- See
devops-hermes-compressionskill for full reference
Rule 9: Compression Threshold for 128K Models
- For 128K context window:
threshold: 0.65(fires at ~85K tokens) - Do NOT use
threshold: 0.25— this fires at 65K, causing premature context loss - Do NOT use
threshold: 0.80— this delays until ~105K, leaving only 23K margin max_context_window: 131072MUST match the model's actual capacity (all GPUs = 128K)- See
devops-hermes-compressionskill for full reference
Rule 10: Default Model Must Be syslog-auto (All Agents)
- Hermes agents:
model.default: syslog-auto,custom_providers[0].model: syslog-auto - pi agents:
defaultModel: syslog-autoinsettings.json, first model inmodels.json syslog-autois the LiteLLM routing model — it load-balances between strix-moe and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against:- Model name typos that cause 403 errors and silent worker failures
- Single GPU downtime (routing falls back automatically)
- Key/model authorization mismatches
- Exception: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models for specialized tasks, but MUST validate those models exist in the key's authorized list
Rule 11: Validate Model IDs Before Deployment (pi Agents)
- After configuring a pi agent's
models.json, verify every model ID:curl -s http://192.168.68.116:4000/v1/models \ -H "Authorization: Bearer <AGENT_KEY>" | jq '.data[].id' - All model IDs in
models.jsonMUST appear in the LiteLLM response - The agent's API key may have a SUBSET of the full model catalog — check per-key
- A non-existent model ID causes 403 errors that silently break the pi RPC worker
(no
agent_endemitted, worker stays "busy", Zulip messages pile up unprocessed)
Rule 12: Context-Issue Diagnostic Checklist (ADDED 2026-07-16, WAL #1300)
When an agent shows "context issues" (premature compression, 401s, 504s, DeepSeek fallback), verify ALL FOUR of these against the live config. They are the only root causes found in production:
- max_context_window correct? — BOTH
compression.max_context_windowANDcontext.max_context_windowMUST be131072(all GPUs are 128K). A value of262144causes instability near 100K and must NOT be used. ~83K instead of ~170K. Check:grep -n max_context_window ~/.hermes/config.yaml - base_url uses /v1 NOT /litellm/v1? —
custom_providers[0].base_url,delegation.base_url, and ALLauxiliary.*.base_urlMUST behttp://192.168.68.116/v1(Rule 5). nginx/litellm/has a 60s default timeout → 504 on any inference >60s;/v1/has 600s. Check:grep -n 'litellm/v1' ~/.hermes/config.yaml(must return NOTHING) - LITELLM_API_KEY valid? — The key must be a real LiteLLM key (
sk-+ 64 hex, 67 chars). Malformed values (e.g.sk-_SWAl_Vu_…, 47 chars) return 401 → DeepSeek fallback. Verify:curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models(must be 200) - custom_providers aligned? —
model: syslog-auto(Rule 10),api_mode: chat_completions(NOTresponses). A wrong api_mode causes silent request failures.
One-line agent health check (run on the agent host):
PID=$(pgrep -f "python -m hermes_cli.main gateway run" | head -1)
cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | sed 's/=.*/<set>/'
curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $(cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | cut -d= -f2)" http://192.168.68.116/v1/models
Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300)
Agents inject LITELLM_API_KEY via ONE of two mechanisms. Both are valid; the contract
requirement is that the key is a valid LiteLLM virtual key (HTTP 200 on /v1/models).
Pattern A — systemd drop-in (Koby, Koonimo, and any agent without infisical wrapper):
A systemd drop-in /etc/systemd/system/hermes-gateway.service.d/litellm-key.conf sets the key:
[Service]
Environment="LITELLM_API_KEY=sk-<VALID_KEY>"
The service unit hermes-gateway.service runs python -m hermes_cli.main gateway run --replace
directly (no infisical). Apply with systemctl daemon-reload && systemctl restart hermes-gateway.
- Koonimo (CT113/.114): service =
hermes-gateway.service, drop-in has the key. - Koby (CT111/.129): service =
hermes-gateway.service(created 2026-07-16), ExecStart uses--replaceto win the lock against strayhermes gateway restartinvocations. Key also in/etc/environment.
Pattern B — infisical-gateway.sh wrapper (Mumuni):
The wrapper sources ~/.hermes/.env then exports LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY".
See litellm-api-keys.prose.md § Machine Identity for Vault Writes for vault sync.
Verification (all agents):
GP=$(pgrep -f "python -m hermes_cli.main gateway run" | head -1)
K=$(cat /proc/$GP/environ | tr '\0' '\n' | grep '^LITELLM_API_KEY=' | cut -d= -f2)
curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.168.68.116/v1/models # must be 200
/etc/environmentis NO LONGER the canonical key source (stale values there caused 401s).- Do NOT leave a hardcoded stale key in
/etc/environment— it shadows the drop-in/wrapper.
Execution
- Check current config — Read the target agent's config.yaml
- Compare against template — Identify missing or divergent sections
- Apply shared infra — Lock web/MCP/compression sections to template values
- Apply agent key — Set from agent_keys table above
- Set model choice — Per agent's workload
- Verify — curl all shared endpoints, test the model with the new key
- Report — What was changed, preserved, custom