Files
prose-contracts/hermes-config-template.prose.md
root cd479caeec Fix: update compression model to syslog-auto across contract and audit (Rule 7)
- Updated hermes-config-template.prose.md: all references to strix-moe for
  compression changed to syslog-auto to match operational decision on 2026-07-23
  (prevents sustained Strix Halo thermal load via weighted pool).
- Updated audit-hermes-config.py Rule 7 to expect syslog-auto instead of
  strix-moe, ensuring Abiba's next run validates against the correct baseline.
2026-07-23 18:02:56 +00:00

21 KiB

kind, name, description
kind name description
template hermes-config-template Standard Hermes configuration template for Syslog Solution LLC agents. Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, RA-H OS MCP) while keeping agent-specific API keys and model choices. UPDATED 2026-07-16: Compression model is the stable alias `strix-moe` (NOT `ornith-1.0-35b`, which LiteLLM does not serve). All 3 GPUs verified at 128K (reduced from 256K 2026-07-17 for stability). Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the 2026-07-16 Mumuni root-cause investigation (WAL #1300). UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).

Maintains

  • template_version: "2.1.0"
  • last_applied: timestamp
  • agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
  • agent_keys: map (see Agent Keys section)
  • infra_endpoints_verified: array

Agent Keys (LiteLLM — Current 2026-07-11)

Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM PostgreSQL DB via POST /key/generate on CT 116. Keys are stored in the DB. The env var LITELLM_API_KEY is injected at runtime via infisical run -- wrapper (project=agents, env=production). /etc/environment and ~/.hermes/.env are NO LONGER used for agent keys — stripped and tagged # [INFISICAL] post-migration. Sub-agent profiles inherit auth from the main config — no separate keys needed.

Agent Key Alias Host SSH Sub-Agents
Tanko tanko-* 192.168.68.122 jerome@.122
Mumuni mumuni 192.168.68.123 root@.123 6 profiles ✱
Abiba abiba-pi 192.168.68.24 local
Koby koby CT 111 (tdunna) Zulip
Koonimo koonimo CT 114 (baggy) SSH root
Kagenz0 kagenz0-* ? Zulip

CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).

✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research, syslog-review, syslog-writer — all at /root/.hermes/profiles/<name>/config.yaml

Infrastructure Stack

Component Endpoint Purpose
Firecrawl http://192.168.68.7:3002/ Web content extraction
SearXNG http://storepve:8888 Privacy-respecting web search
LiteLLM http://192.168.68.116/v1 Unified model gateway (via nginx)
LiteLLM (NetBird) https://litellm.sysloggh.net/v1 Alternative (may have 502 issues)
RA-H OS MCP http://192.168.68.65:3100/mcp Knowledge graph bridge
Context7 MCP http://localhost:8079/mcp Documentation queries

API Key Rules

  • api_key_env: LITELLM_API_KEY — Use env var for main model auth (preferred) Key is injected at runtime via infisical run -- wrapper — never in /etc/environment
  • api_key: '' — Sub-agents leave empty to inherit from main config's custom_provider
  • api_key: sk-... — Hardcoded key only as fallback when env var not possible
  • Store LITELLM_API_KEY in Infisical vault (project=agents, env=production)
  • Sub-agents NEVER get their own key — they share the host agent's key
  • Restart Hermes gateway after updating vault secret (key auto-injected via wrapper)

Sub-Agent Profiles (Mumuni pattern)

Mumuni has 6 sub-agent profiles in /root/.hermes/profiles/<name>/config.yaml:

profiles/
├── syslog-code/config.yaml      # Code generation
├── syslog-devops/config.yaml     # DevOps/infrastructure
├── syslog-email/config.yaml      # Email processing
├── syslog-research/config.yaml   # Research & analysis
├── syslog-review/config.yaml     # Code review
└── syslog-writer/config.yaml     # Content writing

Sub-agent profile rules:

  1. api_key must be emptyapi_key: '' or omitted entirely
  2. base_url must be empty — inherits from main config's custom_provider
  3. provider is auto or harness — routes through the shared LiteLLM gateway
  4. model is agent-specific — each sub-agent can have its own default model
  5. Auxiliary tasks (vision, compression, etc.) also leave api_key empty
  6. Never hardcode a key in sub-agent profiles

This ensures all 6 sub-agents use the same LiteLLM key injected via infisical run -- wrapper. When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs) work immediately after restart.

Template — Required Sections

# ─── Model Selection ───
model:
  default: <agent_model>          # e.g., strix-moe, qwen3.6-27B-code, syslog-auto
  provider: harness
  base_url: http://192.168.68.116/v1
  api_key_env: LITELLM_API_KEY    # Injected via infisical run -- wrapper
  max_tokens: 4096                # ⚠️ CRITICAL: Prevents unbounded generation
  context_length: 131072          # For syslog-auto (all GPUs at 128K for stability).
                                  # Set 65536 if using gemma-4-12b directly (tight VRAM).

fallback_providers:
  provider: deepseek
  model: deepseek-chat
  api_key_env: DEEPSEEK_API_KEY

# ─── Web Stack (SHARED INFRA — DO NOT CHANGE) ───
web:
  backend: firecrawl
  search_backend: searxng
  extract_backend: firecrawl
  firecrawl:
    base_url: http://192.168.68.7:3002/

# ─── MCP Servers (SHARED INFRA) ───
mcp_servers:
  context7:
    connect_timeout: 60
    timeout: 300
    url: http://localhost:8079/mcp
  ra-h-os:
    url: http://192.168.68.65:3100/mcp
    timeout: 120
    connect_timeout: 60

# ─── Compression ───
compression:
  enabled: true
  model: syslog-auto               # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
  provider: harness
  max_context_window: 131072      # MUST match actual GPU capacity. All 3 GPUs are 128K (Jul 17).
  threshold: 0.65                 # Fires at ~170K for 262K window, ~85K for 128K
  target_ratio: 0.30
  protect_last_n: 40
  hygiene_hard_message_limit: 350
  protect_first_n: 3
  abort_on_summary_failure: false

# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# All auxiliary services MUST use identical model, base_url, and api_key_env:
#   model: gpu-light                 # stable alias (NOT raw "gemma-4-12b")
#   base_url: http://192.168.68.116/v1
#   api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4)
# in agent configs — use the stable aliases so model swaps don't break agents.
auxiliary:
  vision:
    provider: harness
    model: gpu-light            # stable alias for RTX 5070 (was raw gemma-4-12b)
    base_url: http://192.168.68.116/v1
    api_key_env: LITELLM_API_KEY
    timeout: 60
    download_timeout: 30
  web_extract:
    provider: harness
    model: gpu-light            # stable alias for RTX 5070
    base_url: http://192.168.68.116/v1
    api_key_env: LITELLM_API_KEY
    timeout: 30
  compression:
    provider: harness
    model: syslog-auto              # MUST match compression.model above. Stable alias for Strix Halo (weighted pool).
    base_url: http://192.168.68.116/v1   # Rule 5: /v1 NOT /litellm/v1
    api_key_env: LITELLM_API_KEY
    timeout: 300                 # gpu-fleet: 300s for large-history summarization (was 60)

# ─── Delegation / Heavy Aux (use gpu-dense = RTX 3090) ───
# delegation.model and x_search.model use gpu-dense (NOT raw qwen3.6-27B-code).

delegation:
  model: gpu-dense              # stable alias for RTX 3090 (was raw qwen3.6-27B-code)
  provider: harness
  base_url: http://192.168.68.116/v1
  api_key_env: LITELLM_API_KEY

# ─── Custom Provider ───
custom_providers:
  - name: harness
    model: syslog-auto            # weighted pool (default)
    base_url: http://192.168.68.116/v1
    api_key_env: LITELLM_API_KEY
    api_mode: chat_completions

Key Update Procedure

When LiteLLM keys are regenerated (e.g., after infrastructure changes):

  1. If SSH available: Update Infisical vault: infisical secrets set LITELLM_API_KEY=sk-<NEW> --project=agents --env=production, then ssh <host> "systemctl restart hermes-gateway"
  2. If SSH unavailable: Send Zulip DM via abiba-bot with update command
  3. After update: Restart Hermes on the agent host
  4. Verify: curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models

Configuration Rules

Rule 1: Shared Infra Is Locked

The following MUST be identical across ALL profiles:

  • web.backend, web.search_backend, web.extract_backend
  • web.firecrawl.base_url
  • mcp_servers.ra-h-os.url
  • custom_providers[0].base_url

Rule 2: Model Choice Is Free

  • model.default — per agent
  • fallback_providers.model — per agent
  • custom_providers[0].model — per agent

Rule 3: API Keys via Environment

  • Prefer api_key_env: LITELLM_API_KEY over hardcoded keys
  • Hardcoded keys in config.yaml become stale after key rotation
  • Infisical vault secrets persist across config updates / reinstalls
  • NEW (July 2026): Always keep a local .env fallback. Infisical service tokens can expire/404 (tanko incident: token not found, gateway ran without key for hours). The .env file should have the key uncommented as a fallback:
    LITELLM_API_KEY=sk-...
    # [INFISICAL] Also sourced from vault.sysloggh.net
    
  • Restart Hermes after env var updates

Rule 4: Sub-Agent Profiles Inherit Auth

  • Sub-agent profiles (/root/.hermes/profiles/*/config.yaml) must have:
    • api_key: '' — inherit from main config's custom_provider
    • base_url: '' — inherit from main config
    • Auxiliary tasks: api_key: '', provider: harness
  • Never hardcode a key in sub-agent profiles
  • When main config uses api_key_env, sub-agents automatically use it
  • This means key rotation only touches ONE vault secret (LITELLM_API_KEY)

Rule 5: Main Config Base URL

|- Use direct IP: http://192.168.68.116/v1 |- NOT the NetBird URL (litellm.sysloggh.net) — can cause 502 when NetBird is down |- NOT the old path (/litellm/v1) — nginx now routes /v1 directly

Rule 6: max_tokens Is Required (Thermal Safety)

  • Every Hermes config MUST set model.max_tokens: 4096 — this is non-negotiable
  • Prevents unbounded generation that caused the July 2 Strix Halo GPU thermal incident
  • The value flows through: config → agent.max_tokens → transport build_kwargs → API max_tokens
  • Even though the server now has -n 8192 hard cap (set by Abiba), the client cap is the first line of defense
  • Apply to BOTH main config AND all sub-agent profiles
  • For agents needing longer outputs: raise to 8192, but never omit

Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)

  • Vision and web_extract use gemma-4-12b (RTX 5070 — 12GB, vision-optimized)
  • Compression uses strix-moe (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized)
  • strix-moe is the only valid compression model name — LiteLLM does NOT serve ornith-1.0-35b (it serves strix-moe, qwen3.6-35B-udq4, gpu-dense, gpu-light, syslog-auto, gemma-4-12b, qwen3.6-27B-code). Old configs with ornith-1.0-35b cause 403/model-not-found on compression calls.
  • OPERATIONAL DECISION (2026-07-23): Use syslog-auto for compression across all agents. The syslog-auto alias routes to the Strix Halo, but uses the weighted pool instead of pinning to strix-moe directly. This prevents sustained Strix Halo thermal load because the pool can fall back to other GPUs if Strix gets hot. Both compression.model and auxiliary.compression.model MUST be syslog-auto.
  • All auxiliary services MUST use identical routing:
    • base_url: http://192.168.68.116/v1 (Rule 5: /v1, NOT /litellm/v1)
    • api_key_env: LITELLM_API_KEY
  • Do NOT use syslog-auto for auxiliary tasks — it routes unpredictably
  • Compression on Strix Halo: The strix-moe alias routes to Strix Halo (64GB UMA, 128K context) — the designated compression GPU. This frees the RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
  • The compression: block's model MUST match auxiliary: compression: model
  • The compression: max_context_window: 131072 MUST match actual GPU capacity (128K)

Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)

  • RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code): Heavy reasoning, code gen, long conversations
  • RTX 5070 (12GB, 128K ctx, gemma-4-12b): Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
  • Strix Halo (64GB, 128K ctx, syslog-auto): Context compression, summarization, long docs
  • Agent profiles MUST route auxiliary tasks to the correct GPU:
    • auxiliary.vision.model: gemma-4-12b (RTX 5070)
    • auxiliary.web_extract.model: gemma-4-12b (RTX 5070)
    • auxiliary.compression.model: syslog-auto (Strix Halo)
  • Default model (model.default) and custom_provider remain syslog-auto for auto-routing
  • For 128K context window: threshold: 0.65 (fires at ~85K tokens)
  • Do NOT use threshold: 0.25 — this fires at 65K, causing premature context loss
  • Do NOT use threshold: 0.80 — this delays until ~105K, leaving only 23K margin
  • max_context_window: 131072 MUST match the model's actual capacity (128K)
  • See devops-hermes-compression skill for full reference

Rule 9: Compression Threshold for 128K Models

  • For 128K context window: threshold: 0.65 (fires at ~85K tokens)
  • Do NOT use threshold: 0.25 — this fires at 65K, causing premature context loss
  • Do NOT use threshold: 0.80 — this delays until ~105K, leaving only 23K margin
  • max_context_window: 131072 MUST match the model's actual capacity (all GPUs = 128K)
  • See devops-hermes-compression skill for full reference

Rule 10: Default Model Must Be syslog-auto (All Agents)

  • Hermes agents: model.default: syslog-auto, custom_providers[0].model: syslog-auto
  • pi agents: defaultModel: syslog-auto in settings.json, first model in models.json
  • syslog-auto is the LiteLLM routing model — it load-balances between strix-moe and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against:
    • Model name typos that cause 403 errors and silent worker failures
    • Single GPU downtime (routing falls back automatically)
    • Key/model authorization mismatches
  • Exception: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models for specialized tasks, but MUST validate those models exist in the key's authorized list

Rule 11: Validate Model IDs Before Deployment (pi Agents)

  • After configuring a pi agent's models.json, verify every model ID:
    curl -s http://192.168.68.116:4000/v1/models \
      -H "Authorization: Bearer <AGENT_KEY>" | jq '.data[].id'
    
  • All model IDs in models.json MUST appear in the LiteLLM response
  • The agent's API key may have a SUBSET of the full model catalog — check per-key
  • A non-existent model ID causes 403 errors that silently break the pi RPC worker (no agent_end emitted, worker stays "busy", Zulip messages pile up unprocessed)

Rule 12: Context-Issue Diagnostic Checklist (ADDED 2026-07-16, WAL #1300)

When an agent shows "context issues" (premature compression, 401s, 504s, DeepSeek fallback), verify ALL FOUR of these against the live config. They are the only root causes found in production:

  1. max_context_window correct? — BOTH compression.max_context_window AND context.max_context_window MUST be 131072 (all GPUs are 128K). A value of 262144 causes instability near 100K and must NOT be used. ~83K instead of ~170K. Check: grep -n max_context_window ~/.hermes/config.yaml
  2. base_url uses /v1 NOT /litellm/v1?custom_providers[0].base_url, delegation.base_url, and ALL auxiliary.*.base_url MUST be http://192.168.68.116/v1 (Rule 5). nginx /litellm/ has a 60s default timeout → 504 on any inference >60s; /v1/ has 600s. Check: grep -n 'litellm/v1' ~/.hermes/config.yaml (must return NOTHING)
  3. LITELLM_API_KEY valid? — The key must be a real LiteLLM key (sk- + 64 hex, 67 chars). Malformed values (e.g. sk-_SWAl_Vu_…, 47 chars) return 401 → DeepSeek fallback. Verify: curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models (must be 200)
  4. custom_providers aligned?model: syslog-auto (Rule 10), api_mode: chat_completions (NOT responses). A wrong api_mode causes silent request failures.

One-line agent health check (run on the agent host):

# Use grep -v infisical to avoid matching the bash wrapper that contains the same string
PID=$(pgrep -f "python -m hermes_cli.main gateway run" | grep -v infisical | head -1)
cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | sed 's/=.*/<set>/'
curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $(cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | cut -d= -f2)" http://192.168.68.116/v1/models

Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300)

Agents inject LITELLM_API_KEY via ONE of two mechanisms. Both are valid; the contract requirement is that the key is a valid LiteLLM virtual key (HTTP 200 on /v1/models).

Pattern A — systemd drop-in (Koby, Koonimo, and any agent without infisical wrapper): A systemd drop-in /etc/systemd/system/hermes-gateway.service.d/litellm-key.conf sets the key:

[Service]
Environment="LITELLM_API_KEY=sk-<VALID_KEY>"

The service unit hermes-gateway.service runs python -m hermes_cli.main gateway run --replace directly (no infisical). Apply with systemctl daemon-reload && systemctl restart hermes-gateway.

  • Koonimo (CT113/.114): service = hermes-gateway.service, drop-in has the key.
  • Koby (CT111/.129): service = hermes-gateway.service (created 2026-07-16), ExecStart uses --replace to win the lock against stray hermes gateway restart invocations. Key also in /etc/environment.

Pattern B — infisical-gateway.sh wrapper (Mumuni): The wrapper sources ~/.hermes/.env then exports LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY". See litellm-api-keys.prose.md § Machine Identity for Vault Writes for vault sync.

⚠️ Vault empty-key guard: If the vault stores the secret as an empty string, the wrapper will inject an empty key and the gateway will silently get 401 errors on all LiteLLM requests (triggering silent DeepSeek fallback). The .env fallback is present but the vault takes precedence when the secret key exists (even if empty).

Fix: The wrapper MUST validate the key length after injection. If LITELLM_API_KEY is empty or shorter than 20 chars, log a warning and either fail with a clear error message or fall back to the .env value before starting the gateway.

Verification (all agents):

# Use grep -v infisical to avoid matching the bash wrapper that contains the same string
GP=$(pgrep -f "python -m hermes_cli.main gateway run" | grep -v infisical | head -1)
K=$(cat /proc/$GP/environ | tr '\0' '\n' | grep '^LITELLM_API_KEY=' | cut -d= -f2)
curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.168.68.116/v1/models  # must be 200
  • /etc/environment is NO LONGER the canonical key source (stale values there caused 401s).
  • Do NOT leave a hardcoded stale key in /etc/environment — it shadows the drop-in/wrapper.

Rule 14: Provider Name Must Match custom_providers Name (ADDED 2026-07-19, WAL #1471)

  • model.provider MUST be harness (the custom_providers[0].name), NOT the literal string custom
  • When provider: custom, Hermes' _get_named_custom_provider("custom") returns None (no provider is named "custom" — it is named "harness"), causing a fall-through to the generic resolution path (source: env/config) at runtime_provider.py:1156
  • The generic path builds api_key_candidates from model.api_key (empty), host-gated OLLAMA/OPENAI/OPENROUTER keys, and _host_derived_api_key (returns "" for IP addresses)
  • The generic path does NOT resolve model.api_key_env or custom_providers.key_envLITELLM_API_KEY is never read, producing api_key = "no-key-required" → HTTP 401
  • The named custom provider path (source: custom_provider:harness) DOES read key_env — but only triggers when provider matches the custom_providers[0].name
  • All sections MUST use provider: harness: model, compression, auxiliary.vision, auxiliary.web_extract, auxiliary.compression, delegation
  • Only fallback_providers uses a different provider (deepseek) for true fallback diversity
  • Diagnostic: If you see source: env/config in a request dump or log, the provider name is wrong. It should be source: custom_provider:harness.
  • Audit script: Run python3 /root/prose-contracts/audit-hermes-config.py <config.yaml> before and after any config change to catch this and all other rule violations.

Execution

  1. Check current config — Read the target agent's config.yaml
  2. Compare against template — Identify missing or divergent sections
  3. Apply shared infra — Lock web/MCP/compression sections to template values
  4. Apply agent key — Set from agent_keys table above
  5. Set model choice — Per agent's workload
  6. Verify — curl all shared endpoints, test the model with the new key
  7. Report — What was changed, preserved, custom