Files
prose-contracts/hermes-config-template.prose.md
T

24 KiB

kind, name, description
kind name description
template hermes-config-template Standard Hermes configuration template for Syslog Solution LLC agents. Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, RA-H OS MCP) while keeping agent-specific API keys and model choices. UPDATED 2026-08-07: Added Rule 15 (MCP Validation) from the 2026-08-07 keyless-MCP incident. Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the 2026-07-16 Mumuni root-cause investigation (WAL #1300). UPDATED 2026-07-12: GPU workload redistributed. Compression → Strix Halo. RTX 3090 context verified at 128K. Infisical .env fallback required (Rule 3/13).

Maintains

  • template_version: "2.1.0"
  • last_applied: timestamp
  • agents_configured: ["mumuni", "abiba", "koby", "koonimo", "kagenz0"] # tanko removed 2026-08-27 (now on DSH/DeepSeek Harness)
  • agent_keys: map (see Agent Keys section)
  • infra_endpoints_verified: array

Agent Keys (LiteLLM — Current 2026-07-11)

Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM PostgreSQL DB via POST /key/generate on CT 116. Keys are stored in the DB. The env var LITELLM_API_KEY is injected at runtime via infisical run -- wrapper (project=agents, env=production). /etc/environment and ~/.hermes/.env are NO LONGER used for agent keys — stripped and tagged # [INFISICAL] post-migration. Sub-agent profiles inherit auth from the main config — no separate keys needed.

Agent Key Alias Host SSH Sub-Agents
Abiba abiba-pi 192.168.68.24 local —
Koby koby CT 111 (tdunna) Zulip —
Koonimo koonimo CT 113 (baggy) SSH root —
Kagenz0 kagenz0-* ? Zulip —

CT hostnames (tdunna, baggy) differ from agent identities (koby, koonimo).

✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research, syslog-review, syslog-writer — all at /root/.hermes/profiles/<name>/config.yaml

Infrastructure Stack

Component Endpoint Purpose
Firecrawl http://192.168.68.7:3002/ Web content extraction
SearXNG http://192.168.68.7:8888 Privacy-respecting web search
LiteLLM http://192.168.68.116/v1 Unified model gateway (via nginx)
LiteLLM (NetBird) https://litellm.sysloggh.net/v1 Alternative (may have 502 issues)
RA-H OS MCP http://192.168.68.65:3100/mcp Knowledge graph bridge
Context7 MCP http://localhost:8079/mcp Documentation queries

API Key Rules

  • api_key_env: LITELLM_API_KEY — Use env var for main model auth (preferred) Key is injected at runtime via infisical run -- wrapper — never in /etc/environment
  • api_key: '' — Sub-agents leave empty to inherit from main config's custom_provider
  • api_key: sk-... — Hardcoded key only as fallback when env var not possible
  • Store LITELLM_API_KEY in Infisical vault (project=agents, env=production)
  • Sub-agents NEVER get their own key — they share the host agent's key
  • Restart Hermes gateway after updating vault secret (key auto-injected via wrapper)

Sub-Agent Profiles (Mumuni pattern)

Mumuni has 6 sub-agent profiles in /root/.hermes/profiles/<name>/config.yaml:

profiles/
├── syslog-code/config.yaml      # Code generation
├── syslog-devops/config.yaml     # DevOps/infrastructure
├── syslog-email/config.yaml      # Email processing
├── syslog-research/config.yaml   # Research & analysis
├── syslog-review/config.yaml     # Code review
└── syslog-writer/config.yaml     # Content writing

Sub-agent profile rules:

  1. api_key must be empty — api_key: '' or omitted entirely
  2. base_url must be empty — inherits from main config's custom_provider
  3. provider is auto or harness — routes through the shared LiteLLM gateway
  4. model is agent-specific — each sub-agent can have its own default model
  5. Auxiliary tasks (vision, compression, etc.) also leave api_key empty
  6. Never hardcode a key in sub-agent profiles

This ensures all 6 sub-agents use the same LiteLLM key injected via infisical run -- wrapper. When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs) work immediately after restart.

Template — Required Sections

# ─── Model Selection ───
model:
  default: <agent_model>          # e.g., strix-moe, gpu-dense, syslog-auto
  provider: harness
  base_url: http://192.168.68.116/litellm/v1   # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
  api_key_env: LITELLM_API_KEY    # Injected via infisical run -- wrapper
  max_tokens: 4096                # ⚠️ CRITICAL: Prevents unbounded generation
  context_length: 131072          # Conservative floor for syslog-auto (NVIDIA hosts 128K; Strix Halo 256K).
                                  # ⚠️ MANDATORY: Hermes probes unknown models from 256K
                                  # and falls back to 256K when /v1/models lacks a context
                                  # field (llama-server does). Without this override, agents
                                  # silently run syslog-auto at 256K (verified 2026-08-09).
                                  # Set 65536 if pinning a single model directly (tight VRAM).

fallback_providers:
  provider: deepseek
  model: deepseek-chat
  api_key_env: DEEPSEEK_API_KEY

# ─── Web Stack (SHARED INFRA — DO NOT CHANGE) ───
web:
  backend: firecrawl
  search_backend: searxng
  extract_backend: firecrawl
  firecrawl:
    base_url: http://192.168.68.7:3002/

# ─── MCP Servers (SHARED INFRA) ───
mcp_servers:
  context7:
    connect_timeout: 60
    timeout: 300
    url: http://localhost:8079/mcp
  ra-h-os:
    url: http://192.168.68.65:3100/mcp
    timeout: 120
    connect_timeout: 60

# ─── Compression ───
compression:
  enabled: true
  model: syslog-auto               # ⚠️ Must match auxiliary.compression.model. Stable alias (gpu-fleet § Stable Role-Based Aliases). NOT ornith-1.0-35b (LiteLLM does not serve that name).
  provider: harness
  max_context_window: 131072      # MUST stay at the syslog-auto pool floor: NVIDIA hosts are 128K, Strix Halo 256K (2026-09-12).
  threshold: 0.65                 # Fires at ~170K for 262K window, ~85K for 128K
  target_ratio: 0.30
  protect_last_n: 40
  hygiene_hard_message_limit: 350
  protect_first_n: 3
  abort_on_summary_failure: false

# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# All auxiliary services MUST use identical model, base_url, and api_key_env:
#   model: gpu-vision                # stable alias (NOT a raw model name)
#   base_url: http://192.168.68.116/litellm/v1   # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
#   api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
auxiliary:
  vision:
    provider: harness
    model: gpu-vision           # stable alias for RTX 5070
    base_url: http://192.168.68.116/litellm/v1   # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
    api_key_env: LITELLM_API_KEY
    timeout: 60
    download_timeout: 30
  web_extract:
    provider: harness
    model: gpu-vision           # stable alias for RTX 5070
    base_url: http://192.168.68.116/litellm/v1   # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
    api_key_env: LITELLM_API_KEY
    timeout: 30
  compression:
    provider: harness
    model: syslog-auto              # MUST match compression.model above. Stable alias for Strix Halo (weighted pool).
    base_url: http://192.168.68.116/litellm/v1   # Rule 5 (2026-08-09): canonical authenticated path; /v1 also OK
    api_key_env: LITELLM_API_KEY
    timeout: 300                 # gpu-fleet: 300s for large-history summarization (was 60)

# ─── Delegation / Heavy Aux (use gpu-dense = RTX 3090) ───
# delegation.model and x_search.model use gpu-dense (NOT retired raw name).

delegation:
  model: gpu-dense              # stable alias for RTX 3090
  provider: harness
  base_url: http://192.168.68.116/litellm/v1   # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
  api_key_env: LITELLM_API_KEY

# ─── Custom Provider ───
custom_providers:
  - name: harness
    model: syslog-auto            # weighted pool (default)
    base_url: http://192.168.68.116/litellm/v1   # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
    api_key_env: LITELLM_API_KEY
    api_mode: chat_completions

Key Update Procedure

When LiteLLM keys are regenerated (e.g., after infrastructure changes):

  1. If SSH available: Update Infisical vault: infisical secrets set LITELLM_API_KEY=sk-<NEW> --project=agents --env=production, then ssh <host> "systemctl restart hermes-gateway"
  2. If SSH unavailable: Send Zulip DM via abiba-bot with update command
  3. After update: Restart Hermes on the agent host
  4. Verify: curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models

Configuration Rules

Rule 1: Shared Infra Is Locked

The following MUST be identical across ALL profiles:

  • web.backend, web.search_backend, web.extract_backend
  • web.firecrawl.base_url
  • mcp_servers.ra-h-os.url
  • custom_providers[0].base_url

Rule 2: Model Choice Is Free

  • model.default — per agent
  • fallback_providers.model — per agent
  • custom_providers[0].model — per agent

Rule 3: API Keys via Environment

  • Prefer api_key_env: LITELLM_API_KEY over hardcoded keys
  • Hardcoded keys in config.yaml become stale after key rotation
  • Infisical vault secrets persist across config updates / reinstalls
  • NEW (July 2026): Always keep a local .env fallback. Infisical service tokens can expire/404 (tanko incident: token not found, gateway ran without key for hours). The .env file should have the key uncommented as a fallback:
    LITELLM_API_KEY=sk-...
    # [INFISICAL] Also sourced from vault.sysloggh.net
    
  • Restart Hermes after env var updates

Rule 4: Sub-Agent Profiles Inherit Auth

  • Sub-agent profiles (/root/.hermes/profiles/*/config.yaml) must have:
    • api_key: '' — inherit from main config's custom_provider
    • base_url: '' — inherit from main config
    • Auxiliary tasks: api_key: '', provider: harness
  • Never hardcode a key in sub-agent profiles
  • When main config uses api_key_env, sub-agents automatically use it
  • This means key rotation only touches ONE vault secret (LITELLM_API_KEY)

Rule 5: Main Config Base URL (UPDATED 2026-08-09)

|- Use the authenticated LiteLLM path: http://192.168.68.116/litellm/v1 (canonical, captain-approved migration) |- Legacy http://192.68.68.116/v1 also works — nginx fronts BOTH paths with key auth (verified 2026-08-09: 401 without key, 200 with key, on both /v1 and /litellm/v1) |- Both locations have proxy_read_timeout 600s (verified in harness-nginx nginx.conf) — the old "60s timeout on /litellm/" claim was stale and is retracted |- NOT the NetBird URL (litellm.sysloggh.net) — can cause 502 when NetBird is down

Rule 6: max_tokens Is Required (Thermal Safety)

  • Every Hermes config MUST set model.max_tokens: 4096 — this is non-negotiable
  • Prevents unbounded generation that caused the July 2 Strix Halo GPU thermal incident
  • The value flows through: config → agent.max_tokens → transport build_kwargs → API max_tokens
  • Even though the server now has -n 8192 hard cap (set by Abiba), the client cap is the first line of defense
  • Apply to BOTH main config AND all sub-agent profiles
  • For agents needing longer outputs: raise to 8192, but never omit

Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16)

  • Vision and web_extract use gpu-vision (RTX 5070 — 12GB, vision-optimized)
  • Compression uses syslog-auto — the Strix Halo weighted pool (64GB, 256K ctx, compression-optimized); do NOT pin compression.model to strix-moe (audit Rule 7 rejects it)
  • ornith-1.0-35b is NOT a valid compression model name — LiteLLM does not serve it (do not restate the served model list here — CT 116 /opt/inference-harness/litellm_config.yaml is the single source of truth for models, aliases, weights and fallbacks). Old configs with ornith-1.0-35b cause 403/model-not-found on compression calls.
  • OPERATIONAL DECISION (2026-07-23): Use syslog-auto for compression across all agents. The syslog-auto alias routes to the Strix Halo, but uses the weighted pool instead of pinning to strix-moe directly. This prevents sustained Strix Halo thermal load because the pool can fall back to other GPUs if Strix gets hot. Both compression.model and auxiliary.compression.model MUST be syslog-auto.
  • All auxiliary services MUST use identical routing:
    • base_url: http://192.168.68.116/litellm/v1 (Rule 5, 2026-08-09: canonical authenticated; /v1 also OK)
    • api_key_env: LITELLM_API_KEY
  • Do NOT use syslog-auto for vision/web_extract — it routes unpredictably; compression is the deliberate exception (see the OPERATIONAL DECISION above)
  • Compression on Strix Halo: The strix-moe alias routes to Strix Halo (64GB UMA, 256K context) — the designated compression GPU. This frees the RTX 5070 for vision and web search, and the RTX 3090 for heavy reasoning.
  • The compression: block's model MUST match auxiliary: compression: model
  • The compression: max_context_window: 131072 MUST stay at the syslog-auto pool floor (NVIDIA hosts 128K; Strix Halo 256K)

Rule 8: GPU Workload Distribution (UPDATED 2026-07-16)

  • RTX 3090 (24GB, 128K ctx, gpu-dense): Heavy reasoning, code gen, long conversations
  • RTX 5070 (12GB, 128K ctx, gpu-vision): Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K)
  • Strix Halo (64GB, 256K ctx, syslog-auto): Context compression, summarization, long docs
  • Agent profiles MUST route auxiliary tasks to the correct GPU:
    • auxiliary.vision.model: gpu-vision (RTX 5070)
    • auxiliary.web_extract.model: gpu-vision (RTX 5070)
    • auxiliary.compression.model: syslog-auto (Strix Halo)
  • Default model (model.default) and custom_provider remain syslog-auto for auto-routing
  • For 128K context window: threshold: 0.65 (fires at ~85K tokens)
  • Do NOT use threshold: 0.25 — this fires at 65K, causing premature context loss
  • Do NOT use threshold: 0.80 — this delays until ~105K, leaving only 23K margin
  • max_context_window: 131072 MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
  • See devops-hermes-compression skill for full reference

Rule 9: Compression Threshold for 128K Models

  • For 128K context window: threshold: 0.65 (fires at ~85K tokens)
  • Do NOT use threshold: 0.25 — this fires at 65K, causing premature context loss
  • Do NOT use threshold: 0.80 — this delays until ~105K, leaving only 23K margin
  • max_context_window: 131072 MUST stay at the pool floor (NVIDIA hosts 128K; Strix Halo 256K)
  • See devops-hermes-compression skill for full reference

Rule 10: Default Model Must Be syslog-auto (All Agents)

  • Koby Exception: Per captain ruling 2026-08-11, Koby is a DeepSeek-primary external agent; its primary model remains deepseek-v4-flash (via api.deepseek.com to preserve DeepSeek-specific reasoning, while other sections follow Rule 10.
  • Hermes agents: model.default: syslog-auto, custom_providers[0].model: syslog-auto
  • pi agents: defaultModel: syslog-auto in settings.json, first model in models.json
  • syslog-auto is the LiteLLM routing model — it load-balances across the live pool (see CT 116 /opt/inference-harness/litellm_config.yaml for the current members and weights). Using it protects against:
    • Model name typos that cause 403 errors and silent worker failures
    • Single GPU downtime (routing falls back automatically)
    • Key/model authorization mismatches
  • Exception: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models for specialized tasks, but MUST validate those models exist in the key's authorized list

Rule 11: Validate Model IDs Before Deployment (pi Agents)

  • After configuring a pi agent's models.json, verify every model ID:
    curl -s http://192.168.68.116:4000/v1/models \
      -H "Authorization: Bearer <AGENT_KEY>" | jq '.data[].id'
    
  • All model IDs in models.json MUST appear in the LiteLLM response
  • The agent's API key may have a SUBSET of the full model catalog — check per-key
  • A non-existent model ID causes 403 errors that silently break the pi RPC worker (no agent_end emitted, worker stays "busy", Zulip messages pile up unprocessed)

Rule 12: Context-Issue Diagnostic Checklist (ADDED 2026-07-16, WAL #1300)

When an agent shows "context issues" (premature compression, 401s, 504s, DeepSeek fallback), verify ALL FOUR of these against the live config. They are the only root causes found in production:

  1. max_context_window correct? — BOTH compression.max_context_window AND context.max_context_window MUST be 131072 (the syslog-auto pool floor: NVIDIA hosts are 128K; Strix Halo is 256K). A 262144 client window can route to a 128K NVIDIA host and fail, so it must NOT be used. ~83K instead of ~170K. Check: grep -n max_context_window ~/.hermes/config.yaml
  2. base_url uses authenticated path? — custom_providers[0].base_url, delegation.base_url, and ALL auxiliary.*.base_url MUST be http://192.168.68.116/litellm/v1 (Rule 5, canonical) or http://192.168.68.116/v1 (legacy, still authenticated via nginx). BOTH verified 200 with key + 600s proxy_read_timeout on 2026-08-09. Never bare :4000 direct. Check: grep -nE 'base_url: http://192.168.68.116(:4000)?/v1' ~/.hermes/config.yaml — the ONLY paths allowed are /v1 or /litellm/v1 (both via nginx :80). :4000 or missing litellm/v1/v1 prefix = violation.
  3. LITELLM_API_KEY valid? — The key must be a real LiteLLM key (sk- + 64 hex, 67 chars). Malformed values (e.g. sk-_SWAl_Vu_…, 47 chars) return 401 → DeepSeek fallback. Verify: curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $LITELLM_API_KEY" http://192.168.68.116/v1/models (must be 200)
  4. custom_providers aligned? — model: syslog-auto (Rule 10), api_mode: chat_completions (NOT responses). A wrong api_mode causes silent request failures.

One-line agent health check (run on the agent host):

# Use grep -v infisical to avoid matching the bash wrapper that contains the same string
PID=$(pgrep -f "python -m hermes_cli.main gateway run" | grep -v infisical | head -1)
cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | sed 's/=.*/<set>/'
curl -s -o /dev/null -w 'key_health: %{http_code}\n' -H "Authorization: Bearer $(cat /proc/$PID/environ | tr '\0' '\n' | grep ^LITELLM_API_KEY= | cut -d= -f2)" http://192.168.68.116/v1/models

Rule 13: API Key Injection — Two Patterns (UPDATED 2026-07-16, WAL #1300)

Rule 14: Hermes Context Detection Uses max_model_tokens, NOT max_input_tokens

CRITICAL: Hermes context detection reads max_model_tokens (128K), NOT max_input_tokens (64K cap).

  • Abiba and Hermes agents: max_model_tokens: 131072 (128K) — unlimited context
  • Crewmates (ops, tune, verify, auth-keys, build): max_input_tokens: 64000 (64K) — capped
  • If you see max_input_tokens: 64000 in an Abiba/Hermes config, that's a mistake
  • Using max_input_tokens for Hermes agents causes premature context loss
  • Check: grep -n 'max_model_tokens\|max_input_tokens' ~/.hermes/config.yaml
  • Expected output: max_model_tokens: 131072 (not max_input_tokens)

Agents inject LITELLM_API_KEY via ONE of two mechanisms. Both are valid; the contract requirement is that the key is a valid LiteLLM virtual key (HTTP 200 on /v1/models).

Pattern A — systemd drop-in (Koby, Koonimo, and any agent without infisical wrapper): A systemd drop-in /etc/systemd/system/hermes-gateway.service.d/litellm-key.conf sets the key:

[Service]
Environment="LITELLM_API_KEY=sk-<VALID_KEY>"

The service unit hermes-gateway.service runs python -m hermes_cli.main gateway run --replace directly (no infisical). Apply with systemctl daemon-reload && systemctl restart hermes-gateway.

  • Koonimo (CT113/.114): service = hermes-gateway.service, drop-in has the key.
  • Koby (CT111/.129): service = hermes-gateway.service (created 2026-07-16), ExecStart uses --replace to win the lock against stray hermes gateway restart invocations. Key also in /etc/environment.

Pattern B — infisical-gateway.sh wrapper (Mumuni): The wrapper sources ~/.hermes/.env then exports LITELLM_API_KEY="$<AGENT>_LITELLM_API_KEY". See litellm-api-keys.prose.md § Machine Identity for Vault Writes for vault sync.

⚠️ Vault empty-key guard: If the vault stores the secret as an empty string, the wrapper will inject an empty key and the gateway will silently get 401 errors on all LiteLLM requests (triggering silent DeepSeek fallback). The .env fallback is present but the vault takes precedence when the secret key exists (even if empty).

Fix: The wrapper MUST validate the key length after injection. If LITELLM_API_KEY is empty or shorter than 20 chars, log a warning and either fail with a clear error message or fall back to the .env value before starting the gateway.

Verification (all agents):

# Use grep -v infisical to avoid matching the bash wrapper that contains the same string
GP=$(pgrep -f "python -m hermes_cli.main gateway run" | grep -v infisical | head -1)
K=$(cat /proc/$GP/environ | tr '\0' '\n' | grep '^LITELLM_API_KEY=' | cut -d= -f2)
curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $K" http://192.168.68.116/v1/models  # must be 200
  • /etc/environment is NO LONGER the canonical key source (stale values there caused 401s).
  • Do NOT leave a hardcoded stale key in /etc/environment — it shadows the drop-in/wrapper.

Rule 14: Provider Name Must Match custom_providers Name (ADDED 2026-07-19, WAL #1471)

  • model.provider MUST be harness (the custom_providers[0].name), NOT the literal string custom
  • When provider: custom, Hermes' _get_named_custom_provider("custom") returns None (no provider is named "custom" — it is named "harness"), causing a fall-through to the generic resolution path (source: env/config) at runtime_provider.py:1156
  • The generic path builds api_key_candidates from model.api_key (empty), host-gated OLLAMA/OPENAI/OPENROUTER keys, and _host_derived_api_key (returns "" for IP addresses)
  • The generic path does NOT resolve model.api_key_env or custom_providers.key_env — LITELLM_API_KEY is never read, producing api_key = "no-key-required" → HTTP 401
  • The named custom provider path (source: custom_provider:harness) DOES read key_env — but only triggers when provider matches the custom_providers[0].name
  • All sections MUST use provider: harness: model, compression, auxiliary.vision, auxiliary.web_extract, auxiliary.compression, delegation
  • Only fallback_providers uses a different provider (deepseek) for true fallback diversity
  • Diagnostic: If you see source: env/config in a request dump or log, the provider name is wrong. It should be source: custom_provider:harness.
  • Audit script: Run python3 /root/prose-contracts/audit-hermes-config.py <config.yaml> before and after any config change to catch this and all other rule violations.

Rule 15: MCP Endpoint and Header Validation (ADDED 2026-08-07)

  • Every MCP server entry must point at the correct endpoint:
  • MCP entries must carry a REAL key value in the header.
    • Avoid using env-var names like LITELLM_API_KEY in the header; they do not resolve for MCP endpoints and result in "Malformed API Key" floods.
    • Ensure the header value is the actual key (e.g., sk-...).

Execution

  1. Check current config — Read the target agent's config.yaml
  2. Compare against template — Identify missing or divergent sections
  3. Apply shared infra — Lock web/MCP/compression sections to template values
  4. Apply agent key — Set from agent_keys table above
  5. Set model choice — Per agent's workload
  6. Verify — curl all shared endpoints, test the model with the new key
  7. Report — What was changed, preserved, custom