Files
prose-contracts/hermes-config-template.prose.md
jerome 0a3b62a598 hermes-config-template: update to v2.1.0 — max_tokens, auxiliary consistency, compression threshold
- Add model.max_tokens: 4096 (thermal safety — Rule 6)
- Add context_length: 262144 to template
- Add Rule 7: Auxiliary model consistency (gemma-4-12b + base_url + api_key_env)
- Add Rule 8: Compression threshold for 256K models (0.65, not 0.25)
- Fix compression section: threshold 0.65, max_context_window 262144, protect_last_n 40
- Remove stray api_key_env from compression block (not propagated)
- Update agent Keys table to show key aliases instead of redacted values
- Add base_url and timeout to all auxiliary service configs
- Add auxiliary: compression: section matching main compression config
- Update version 2.0.0 → 2.1.0
2026-07-04 17:54:43 -04:00

9.9 KiB

kind, name, description
kind name description
template hermes-config-template Standard Hermes configuration template for Syslog Solution LLC agents. Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, RA-H OS MCP) while keeping agent-specific API keys and model choices. Updated 2026-06-30: new LiteLLM keys, nginx routing, router back online.

Maintains

  • template_version: "2.1.0"
  • last_applied: timestamp
  • agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"]
  • agent_keys: map (see Agent Keys section)
  • infra_endpoints_verified: array

Agent Keys (LiteLLM — Current 2026-07-04)

Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM PostgreSQL DB via POST /key/generate on CT 116. Keys are stored in the DB, not in config files. The env var LITELLM_API_KEY is set in /etc/environment on each agent host AND in ~/.hermes/.env for gateway env propagation. Sub-agent profiles inherit auth from the main config — no separate keys needed.

Agent Key Alias Host SSH Sub-Agents
Tanko tanko-* 192.168.68.122 jerome@.122
Mumuni mumuni-jul2026 192.168.68.123 root@.123 6 profiles ✱
Abiba abiba-* 192.168.68.24 local
Koby koby-* ? Zulip
Koonimo koonimo-* 192.168.68.114 Zulip
Kagenz0 kagenz0-* ? Zulip

✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research, syslog-review, syslog-writer — all at /root/.hermes/profiles/<name>/config.yaml

Infrastructure Stack

Component Endpoint Purpose
Firecrawl http://192.168.68.7:3002/ Web content extraction
SearXNG http://storepve:8888 Privacy-respecting web search
LiteLLM http://192.168.68.116/v1 Unified model gateway (via nginx)
LiteLLM (NetBird) https://litellm.sysloggh.net/v1 Alternative (may have 502 issues)
RA-H OS MCP http://192.168.68.65:3100/mcp Knowledge graph bridge
Context7 MCP http://localhost:8079/mcp Documentation queries

API Key Rules

  • api_key_env: LITELLM_API_KEY — Use env var for main model auth (preferred)
  • api_key: '' — Sub-agents leave empty to inherit from main config's custom_provider
  • api_key: sk-... — Hardcoded key only as fallback when env var not possible
  • Set LITELLM_API_KEY in /etc/environment on each host
  • Sub-agents NEVER get their own key — they share the host agent's key
  • Restart Hermes after updating /etc/environment

Sub-Agent Profiles (Mumuni pattern)

Mumuni has 6 sub-agent profiles in /root/.hermes/profiles/<name>/config.yaml:

profiles/
├── syslog-code/config.yaml      # Code generation
├── syslog-devops/config.yaml     # DevOps/infrastructure
├── syslog-email/config.yaml      # Email processing
├── syslog-research/config.yaml   # Research & analysis
├── syslog-review/config.yaml     # Code review
└── syslog-writer/config.yaml     # Content writing

Sub-agent profile rules:

  1. api_key must be emptyapi_key: '' or omitted entirely
  2. base_url must be empty — inherits from main config's custom_provider
  3. provider is auto or harness — routes through the shared LiteLLM gateway
  4. model is agent-specific — each sub-agent can have its own default model
  5. Auxiliary tasks (vision, compression, etc.) also leave api_key empty
  6. Never hardcode a key in sub-agent profiles

This ensures all 6 sub-agents use the same LiteLLM key set in /etc/environment. When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs) work immediately after restart.

Template — Required Sections

# ─── Model Selection ───
model:
  default: <agent_model>          # e.g., ornith-1.0-35b, qwen3.6-27B-code
  provider: harness
  base_url: http://192.168.68.116/v1
  api_key_env: LITELLM_API_KEY    # Set in /etc/environment AND ~/.hermes/.env
  max_tokens: 4096                # ⚠️ CRITICAL: Prevents unbounded generation
  context_length: 262144          # Must match model's max context window

fallback_providers:
  provider: deepseek
  model: deepseek-chat
  api_key_env: DEEPSEEK_API_KEY

# ─── Web Stack (SHARED INFRA — DO NOT CHANGE) ───
web:
  backend: firecrawl
  search_backend: searxng
  extract_backend: firecrawl
  firecrawl:
    base_url: http://192.168.68.7:3002/

# ─── MCP Servers (SHARED INFRA) ───
mcp_servers:
  context7:
    connect_timeout: 60
    timeout: 300
    url: http://localhost:8079/mcp
  ra-h-os:
    url: http://192.168.68.65:3100/mcp
    timeout: 120
    connect_timeout: 60

# ─── Compression ───
compression:
  enabled: true
  model: gemma-4-12b              # ⚠️ Must match auxiliary.compression.model
  provider: harness
  max_context_window: 262144      # Must match model's actual capacity
  threshold: 0.65                 # Fires at ~170K for 262K window (not 0.25!)
  target_ratio: 0.30
  protect_last_n: 40
  hygiene_hard_message_limit: 350
  protect_first_n: 3
  abort_on_summary_failure: false

# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
# All auxiliary services MUST use identical model, base_url, and api_key_env:
#   model: gemma-4-12b
#   base_url: http://192.168.68.116/v1
#   api_key_env: LITELLM_API_KEY
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
# gemma-4-12b is a lightweight 12B model on the RTX 5070, freeing the Strix Halo
# for agent reasoning.
auxiliary:
  vision:
    provider: harness
    model: gemma-4-12b
    base_url: http://192.168.68.116/v1
    api_key_env: LITELLM_API_KEY
    timeout: 60
    download_timeout: 30
  web_extract:
    provider: harness
    model: gemma-4-12b
    base_url: http://192.168.68.116/v1
    api_key_env: LITELLM_API_KEY
    timeout: 30
  compression:
    provider: harness
    model: gemma-4-12b
    base_url: http://192.168.68.116/v1
    api_key_env: LITELLM_API_KEY
    timeout: 60

# ─── Custom Provider ───
custom_providers:
  - name: harness
    model: <agent_model>
    base_url: http://192.168.68.116/v1
    api_key_env: LITELLM_API_KEY
    api_mode: chat_completions

Key Update Procedure

When LiteLLM keys are regenerated (e.g., after infrastructure changes):

  1. If SSH available: ssh <host> "sudo sed -i 's/LITELLM_API_KEY=.*/LITELLM_API_KEY=sk-<NEW>/' /etc/environment"
  2. If SSH unavailable: Send Zulip DM via abiba-bot with update command
  3. After update: Restart Hermes on the agent host
  4. Verify: curl -H "Authorization: Bearer sk-<KEY>" http://192.168.68.116/v1/models

Configuration Rules

Rule 1: Shared Infra Is Locked

The following MUST be identical across ALL profiles:

  • web.backend, web.search_backend, web.extract_backend
  • web.firecrawl.base_url
  • mcp_servers.ra-h-os.url
  • custom_providers[0].base_url

Rule 2: Model Choice Is Free

  • model.default — per agent
  • fallback_providers.model — per agent
  • custom_providers[0].model — per agent

Rule 3: API Keys via Environment

  • Prefer api_key_env: LITELLM_API_KEY over hardcoded keys
  • Hardcoded keys in config.yaml become stale after key rotation
  • /etc/environment persists across config updates
  • Restart Hermes after env var updates

Rule 4: Sub-Agent Profiles Inherit Auth

  • Sub-agent profiles (/root/.hermes/profiles/*/config.yaml) must have:
    • api_key: '' — inherit from main config's custom_provider
    • base_url: '' — inherit from main config
    • Auxiliary tasks: api_key: '', provider: harness
  • Never hardcode a key in sub-agent profiles
  • When main config uses api_key_env, sub-agents automatically use it
  • This means key rotation only touches ONE file (/etc/environment)

Rule 5: Main Config Base URL

|- Use direct IP: http://192.168.68.116/v1 |- NOT the NetBird URL (litellm.sysloggh.net) — can cause 502 when NetBird is down |- NOT the old path (/litellm/v1) — nginx now routes /v1 directly

Rule 6: max_tokens Is Required (Thermal Safety)

  • Every Hermes config MUST set model.max_tokens: 4096 — this is non-negotiable
  • Prevents unbounded generation that caused the July 2 Strix Halo GPU thermal incident
  • The value flows through: config → agent.max_tokens → transport build_kwargs → API max_tokens
  • Even though the server now has -n 8192 hard cap (set by Abiba), the client cap is the first line of defense
  • Apply to BOTH main config AND all sub-agent profiles
  • For agents needing longer outputs: raise to 8192, but never omit

Rule 7: Auxiliary Model Consistency

  • All auxiliary services (vision, web_extract, compression) MUST use the same model:
    • model: gemma-4-12b
    • base_url: http://192.168.68.116/v1
    • api_key_env: LITELLM_API_KEY
  • Do NOT use syslog-auto for auxiliary tasks — it routes to the primary 35B reasoning GPU
  • gemma-4-12b is a lightweight 12B model on the RTX 5070, keeping the Strix Halo free for reasoning
  • The compression: block's model MUST match auxiliary: compression: model — they are two different configs for the same service

Rule 8: Compression Threshold for 256K Models

  • For 262K context window: threshold: 0.65 (fires at ~170K tokens)
  • Do NOT use threshold: 0.25 — this fires at 65K, causing premature context loss
  • Do NOT use threshold: 0.80 — this delays until 209K, risking the gateway hygiene layer
  • max_context_window: 262144 MUST match the model's actual capacity
  • See devops-hermes-compression skill for full reference

Execution

  1. Check current config — Read the target agent's config.yaml
  2. Compare against template — Identify missing or divergent sections
  3. Apply shared infra — Lock web/MCP/compression sections to template values
  4. Apply agent key — Set from agent_keys table above
  5. Set model choice — Per agent's workload
  6. Verify — curl all shared endpoints, test the model with the new key
  7. Report — What was changed, preserved, custom