--- kind: template name: hermes-config-template description: > Standard Hermes configuration template for Syslog Solution LLC agents. Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, RA-H OS MCP) while keeping agent-specific API keys and model choices. Updated 2026-06-30: new LiteLLM keys, nginx routing, router back online. --- ## Maintains - template_version: "2.1.0" - last_applied: timestamp - agents_configured: ["tanko", "mumuni", "abiba", "koby", "koonimo", "kagenz0"] - agent_keys: map (see Agent Keys section) - infra_endpoints_verified: array ## Agent Keys (LiteLLM — Current 2026-07-04) Each agent has a unique LiteLLM API key (virtual key) generated against the LiteLLM PostgreSQL DB via `POST /key/generate` on CT 116. Keys are stored in the DB, not in config files. The env var `LITELLM_API_KEY` is set in `/etc/environment` on each agent host AND in `~/.hermes/.env` for gateway env propagation. Sub-agent profiles inherit auth from the main config — no separate keys needed. | Agent | Key Alias | Host | SSH | Sub-Agents | |-------|-----------|------|-----|-----------| | Tanko | `tanko-*` | 192.168.68.122 | jerome@.122 | — | | Mumuni | `mumuni-jul2026` | 192.168.68.123 | root@.123 | 6 profiles ✱ | | Abiba | `abiba-*` | 192.168.68.24 | local | — | | Koby | `koby-*` | ? | Zulip | — | | Koonimo | `koonimo-*` | 192.168.68.114 | Zulip | — | | Kagenz0 | `kagenz0-*` | ? | Zulip | — | ✱ Mumuni sub-agents: syslog-code, syslog-devops, syslog-email, syslog-research, syslog-review, syslog-writer — all at `/root/.hermes/profiles//config.yaml` ## Infrastructure Stack | Component | Endpoint | Purpose | |---|---|---| | Firecrawl | `http://192.168.68.7:3002/` | Web content extraction | | SearXNG | `http://storepve:8888` | Privacy-respecting web search | | LiteLLM | `http://192.168.68.116/v1` | Unified model gateway (via nginx) | | LiteLLM (NetBird) | `https://litellm.sysloggh.net/v1` | Alternative (may have 502 issues) | | RA-H OS MCP | `http://192.168.68.65:3100/mcp` | Knowledge graph bridge | | Context7 MCP | `http://localhost:8079/mcp` | Documentation queries | ## API Key Rules - `api_key_env: LITELLM_API_KEY` — Use env var for main model auth (preferred) - `api_key: ''` — Sub-agents leave empty to inherit from main config's custom_provider - `api_key: sk-...` — Hardcoded key only as fallback when env var not possible - Set `LITELLM_API_KEY` in `/etc/environment` on each host - Sub-agents NEVER get their own key — they share the host agent's key - Restart Hermes after updating `/etc/environment` ### Sub-Agent Profiles (Mumuni pattern) Mumuni has 6 sub-agent profiles in `/root/.hermes/profiles//config.yaml`: ``` profiles/ ├── syslog-code/config.yaml # Code generation ├── syslog-devops/config.yaml # DevOps/infrastructure ├── syslog-email/config.yaml # Email processing ├── syslog-research/config.yaml # Research & analysis ├── syslog-review/config.yaml # Code review └── syslog-writer/config.yaml # Content writing ``` Sub-agent profile rules: 1. **`api_key` must be empty** — `api_key: ''` or omitted entirely 2. **`base_url` must be empty** — inherits from main config's custom_provider 3. **`provider` is `auto` or `harness`** — routes through the shared LiteLLM gateway 4. **`model` is agent-specific** — each sub-agent can have its own default model 5. **Auxiliary tasks** (vision, compression, etc.) also leave `api_key` empty 6. **Never hardcode a key** in sub-agent profiles This ensures all 6 sub-agents use the same LiteLLM key set in `/etc/environment`. When the key is rotated, only the env var needs updating — all 7 configs (main + 6 subs) work immediately after restart. ## Template — Required Sections ```yaml # ─── Model Selection ─── model: default: # e.g., ornith-1.0-35b, qwen3.6-27B-code provider: harness base_url: http://192.168.68.116/v1 api_key_env: LITELLM_API_KEY # Set in /etc/environment AND ~/.hermes/.env max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation context_length: 262144 # Must match model's max context window fallback_providers: provider: deepseek model: deepseek-chat api_key_env: DEEPSEEK_API_KEY # ─── Web Stack (SHARED INFRA — DO NOT CHANGE) ─── web: backend: firecrawl search_backend: searxng extract_backend: firecrawl firecrawl: base_url: http://192.168.68.7:3002/ # ─── MCP Servers (SHARED INFRA) ─── mcp_servers: context7: connect_timeout: 60 timeout: 300 url: http://localhost:8079/mcp ra-h-os: url: http://192.168.68.65:3100/mcp timeout: 120 connect_timeout: 60 # ─── Compression ─── compression: enabled: true model: gemma-4-12b # ⚠️ Must match auxiliary.compression.model provider: harness max_context_window: 262144 # Must match model's actual capacity threshold: 0.65 # Fires at ~170K for 262K window (not 0.25!) target_ratio: 0.30 protect_last_n: 40 hygiene_hard_message_limit: 350 protect_first_n: 3 abort_on_summary_failure: false # ─── Auxiliary Tasks (CONSISTENCY RULE) ─── # All auxiliary services MUST use identical model, base_url, and api_key_env: # model: gemma-4-12b # base_url: http://192.168.68.116/v1 # api_key_env: LITELLM_API_KEY # Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU. # gemma-4-12b is a lightweight 12B model on the RTX 5070, freeing the Strix Halo # for agent reasoning. auxiliary: vision: provider: harness model: gemma-4-12b base_url: http://192.168.68.116/v1 api_key_env: LITELLM_API_KEY timeout: 60 download_timeout: 30 web_extract: provider: harness model: gemma-4-12b base_url: http://192.168.68.116/v1 api_key_env: LITELLM_API_KEY timeout: 30 compression: provider: harness model: gemma-4-12b base_url: http://192.168.68.116/v1 api_key_env: LITELLM_API_KEY timeout: 60 # ─── Custom Provider ─── custom_providers: - name: harness model: base_url: http://192.168.68.116/v1 api_key_env: LITELLM_API_KEY api_mode: chat_completions ``` ## Key Update Procedure When LiteLLM keys are regenerated (e.g., after infrastructure changes): 1. **If SSH available**: `ssh "sudo sed -i 's/LITELLM_API_KEY=.*/LITELLM_API_KEY=sk-/' /etc/environment"` 2. **If SSH unavailable**: Send Zulip DM via abiba-bot with update command 3. **After update**: Restart Hermes on the agent host 4. **Verify**: `curl -H "Authorization: Bearer sk-" http://192.168.68.116/v1/models` ## Configuration Rules ### Rule 1: Shared Infra Is Locked The following MUST be identical across ALL profiles: - `web.backend`, `web.search_backend`, `web.extract_backend` - `web.firecrawl.base_url` - `mcp_servers.ra-h-os.url` - `custom_providers[0].base_url` ### Rule 2: Model Choice Is Free - `model.default` — per agent - `fallback_providers.model` — per agent - `custom_providers[0].model` — per agent ### Rule 3: API Keys via Environment - Prefer `api_key_env: LITELLM_API_KEY` over hardcoded keys - Hardcoded keys in config.yaml become stale after key rotation - `/etc/environment` persists across config updates - Restart Hermes after env var updates ### Rule 4: Sub-Agent Profiles Inherit Auth - Sub-agent profiles (`/root/.hermes/profiles/*/config.yaml`) must have: - `api_key: ''` — inherit from main config's custom_provider - `base_url: ''` — inherit from main config - Auxiliary tasks: `api_key: ''`, `provider: harness` - Never hardcode a key in sub-agent profiles - When main config uses `api_key_env`, sub-agents automatically use it - This means key rotation only touches ONE file (`/etc/environment`) ### Rule 5: Main Config Base URL |- Use direct IP: `http://192.168.68.116/v1` |- NOT the NetBird URL (`litellm.sysloggh.net`) — can cause 502 when NetBird is down |- NOT the old path (`/litellm/v1`) — nginx now routes `/v1` directly ### Rule 6: max_tokens Is Required (Thermal Safety) - **Every Hermes config MUST set `model.max_tokens: 4096`** — this is non-negotiable - Prevents unbounded generation that caused the July 2 Strix Halo GPU thermal incident - The value flows through: `config → agent.max_tokens → transport build_kwargs → API max_tokens` - Even though the server now has `-n 8192` hard cap (set by Abiba), the client cap is the first line of defense - Apply to BOTH main config AND all sub-agent profiles - For agents needing longer outputs: raise to 8192, but never omit ### Rule 7: Auxiliary Model Consistency - All auxiliary services (vision, web_extract, compression) MUST use the same model: - `model: gemma-4-12b` - `base_url: http://192.168.68.116/v1` - `api_key_env: LITELLM_API_KEY` - **Do NOT use `syslog-auto`** for auxiliary tasks — it routes to the primary 35B reasoning GPU - gemma-4-12b is a lightweight 12B model on the RTX 5070, keeping the Strix Halo free for reasoning - The `compression:` block's `model` MUST match `auxiliary: compression: model` — they are two different configs for the same service ### Rule 8: Compression Threshold for 256K Models - For 262K context window: `threshold: 0.65` (fires at ~170K tokens) - Do NOT use `threshold: 0.25` — this fires at 65K, causing premature context loss - Do NOT use `threshold: 0.80` — this delays until 209K, risking the gateway hygiene layer - `max_context_window: 262144` MUST match the model's actual capacity - See `devops-hermes-compression` skill for full reference ## Execution 1. **Check current config** — Read the target agent's config.yaml 2. **Compare against template** — Identify missing or divergent sections 3. **Apply shared infra** — Lock web/MCP/compression sections to template values 4. **Apply agent key** — Set from agent_keys table above 5. **Set model choice** — Per agent's workload 6. **Verify** — curl all shared endpoints, test the model with the new key 7. **Report** — What was changed, preserved, custom