diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index a1bdf88..76d48f4 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -5,6 +5,11 @@ description: > Standard Hermes configuration template for Syslog Solution LLC agents. Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, RA-H OS MCP) while keeping agent-specific API keys and model choices. + UPDATED 2026-09-27: Clarified the Auxiliary Tasks policy — light aux (vision, + web_extract/browsing) -> gpu-vision (RTX 5070); context-heavy aux (compression) -> + syslog-auto (2026-07-23 decision, Rule 7). Removed the false "one model for all + auxiliary" / "never syslog-auto" claim; stated gpu-dense + strix-moe are the reasoning + hosts and aux should not be pinned to them. Now matches audit-hermes-config.py line-for-line. UPDATED 2026-08-07: Added litellm MCP server entry; updated Rule 15 (MCP Validation) to enforce REAL key headers (not env-vars) from the 2026-08-07 keyless-MCP incident. Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the @@ -167,13 +172,16 @@ compression: abort_on_summary_failure: false # ─── Auxiliary Tasks (CONSISTENCY RULE) ─── -# All auxiliary services MUST use identical model, base_url, and api_key_env: -# model: gpu-vision # stable alias (NOT a raw model name) +# Auxiliary tasks split into TWO model classes — do NOT assume one model for all: +# Light auxiliary (vision, web_extract/browsing) -> model: gpu-vision # RTX 5070 +# Keeps the reasoning hosts (gpu-dense / strix-moe) free for agent prompts. +# Context-heavy auxiliary (compression) -> model: syslog-auto # weighted pool +# Deliberate per the 2026-07-23 OPERATIONAL DECISION in Rule 7: summarization +# runs against long histories and must be able to use the pool. +# Do NOT pin auxiliary work to the reasoning hosts (gpu-dense / strix-moe). +# All auxiliary services share identical ROUTING (base_url + api_key_env), not model: # base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK # api_key_env: LITELLM_API_KEY -# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU. -# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning. -# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead. # NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired # and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents. auxiliary: