fix(hermes): clarify auxiliary model policy to match audit
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 9s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
The 'Auxiliary Tasks (CONSISTENCY RULE)' header claimed ALL auxiliary services must use an identical model (gpu-vision) and that syslog-auto must never be used for auxiliary tasks. Both are false for compression, which the script (and Rule 7/8) require to be syslog-auto. An agent following the template produced a config the audit then failed. - Split auxiliary into two classes: light (vision, web_extract/browsing) -> gpu-vision (RTX 5070); context-heavy (compression) -> syslog-auto, citing the existing 2026-07-23 OPERATIONAL DECISION in Rule 7. - Remove the absolute 'Do NOT use syslog-auto for auxiliary tasks' line. - State gpu-dense + strix-moe are the reasoning hosts; do not pin aux to them. - Note the change in the frontmatter UPDATED log. audit-hermes-config.py unchanged (it is the enforcement contract); prose now matches it line-for-line on vision/web_extract/compression. Verified: prose-lint.sh PASS, secret-scan.sh clean, 14/14 tests in tests/test_audit_hermes_config_alias.py pass, audit PASSES on the template's stated policy.
This commit is contained in:
@@ -5,6 +5,11 @@ description: >
|
|||||||
Standard Hermes configuration template for Syslog Solution LLC agents.
|
Standard Hermes configuration template for Syslog Solution LLC agents.
|
||||||
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
Enforces shared infrastructure setup (Firecrawl, SearXNG, local models,
|
||||||
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
RA-H OS MCP) while keeping agent-specific API keys and model choices.
|
||||||
|
UPDATED 2026-09-27: Clarified the Auxiliary Tasks policy — light aux (vision,
|
||||||
|
web_extract/browsing) -> gpu-vision (RTX 5070); context-heavy aux (compression) ->
|
||||||
|
syslog-auto (2026-07-23 decision, Rule 7). Removed the false "one model for all
|
||||||
|
auxiliary" / "never syslog-auto" claim; stated gpu-dense + strix-moe are the reasoning
|
||||||
|
hosts and aux should not be pinned to them. Now matches audit-hermes-config.py line-for-line.
|
||||||
UPDATED 2026-08-07: Added litellm MCP server entry; updated Rule 15 (MCP Validation)
|
UPDATED 2026-08-07: Added litellm MCP server entry; updated Rule 15 (MCP Validation)
|
||||||
to enforce REAL key headers (not env-vars) from the 2026-08-07 keyless-MCP incident.
|
to enforce REAL key headers (not env-vars) from the 2026-08-07 keyless-MCP incident.
|
||||||
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
Added Rule 12 (Context-Issue Diagnostic) + Rule 13 (.env fallback enforcement) from the
|
||||||
@@ -167,13 +172,16 @@ compression:
|
|||||||
abort_on_summary_failure: false
|
abort_on_summary_failure: false
|
||||||
|
|
||||||
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
|
# ─── Auxiliary Tasks (CONSISTENCY RULE) ───
|
||||||
# All auxiliary services MUST use identical model, base_url, and api_key_env:
|
# Auxiliary tasks split into TWO model classes — do NOT assume one model for all:
|
||||||
# model: gpu-vision # stable alias (NOT a raw model name)
|
# Light auxiliary (vision, web_extract/browsing) -> model: gpu-vision # RTX 5070
|
||||||
|
# Keeps the reasoning hosts (gpu-dense / strix-moe) free for agent prompts.
|
||||||
|
# Context-heavy auxiliary (compression) -> model: syslog-auto # weighted pool
|
||||||
|
# Deliberate per the 2026-07-23 OPERATIONAL DECISION in Rule 7: summarization
|
||||||
|
# runs against long histories and must be able to use the pool.
|
||||||
|
# Do NOT pin auxiliary work to the reasoning hosts (gpu-dense / strix-moe).
|
||||||
|
# All auxiliary services share identical ROUTING (base_url + api_key_env), not model:
|
||||||
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
# base_url: http://192.168.68.116/litellm/v1 # Rule 5 (2026-08-09): canonical authenticated; /v1 also OK
|
||||||
# api_key_env: LITELLM_API_KEY
|
# api_key_env: LITELLM_API_KEY
|
||||||
# Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU.
|
|
||||||
# gpu-vision = RTX 5070 (12B), freeing the Strix Halo for agent reasoning.
|
|
||||||
# Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead.
|
|
||||||
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
|
# NEVER use retired model names (qwen3.6-27B-code, qwen3.6-35B-udq4; gemma-4-12b is retired
|
||||||
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
|
# and no longer resolves) in agent configs — use the stable aliases so model swaps don't break agents.
|
||||||
auxiliary:
|
auxiliary:
|
||||||
|
|||||||
Reference in New Issue
Block a user