From 9a789ab76d3493baa4e71b96a186c70a74c090e8 Mon Sep 17 00:00:00 2001 From: root Date: Mon, 14 Sep 2026 12:32:06 +0000 Subject: [PATCH] fix(hermes): separate policy observations from fault findings Contract design defect: 'uses a non-harness provider' (POLICY) and 'cannot authenticate' (FAULT) were printed as the same violation class. A policy observation must never be phrased as if the agent were broken. Changes: 1. Added 'Violation Classification' section to all three contracts 2. Separated POLICY (observation only) from FAULT (requires request-level evidence) 3. Rules: - Do NOT infer runtime credential resolution from config text alone - Require request-level evidence before calling a FAULT: observed auth failure or absence of successful calls - If calls are succeeding, output is 'POLICY: uses directly; calls succeeding' - not a violation - State what you OBSERVED, not what the field implies Files changed (3): - hermes-key-enforcement.prose.md - hermes-config-template.prose.md - hermes-agent-baseline.prose.md --- hermes-agent-baseline.prose.md | 20 ++++++++++++++++++++ hermes-config-template.prose.md | 20 ++++++++++++++++++++ hermes-key-enforcement.prose.md | 20 ++++++++++++++++++++ 3 files changed, 60 insertions(+) diff --git a/hermes-agent-baseline.prose.md b/hermes-agent-baseline.prose.md index 8f814b8..99686e0 100644 --- a/hermes-agent-baseline.prose.md +++ b/hermes-agent-baseline.prose.md @@ -121,6 +121,26 @@ auxiliary: timeout: 120 ``` +## Violation Classification + +When reporting findings, separate POLICY observations from FAULT findings: + +### POLICY (observation only, not a fault) +- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter) +- Config text has a field that looks unusual but the agent's calls are succeeding +- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour" + +### FAULT (requires request-level evidence) +- Agent's calls are failing with auth errors (401/403 in logs) +- Agent's config has no valid API key AND calls are failing +- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00" + +### Rules +1. Do NOT infer the runtime's credential resolution from config text alone. +2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window. +3. If calls are succeeding, the correct output is "POLICY: uses directly; calls succeeding" - not a violation. +4. State what you OBSERVED, not what the field implies. + ## Known Bug: `api_key_env` Ignored by Auxiliary Client **Bug location**: `agent/auxiliary_client.py` → `_resolve_task_provider_model()` (line ~5478) diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index b04d93e..0a1a6ef 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -217,6 +217,26 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes): 3. **After update**: Restart Hermes on the agent host 4. **Verify**: `curl -H "Authorization: Bearer sk-" http://192.168.68.116/v1/models` +## Violation Classification + +When reporting findings, separate POLICY observations from FAULT findings: + +### POLICY (observation only, not a fault) +- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter) +- Config text has a field that looks unusual but the agent's calls are succeeding +- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour" + +### FAULT (requires request-level evidence) +- Agent's calls are failing with auth errors (401/403 in logs) +- Agent's config has no valid API key AND calls are failing +- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00" + +### Rules +1. Do NOT infer the runtime's credential resolution from config text alone. +2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window. +3. If calls are succeeding, the correct output is "POLICY: uses directly; calls succeeding" - not a violation. +4. State what you OBSERVED, not what the field implies. + ## Configuration Rules ### Rule 1: Shared Infra Is Locked diff --git a/hermes-key-enforcement.prose.md b/hermes-key-enforcement.prose.md index ee38b0d..06fb821 100644 --- a/hermes-key-enforcement.prose.md +++ b/hermes-key-enforcement.prose.md @@ -135,6 +135,26 @@ scripts/hermes-reachability-check.sh "api_key: sk-" "/root/.hermes/" # The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only. ``` +## Violation Classification + +When reporting findings, separate POLICY observations from FAULT findings: + +### POLICY (observation only, not a fault) +- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter) +- Config text has a field that looks unusual but the agent's calls are succeeding +- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour" + +### FAULT (requires request-level evidence) +- Agent's calls are failing with auth errors (401/403 in logs) +- Agent's config has no valid API key AND calls are failing +- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00" + +### Rules +1. Do NOT infer the runtime's credential resolution from config text alone. +2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window. +3. If calls are succeeding, the correct output is "POLICY: uses directly; calls succeeding" - not a violation. +4. State what you OBSERVED, not what the field implies. + ## Detection Query Run on any Hermes host to detect violations: