diff --git a/hermes-agent-baseline.prose.md b/hermes-agent-baseline.prose.md index 8f814b8..99686e0 100644 --- a/hermes-agent-baseline.prose.md +++ b/hermes-agent-baseline.prose.md @@ -121,6 +121,26 @@ auxiliary: timeout: 120 ``` +## Violation Classification + +When reporting findings, separate POLICY observations from FAULT findings: + +### POLICY (observation only, not a fault) +- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter) +- Config text has a field that looks unusual but the agent's calls are succeeding +- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour" + +### FAULT (requires request-level evidence) +- Agent's calls are failing with auth errors (401/403 in logs) +- Agent's config has no valid API key AND calls are failing +- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00" + +### Rules +1. Do NOT infer the runtime's credential resolution from config text alone. +2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window. +3. If calls are succeeding, the correct output is "POLICY: uses directly; calls succeeding" - not a violation. +4. State what you OBSERVED, not what the field implies. + ## Known Bug: `api_key_env` Ignored by Auxiliary Client **Bug location**: `agent/auxiliary_client.py` → `_resolve_task_provider_model()` (line ~5478) diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index b04d93e..0a1a6ef 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -217,6 +217,26 @@ When LiteLLM keys are regenerated (e.g., after infrastructure changes): 3. **After update**: Restart Hermes on the agent host 4. **Verify**: `curl -H "Authorization: Bearer sk-" http://192.168.68.116/v1/models` +## Violation Classification + +When reporting findings, separate POLICY observations from FAULT findings: + +### POLICY (observation only, not a fault) +- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter) +- Config text has a field that looks unusual but the agent's calls are succeeding +- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour" + +### FAULT (requires request-level evidence) +- Agent's calls are failing with auth errors (401/403 in logs) +- Agent's config has no valid API key AND calls are failing +- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00" + +### Rules +1. Do NOT infer the runtime's credential resolution from config text alone. +2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window. +3. If calls are succeeding, the correct output is "POLICY: uses directly; calls succeeding" - not a violation. +4. State what you OBSERVED, not what the field implies. + ## Configuration Rules ### Rule 1: Shared Infra Is Locked diff --git a/hermes-key-enforcement.prose.md b/hermes-key-enforcement.prose.md index ee38b0d..06fb821 100644 --- a/hermes-key-enforcement.prose.md +++ b/hermes-key-enforcement.prose.md @@ -135,6 +135,26 @@ scripts/hermes-reachability-check.sh "api_key: sk-" "/root/.hermes/" # The correct pattern: remote side always succeeds (grep ...; true), so ssh status = connection only. ``` +## Violation Classification + +When reporting findings, separate POLICY observations from FAULT findings: + +### POLICY (observation only, not a fault) +- Agent uses a non-internal-harness provider (e.g., direct DeepSeek, Tencent, OpenRouter) +- Config text has a field that looks unusual but the agent's calls are succeeding +- Example: "POLICY: Koonimo uses deepseek directly; calls succeeding in last hour" + +### FAULT (requires request-level evidence) +- Agent's calls are failing with auth errors (401/403 in logs) +- Agent's config has no valid API key AND calls are failing +- Example: "FAULT: Koby's LiteLLM key expired; 401 observed at 2026-09-14 11:42:00" + +### Rules +1. Do NOT infer the runtime's credential resolution from config text alone. +2. Require request-level evidence before calling something a FAULT: an observed auth failure in the agent's log, or the absence of successful calls in the window. +3. If calls are succeeding, the correct output is "POLICY: uses directly; calls succeeding" - not a violation. +4. State what you OBSERVED, not what the field implies. + ## Detection Query Run on any Hermes host to detect violations: