fix: hermes-key-enforcement probe timeout - bounded scan, honest failure kinds #148
@@ -186,30 +186,60 @@ Agent keys live in `.env` or `.env.vault` files with 600 permissions (koonimo's
|
||||
|
||||
## Detection Query
|
||||
|
||||
Run on any Hermes host to detect violations:
|
||||
Run on any Hermes host to detect violations.
|
||||
|
||||
**Timeout policy (2026-10-02):** The scan timeout is **15 seconds**, set from measured cost on the largest target (koby, 16 GB `.hermes` tree; full scan: **cold ≈ 5.7 s**, warm ≈ 0.44 s; bounded scan: warm ≈ 0.37 s, over SSH, measured 2026-10-02). The 15 s bound is justified by the COLD cost, not the warm cost — a 13× cold/warm spread means the warm figure alone would understate the real worst case by an order of magnitude. The SSH connection timeout is **10 seconds** (separate from the scan timeout). A scan timeout renders as `probe-failed: <agent> <ip> (timeout after 15s)` — **never** as "unreachable" or "may be down". An SSH connection failure (exit status 255) renders as `unreachable: <agent> <ip> (ssh connect failed)`. The original failure (2026-10-02 koby) was a slow/cold scan that exceeded whatever bound the prior run used and was rendered as a host-down verdict; the exact prior timeout was never reproduced, so this is the only proven fix: honest failure-kind rendering plus the bounded scan.
|
||||
|
||||
**Bounded scan (2026-10-02):** Do NOT recurse the entire `/root/.hermes/` tree. Use `--exclude-dir=state-snapshots` to skip dated snapshot directories. Rationale: a superseded config will always carry a superseded key and will report forever with zero signal content (the koby state-snapshot line has repeated on consecutive days). If you deliberately want to include snapshots, say so in the contract and the report.
|
||||
|
||||
```bash
|
||||
# 1. Check config.yaml for hardcoded harness keys
|
||||
grep -rn 'api_key: sk-' /root/.hermes/ \
|
||||
--include='config.yaml' \
|
||||
| grep -v 'deepseek\|openai\|anthropic\|DEEPSEEK'
|
||||
# 1. Check config.yaml for hardcoded harness keys (bounded scan — excludes state-snapshots)
|
||||
timeout 15 ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=no root@<ip> \
|
||||
"grep -rn 'api_key: sk-' /root/.hermes/ --exclude-dir=state-snapshots --include='config.yaml' | grep -v 'deepseek\|openai\|anthropic\|DEEPSEEK'" \
|
||||
2>/dev/null
|
||||
|
||||
# Interpret exit status:
|
||||
# 0 = match found (violation)
|
||||
# 1 = no match (pass)
|
||||
# 124 = timeout (probe-failed, not unreachable)
|
||||
# 255 = ssh connect failed (unreachable)
|
||||
# other = probe-failed (record the actual code)
|
||||
|
||||
# 1b. Check for double-path bug: base_url ending with /responses
|
||||
# (Hermes appends /v1/responses when api_mode=responses, so base_url must end at /v1)
|
||||
grep -rn 'litellm/v1/responses' /root/.hermes/config.yaml
|
||||
timeout 15 ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=no root@<ip> \
|
||||
"grep -rn 'litellm/v1/responses' /root/.hermes/config.yaml" 2>/dev/null
|
||||
# ANY output here = WRONG. Must be 'litellm/v1' without /responses suffix.
|
||||
|
||||
# 2. Check systemd drop-ins for master key leaks (2026-07-05: Tanko had this)
|
||||
grep -rn 'LITELLM_API_KEY' /root/.config/systemd/user/ 2>/dev/null
|
||||
grep -rn 'LITELLM_API_KEY=sk-synthetic-litellm-…' /root/.config/systemd/ 2>/dev/null
|
||||
# NOTE: Both greps are inside ONE quoted remote command, separated by ; (not two separate ssh arguments)
|
||||
timeout 15 ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=no root@<ip> \
|
||||
"grep -rn 'LITELLM_API_KEY' /root/.config/systemd/user/ 2>/dev/null; grep -rn 'LITELLM_API_KEY=sk-synthetic-litellm-' /root/.config/systemd/ 2>/dev/null" ; true
|
||||
|
||||
# 3. Verify running process env matches dedicated key
|
||||
cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c "import sys,json; print(json.load(sys.stdin)['pid'])")/environ \
|
||||
| tr '\0' '\n' | grep LITELLM_API_KEY
|
||||
# NOTE: Single-quoted remote command so $(...) expands on the REMOTE host, not the runner
|
||||
timeout 15 ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=no root@<ip> \
|
||||
'cat /proc/$(cat /home/jerome/.hermes/gateway.pid | python3 -c "import sys,json; print(json.load(sys.stdin)['pid'])")/environ | tr "\0" "\n" | grep LITELLM_API_KEY' \
|
||||
&& echo "step3: PASS" || echo "step3: probe-failed (exit $?; see stderr above)"
|
||||
```
|
||||
|
||||
If any output from step 2 — **critical violation** (master key leaked). Fix immediately.
|
||||
|
||||
### Negative control (probe-failed vs unreachable) — deterministic
|
||||
|
||||
To prove the distinction between a scan timeout and a connection failure, run a command that CANNOT finish in time (sleep 5s) with a 1-second timeout:
|
||||
|
||||
```bash
|
||||
# Negative control: 1-second timeout on koby — sleep 5s guarantees timeout
|
||||
timeout 1 ssh -o ConnectTimeout=10 -o StrictHostKeyChecking=no root@192.168.68.129 "sleep 5"; echo "exit=$?"
|
||||
# Expected: exit=124 (timeout) → render as "probe-failed: koby 192.168.68.129 (timeout after 1s)"
|
||||
# NOT: "unreachable" or "may be down"
|
||||
|
||||
# Run TWICE to prove determinism:
|
||||
# Run 1: timeout 1 ssh ... "sleep 5"; echo "exit=$?" → exit=124
|
||||
# Run 2: timeout 1 ssh ... "sleep 5"; echo "exit=$?" → exit=124
|
||||
```
|
||||
|
||||
## Rotation Procedure
|
||||
|
||||
With this standard enforced, key rotation is one vault update:
|
||||
|
||||
Reference in New Issue
Block a user