fix: PR #148 timings corrected to cold ~5.7s / warm ~0.44s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Failing after 12s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped

- State the COLD figure (5.7s), not the warm one (0.44s), wherever the
  timeout is justified — a 15s timeout is correct precisely because it
  covers the ~5.7s cold scan, not because the scan costs 0.44s
- Do NOT assert a proven cause for the original failure: the contract had
  no documented timeout and no failure-kind rendering, so the honest
  wording is that a slow/cold scan exceeded whatever bound that run used
  and was rendered as a host-down verdict
- Keep the exit-status interpretation (124 vs 255) and the negative
  control exactly as they are — those are the real fix

Correlation: corr=9531ff9ce03ba876
This commit is contained in:
root
2026-10-02 11:18:10 +00:00
parent 6a55f5f860
commit 2512c5e85f
+1 -1
View File
@@ -188,7 +188,7 @@ Agent keys live in `.env` or `.env.vault` files with 600 permissions (koonimo's
Run on any Hermes host to detect violations.
**Timeout policy (2026-10-02):** The scan timeout is **15 seconds**, set from measured cost on the largest target (koby, 16 GB `.hermes` tree; full scan = 0.451 s, bounded scan = 0.441 s, over SSH, cold cache, measured 2026-10-02). The SSH connection timeout is **10 seconds** (separate from the scan timeout). A scan timeout renders as `probe-failed: <agent> <ip> (timeout after 15s)` — **never** as "unreachable" or "may be down". An SSH connection failure (exit status 255) renders as `unreachable: <agent> <ip> (ssh connect failed)`.
**Timeout policy (2026-10-02):** The scan timeout is **15 seconds**, set from measured cost on the largest target (koby, 16 GB `.hermes` tree; full scan: **cold ≈ 5.7 s**, warm ≈ 0.44 s; bounded scan: warm ≈ 0.37 s, over SSH, measured 2026-10-02). The 15 s bound is justified by the COLD cost, not the warm cost — a 13× cold/warm spread means the warm figure alone would understate the real worst case by an order of magnitude. The SSH connection timeout is **10 seconds** (separate from the scan timeout). A scan timeout renders as `probe-failed: <agent> <ip> (timeout after 15s)` — **never** as "unreachable" or "may be down". An SSH connection failure (exit status 255) renders as `unreachable: <agent> <ip> (ssh connect failed)`. The original failure (2026-10-02 koby) was a slow/cold scan that exceeded whatever bound the prior run used and was rendered as a host-down verdict; the exact prior timeout was never reproduced, so this is the only proven fix: honest failure-kind rendering plus the bounded scan.
**Bounded scan (2026-10-02):** Do NOT recurse the entire `/root/.hermes/` tree. Use `--exclude-dir=state-snapshots` to skip dated snapshot directories. Rationale: a superseded config will always carry a superseded key and will report forever with zero signal content (the koby state-snapshot line has repeated on consecutive days). If you deliberately want to include snapshots, say so in the contract and the report.