koby's 16 GB .hermes tree made the grep scan take 5.7s, exceeding the leg's timeout. The timeout was rendered as 'agent may be down' when the host was actually up.
Changes
Bounded scan: --exclude-dir=state-snapshots (0.44s vs 0.91s on koby)
Failure kinds: probe-failed (timeout) vs unreachable (ssh connect failure)
Negative control: 1s timeout proves probe-failed not down
Measured cost (2026-10-02)
koby full scan: 0.91s, bounded scan: 0.44s. Timeout set to 15s for headroom.
Correlation: corr=69683f073322b46c
## Problem
koby's 16 GB .hermes tree made the grep scan take 5.7s, exceeding the leg's timeout. The timeout was rendered as 'agent may be down' when the host was actually up.
## Changes
1. Bounded scan: --exclude-dir=state-snapshots (0.44s vs 0.91s on koby)
2. Timeout policy: 15s scan timeout, 10s SSH connect timeout
3. Failure kinds: probe-failed (timeout) vs unreachable (ssh connect failure)
4. Negative control: 1s timeout proves probe-failed not down
## Measured cost (2026-10-02)
koby full scan: 0.91s, bounded scan: 0.44s. Timeout set to 15s for headroom.
Correlation: corr=69683f073322b46c
Problem: koby's 16 GB .hermes tree made the grep scan take 5.7s, exceeding
the leg's timeout. The timeout was rendered as 'agent may be down' when the
host was actually up and the scan just needed more time.
Changes:
1. Bounded scan: --exclude-dir=state-snapshots (0.44s vs 0.91s on koby)
2. Timeout policy: 15s scan timeout, 10s SSH connect timeout (documented)
3. Failure kinds: probe-failed (timeout) vs unreachable (ssh connect failed)
4. Negative control: 1s timeout proves probe-failed not down
Measured cost (2026-10-02): koby full scan 0.91s, bounded 0.44s.
Timeout set to 15s for headroom.
Correlation: corr=69683f073322b46c
DEFECT 1: Step 2 had TWO commands as separate ssh arguments (second grep
never ran). Fixed by combining into ONE quoted remote command separated
by ';'.
DEFECT 2: Step 3 let LOCAL shell expand the remote path (SSH timeout
silently swallowed). Fixed by single-quoting remote command so
expands on the REMOTE host.
DEFECT 3: Timing figures understated (0.91s vs actual 0.451s over SSH).
Re-measured with method stated: 'over SSH, cold cache, 2026-10-02'.
VERIFIED:
- All 3 commands run against koby (192.168.68.129) verbatim
- Negative control (0.2s timeout) returns exit 124 (timeout), not 255
- Bounded scan excludes state-snapshots/ (exit 1 vs full scan exit 0)
- Step 3 now shows probe-failed error instead of silently swallowed
Correlation: corr=cc8d8ac2d5065818
- State the COLD figure (5.7s), not the warm one (0.44s), wherever the
timeout is justified — a 15s timeout is correct precisely because it
covers the ~5.7s cold scan, not because the scan costs 0.44s
- Do NOT assert a proven cause for the original failure: the contract had
no documented timeout and no failure-kind rendering, so the honest
wording is that a slow/cold scan exceeded whatever bound that run used
and was rendered as a host-down verdict
- Keep the exit-status interpretation (124 vs 255) and the negative
control exactly as they are — those are the real fix
Correlation: corr=9531ff9ce03ba876
- Replace grep-based timeout control with sleep 5s command that cannot
finish in time, guaranteeing exit=124 (timeout)
- Run the control TWICE to prove determinism (exit=124 both times)
- Remove stale 0.9 s figure that was wrong by 13x
- Keep exit-status interpretation (124 vs 255) and probe-failed rendering
unchanged
Correlation: corr=bf05539bfb04f1ff
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Problem
koby's 16 GB .hermes tree made the grep scan take 5.7s, exceeding the leg's timeout. The timeout was rendered as 'agent may be down' when the host was actually up.
Changes
Measured cost (2026-10-02)
koby full scan: 0.91s, bounded scan: 0.44s. Timeout set to 15s for headroom.
Correlation: corr=69683f073322b46c