Closes the flap that produced "0/4 agent gateways reachable" three hours after the same contract reported "4/4", while both gateways were live.
The reported defect
On 2026-09-13 the agent-health-check contract reported 0/4 agent gateways reachable while koby's and koonimo's gateway processes were running and root SSH to their hosts worked (firstmate verified the processes at the time). The leg could not tell "I could not reach it" from "it is not running", and both were rendered as an outage.
Changes (scripts/agent-health-check.py, 64+/26-)
_ssh_retry(): one retry at a longer timeout (15s then 25s), returning (output, probe_failed, fail_kind) instead of collapsing every failure into "not reachable". Used by the gateway-PID lookup, the gateway state file, the streaming check and the recent-errors read.
A failed probe now says so: probe-failed: ssh <user>@<host> <kind> (retried at 25s: also <kind>) ... gateway status UNDETERMINED, recorded under probe-failed:<name>:<kind>. "Undetermined" is a different claim from "down", and the previous unreachable:<name>/gateway-down:<name> collapse is what made a transient SSH blip look like a fleet outage.
Every line names the probe target (ssh <user>@<host>, plus the CT), so a reader can re-run the exact command.
Koby stays report-only: its legs are detected and reported, never counted as fleet failures and never repaired (captain's 2026-08-17 ruling), and the report-only line now names the probe it ran.
Verification performed by firstmate before this PR
Read the whole three-dot diff: the retry is applied to every SSH leg in the agent section, and the report-only branch is intact with no repair path.
Ran the branch's script unmodified: it completes with All checks passed, and koby's report-only legs render as 🔍 report-only (koby): ... reported, not counted/repaired rather than as failures.
Reviewers: (1) confirm no path can print an agent as unreachable/down without the probe target and without the retry having run; (2) confirm a genuine gateway outage still reports clearly - walk the "process really is absent" case and say whether the output distinguishes it from an unreachable probe; (3) confirm koby's report-only handling is unchanged in substance (detected, reported, never counted, never repaired); (4) check whether anything consumes the failure KEYS that changed (unreachable:<name> -> probe-failed:<name>:<kind>), since a rename can silently break an alert or digest consumer; (5) PASS/FAIL with the PR number. Do NOT merge.
Closes the flap that produced "0/4 agent gateways reachable" three hours after the same contract reported "4/4", while both gateways were live.
## The reported defect
On 2026-09-13 the agent-health-check contract reported `0/4 agent gateways reachable` while koby's and koonimo's gateway processes were running and root SSH to their hosts worked (firstmate verified the processes at the time). The leg could not tell "I could not reach it" from "it is not running", and both were rendered as an outage.
## Changes (`scripts/agent-health-check.py`, 64+/26-)
- **`_ssh_retry()`**: one retry at a longer timeout (15s then 25s), returning `(output, probe_failed, fail_kind)` instead of collapsing every failure into "not reachable". Used by the gateway-PID lookup, the gateway state file, the streaming check and the recent-errors read.
- **A failed probe now says so**: `probe-failed: ssh <user>@<host> <kind> (retried at 25s: also <kind>) ... gateway status UNDETERMINED`, recorded under `probe-failed:<name>:<kind>`. "Undetermined" is a different claim from "down", and the previous `unreachable:<name>`/`gateway-down:<name>` collapse is what made a transient SSH blip look like a fleet outage.
- **Every line names the probe target** (`ssh <user>@<host>`, plus the CT), so a reader can re-run the exact command.
- **Koby stays report-only**: its legs are detected and reported, never counted as fleet failures and never repaired (captain's 2026-08-17 ruling), and the report-only line now names the probe it ran.
## Verification performed by firstmate before this PR
- Read the whole three-dot diff: the retry is applied to every SSH leg in the agent section, and the report-only branch is intact with no repair path.
- Ran the branch's script unmodified: it completes with `All checks passed`, and koby's report-only legs render as `🔍 report-only (koby): ... reported, not counted/repaired` rather than as failures.
Reviewers: (1) confirm no path can print an agent as unreachable/down without the probe target and without the retry having run; (2) confirm a genuine gateway outage still reports clearly - walk the "process really is absent" case and say whether the output distinguishes it from an unreachable probe; (3) confirm koby's report-only handling is unchanged in substance (detected, reported, never counted, never repaired); (4) check whether anything consumes the failure KEYS that changed (`unreachable:<name>` -> `probe-failed:<name>:<kind>`), since a rename can silently break an alert or digest consumer; (5) PASS/FAIL with the PR number. Do NOT merge.
Per defect report 1154.msg (agent-health-gateway-leg-flap-20260913):
1. Add _ssh_retry() helper with one retry at longer timeout (25s)
2. Name probe target explicitly: ssh {user}@{host} <command>
3. Print probe-failed when first attempt fails, then retry
4. Only declare gateway-down after retry fails
5. Make Koby report-only explicit in output
Changes:
- check_agents(): all gateway probes now use _ssh_retry()
- All output lines name the probe target (ssh host:port)
- Koby's report-only status is explicit in output
- Never print bare "gateway down" — always name target and failure kind
Verified: koonimo shows "probe-failed" on first attempt (transient SSH),
retries at 25s, succeeds, reports ✅ koonimo: gw=running
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Closes the flap that produced "0/4 agent gateways reachable" three hours after the same contract reported "4/4", while both gateways were live.
The reported defect
On 2026-09-13 the agent-health-check contract reported
0/4 agent gateways reachablewhile koby's and koonimo's gateway processes were running and root SSH to their hosts worked (firstmate verified the processes at the time). The leg could not tell "I could not reach it" from "it is not running", and both were rendered as an outage.Changes (
scripts/agent-health-check.py, 64+/26-)_ssh_retry(): one retry at a longer timeout (15s then 25s), returning(output, probe_failed, fail_kind)instead of collapsing every failure into "not reachable". Used by the gateway-PID lookup, the gateway state file, the streaming check and the recent-errors read.probe-failed: ssh <user>@<host> <kind> (retried at 25s: also <kind>) ... gateway status UNDETERMINED, recorded underprobe-failed:<name>:<kind>. "Undetermined" is a different claim from "down", and the previousunreachable:<name>/gateway-down:<name>collapse is what made a transient SSH blip look like a fleet outage.ssh <user>@<host>, plus the CT), so a reader can re-run the exact command.Verification performed by firstmate before this PR
All checks passed, and koby's report-only legs render as🔍 report-only (koby): ... reported, not counted/repairedrather than as failures.Reviewers: (1) confirm no path can print an agent as unreachable/down without the probe target and without the retry having run; (2) confirm a genuine gateway outage still reports clearly - walk the "process really is absent" case and say whether the output distinguishes it from an unreachable probe; (3) confirm koby's report-only handling is unchanged in substance (detected, reported, never counted, never repaired); (4) check whether anything consumes the failure KEYS that changed (
unreachable:<name>->probe-failed:<name>:<kind>), since a rename can silently break an alert or digest consumer; (5) PASS/FAIL with the PR number. Do NOT merge.Per defect report 1154.msg (agent-health-gateway-leg-flap-20260913): 1. Add _ssh_retry() helper with one retry at longer timeout (25s) 2. Name probe target explicitly: ssh {user}@{host} <command> 3. Print probe-failed when first attempt fails, then retry 4. Only declare gateway-down after retry fails 5. Make Koby report-only explicit in output Changes: - check_agents(): all gateway probes now use _ssh_retry() - All output lines name the probe target (ssh host:port) - Koby's report-only status is explicit in output - Never print bare "gateway down" — always name target and failure kind Verified: koonimo shows "probe-failed" on first attempt (transient SSH), retries at 25s, succeeds, reports ✅ koonimo: gw=running