1. TIMEOUT KIND FIX: run_command returns (1, '', 'TIMEOUT') when its own
timeout fires. probe_http now checks for this before falling through to
'curl exit <rc>', so a 30s timeout reports 'timeout after 30s' not
'curl exit 1'.
2. BUSY/DEGRADED DETECTION: After both model probes fail, check the
model's host health endpoint (e.g. 192.168.68.8:8080/health for
gpu-dense). If the host answers 200, report 'busy (completion timed
out after retry; host healthy 200)' — do NOT fail the run on that
alone. If the host does not answer, that's a real FAIL.
3. RETRY TIMEOUT INCREASED: Single-host retry timeout raised from 45s to
90s. Worst-case prefill on a single-slot .8 host is ~76s (observed
83K-token prompt at 1078 tok/s), so 90s covers it.
New line shapes:
- Busy: 'gpu-dense: busy (completion timed out after retry; host healthy 200)'
- Real failure: 'probe-failed: gpu-dense timeout after 30s then timeout after 90s (2 attempts)'