scripts/litellm-health-check.py: (1) a wrapper timeout is now reported as "timeout after s" instead of an opaque "curl exit 1"; (2) after a failed retry the script probes the model host's own /health (gpu-dense .8, gpu-vision .110, strix-moe .15) and reports a distinct BUSY/⚠️ verdict when the host answers, which does NOT fail the run; (3) the single-host retry timeout rises to 90s, above a realistic worst-case prefill.
Why
2026-09-19 ~07:15Z litellm-health failed gpu-dense twice with "curl exit 1 then curl exit 1" and the daily digest exited 1. The model was fine: on .8, llama-chat-api was active, its own /health answered 200, and its llama-server log showed ONE prompt being prefilled on the single slot - task 592307 at n_tokens 40,855 / progress 0.49 (~1,078 tok/s), i.e. an ~83K-token prompt needing ~76s of prefill. Our own agents route through syslog-auto (.8 weighted 0.70), so the contention was self-inflicted, and a healthy-but-busy host was being reported as a failed model. Firstmate verified from CT 116 that all three host health endpoints answer in ~1ms.
Firstmate runs the pre-delivery review and the merge.
## What
scripts/litellm-health-check.py: (1) a wrapper timeout is now reported as "timeout after <N>s" instead of an opaque "curl exit 1"; (2) after a failed retry the script probes the model host's own /health (gpu-dense .8, gpu-vision .110, strix-moe .15) and reports a distinct BUSY/⚠️ verdict when the host answers, which does NOT fail the run; (3) the single-host retry timeout rises to 90s, above a realistic worst-case prefill.
## Why
2026-09-19 ~07:15Z litellm-health failed gpu-dense twice with "curl exit 1 then curl exit 1" and the daily digest exited 1. The model was fine: on .8, llama-chat-api was active, its own /health answered 200, and its llama-server log showed ONE prompt being prefilled on the single slot - task 592307 at n_tokens 40,855 / progress 0.49 (~1,078 tok/s), i.e. an ~83K-token prompt needing ~76s of prefill. Our own agents route through syslog-auto (.8 weighted 0.70), so the contention was self-inflicted, and a healthy-but-busy host was being reported as a failed model. Firstmate verified from CT 116 that all three host health endpoints answer in ~1ms.
Firstmate runs the pre-delivery review and the merge.
1. TIMEOUT KIND FIX: run_command returns (1, '', 'TIMEOUT') when its own
timeout fires. probe_http now checks for this before falling through to
'curl exit <rc>', so a 30s timeout reports 'timeout after 30s' not
'curl exit 1'.
2. BUSY/DEGRADED DETECTION: After both model probes fail, check the
model's host health endpoint (e.g. 192.168.68.8:8080/health for
gpu-dense). If the host answers 200, report 'busy (completion timed
out after retry; host healthy 200)' — do NOT fail the run on that
alone. If the host does not answer, that's a real FAIL.
3. RETRY TIMEOUT INCREASED: Single-host retry timeout raised from 45s to
90s. Worst-case prefill on a single-slot .8 host is ~76s (observed
83K-token prompt at 1078 tok/s), so 90s covers it.
New line shapes:
- Busy: 'gpu-dense: busy (completion timed out after retry; host healthy 200)'
- Real failure: 'probe-failed: gpu-dense timeout after 30s then timeout after 90s (2 attempts)'
Three states:
- healthy: passed, exit 0 (unchanged)
- busy (completion timed out after retry AND host /health answered):
⚠️ DEGRADED line, does NOT fail the run, exit 0
- host unreachable or real fault: ❌, exit 1 (unchanged)
Summary now reports degraded count:
- All pass, no degraded: '✅ All checks passed'
- All pass, 1+ degraded: '✅ All checks passed (1 degraded: gpu-dense)'
- Some failed: '❌ Some checks failed' or '❌ Some checks failed (1 degraded: ...)'
Host health mapping verified:
- gpu-dense -> 192.168.68.8:8080/health
- gpu-vision -> 192.168.68.110:8080/health
- strix-moe -> 192.168.68.15:8080/health
The busy line was rendering as:
'busy (completion timed out after retry; host healthy host healthy (200))'
because host_detail already contains 'host healthy (200)' and the prefix
also said 'host healthy'. Fixed to:
'busy (completion timed out after retry; host healthy (200))'
F1 cosmetic fix from PR #123 verify.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
What
scripts/litellm-health-check.py: (1) a wrapper timeout is now reported as "timeout after s" instead of an opaque "curl exit 1"; (2) after a failed retry the script probes the model host's own /health (gpu-dense .8, gpu-vision .110, strix-moe .15) and reports a distinct BUSY/⚠️ verdict when the host answers, which does NOT fail the run; (3) the single-host retry timeout rises to 90s, above a realistic worst-case prefill.
Why
2026-09-19 ~07:15Z litellm-health failed gpu-dense twice with "curl exit 1 then curl exit 1" and the daily digest exited 1. The model was fine: on .8, llama-chat-api was active, its own /health answered 200, and its llama-server log showed ONE prompt being prefilled on the single slot - task 592307 at n_tokens 40,855 / progress 0.49 (~1,078 tok/s), i.e. an ~83K-token prompt needing ~76s of prefill. Our own agents route through syslog-auto (.8 weighted 0.70), so the contention was self-inflicted, and a healthy-but-busy host was being reported as a failed model. Firstmate verified from CT 116 that all three host health endpoints answer in ~1ms.
Firstmate runs the pre-delivery review and the merge.