scripts/litellm-health-check.py: a single-host model probe (gpu-dense, gpu-vision, strix-moe) now retries once at 45s after an initial 30s timeout, and only reports failure if BOTH attempts fail ("probe-failed: (timeout after retry, 45s)").
Why
2026-09-19 ~06:55Z the daily digest failed with "LiteLLM Health: gpu-dense probe-failed (30s timeout)" and exited 1. A direct probe 30 seconds later returned HTTP 200 in 1.04s (syslog-auto 200 in 0.39s), so the model was healthy. On this fleet's single-slot .8 host a cold prefill costs ~13s and any concurrent generation can hold the slot past 30s, so one transient was failing the whole digest. Failing loudly on two real failures is preserved.
Firstmate runs the pre-delivery review and the merge.
## What
scripts/litellm-health-check.py: a single-host model probe (gpu-dense, gpu-vision, strix-moe) now retries once at 45s after an initial 30s timeout, and only reports failure if BOTH attempts fail ("probe-failed: <model> <kind> (timeout after retry, 45s)").
## Why
2026-09-19 ~06:55Z the daily digest failed with "LiteLLM Health: gpu-dense probe-failed (30s timeout)" and exited 1. A direct probe 30 seconds later returned HTTP 200 in 1.04s (syslog-auto 200 in 0.39s), so the model was healthy. On this fleet's single-slot .8 host a cold prefill costs ~13s and any concurrent generation can hold the slot past 30s, so one transient was failing the whole digest. Failing loudly on two real failures is preserved.
Firstmate runs the pre-delivery review and the merge.
Single-host models (gpu-dense, gpu-vision, strix-moe) now retry once at
45s on initial 30s timeout failure before declaring probe-failed. This
prevents a single transient timeout (cold prefill ~13s or concurrent
generation hold) from failing the entire health digest.
Evidence: 2026-09-19 ~06:55Z digest failed gpu-dense at 30s; 06:56Z
direct probe 200 in 1.04s.
The failed-probe-fails-the-run property is preserved: if both attempts
fail, the script still exits non-zero with the target and duration named.
Closes: daily-health-digest false negative on single transient timeout
The failure line now preserves both attempts' failure kinds instead of
hardcoding 'timeout after retry, 45s'. If both attempts fail, the report
shows: 'probe-failed: <model> <first kind> then <retry kind> (2 attempts)'.
This fixes the self-contradictory output when the first attempt timed out
but the retry failed with connection refused, and prevents the duration
from appearing twice when both attempts were timeouts.
Example outputs:
- timeout then timeout: 'probe-failed: gpu-dense timeout after 30s then timeout after 45s (2 attempts)'
- timeout then refused: 'probe-failed: gpu-dense timeout after 30s then connection refused (2 attempts)'
- refused then refused: 'probe-failed: gpu-dense connection refused then connection refused (2 attempts)'
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
What
scripts/litellm-health-check.py: a single-host model probe (gpu-dense, gpu-vision, strix-moe) now retries once at 45s after an initial 30s timeout, and only reports failure if BOTH attempts fail ("probe-failed: (timeout after retry, 45s)").
Why
2026-09-19 ~06:55Z the daily digest failed with "LiteLLM Health: gpu-dense probe-failed (30s timeout)" and exited 1. A direct probe 30 seconds later returned HTTP 200 in 1.04s (syslog-auto 200 in 0.39s), so the model was healthy. On this fleet's single-slot .8 host a cold prefill costs ~13s and any concurrent generation can hold the slot past 30s, so one transient was failing the whole digest. Failing loudly on two real failures is preserved.
Firstmate runs the pre-delivery review and the merge.