fix(litellm): Add retry with longer timeout for single-host model probes
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Single-host models (gpu-dense, gpu-vision, strix-moe) now retry once at 45s on initial 30s timeout failure before declaring probe-failed. This prevents a single transient timeout (cold prefill ~13s or concurrent generation hold) from failing the entire health digest. Evidence: 2026-09-19 ~06:55Z digest failed gpu-dense at 30s; 06:56Z direct probe 200 in 1.04s. The failed-probe-fails-the-run property is preserved: if both attempts fail, the script still exits non-zero with the target and duration named. Closes: daily-health-digest false negative on single transient timeout
This commit is contained in:
@@ -116,8 +116,8 @@ def check_model_probes():
|
||||
results = []
|
||||
|
||||
for model in ["gpu-dense", "gpu-vision", "strix-moe"]:
|
||||
# Single-host aliases: 30s timeout each
|
||||
# gpu-dense (RTX 3090) may need long warmup/prefill - timeout is acceptable on cold-start
|
||||
# Single-host aliases: 30s initial timeout, retry once at 45s on failure
|
||||
# gpu-dense (RTX 3090) may need long warmup/prefill or concurrent generation hold
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
@@ -125,8 +125,17 @@ def check_model_probes():
|
||||
timeout=30)
|
||||
|
||||
if code == 000 and failure_kind:
|
||||
# Report probe failure with kind, do not assert a service verdict
|
||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (30s timeout)"))
|
||||
# Retry once with longer timeout (45s) before declaring failure
|
||||
time.sleep(1)
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
bearer_token=monitor_key,
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=45)
|
||||
|
||||
if code == 000 and failure_kind:
|
||||
# Both attempts failed - report with duration
|
||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (timeout after retry, 45s)"))
|
||||
elif code == 200:
|
||||
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
elif code in (401, 403):
|
||||
|
||||
Reference in New Issue
Block a user