fix(litellm): Report both attempts' failure kinds in probe-failed
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 7s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 4s
The failure line now preserves both attempts' failure kinds instead of hardcoding 'timeout after retry, 45s'. If both attempts fail, the report shows: 'probe-failed: <model> <first kind> then <retry kind> (2 attempts)'. This fixes the self-contradictory output when the first attempt timed out but the retry failed with connection refused, and prevents the duration from appearing twice when both attempts were timeouts. Example outputs: - timeout then timeout: 'probe-failed: gpu-dense timeout after 30s then timeout after 45s (2 attempts)' - timeout then refused: 'probe-failed: gpu-dense timeout after 30s then connection refused (2 attempts)' - refused then refused: 'probe-failed: gpu-dense connection refused then connection refused (2 attempts)'
This commit is contained in:
@@ -124,8 +124,10 @@ def check_model_probes():
|
||||
data='{"model":"' + model + '","messages":[{"role":"user","content":"health ' + str(random.randint(1000, 9999)) + '"}],"max_tokens":4}',
|
||||
timeout=30)
|
||||
|
||||
first_kind = None # Track first attempt's failure kind
|
||||
if code == 000 and failure_kind:
|
||||
# Retry once with longer timeout (45s) before declaring failure
|
||||
first_kind = failure_kind
|
||||
time.sleep(1)
|
||||
code, failure_kind = probe_http("http://" + BACKEND_HOST + "/litellm/v1/chat/completions",
|
||||
method="POST",
|
||||
@@ -134,8 +136,11 @@ def check_model_probes():
|
||||
timeout=45)
|
||||
|
||||
if code == 000 and failure_kind:
|
||||
# Both attempts failed - report with duration
|
||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind + " (timeout after retry, 45s)"))
|
||||
# Both attempts failed - report both kinds
|
||||
if first_kind:
|
||||
results.append((model, False, "probe-failed: " + model + " " + first_kind + " then " + failure_kind + " (2 attempts)"))
|
||||
else:
|
||||
results.append((model, False, "probe-failed: " + model + " " + failure_kind))
|
||||
elif code == 200:
|
||||
results.append((model, True, str(code) + " (target: " + BACKEND_HOST + "/litellm/v1/chat/completions, model=" + model + ")"))
|
||||
elif code in (401, 403):
|
||||
|
||||
Reference in New Issue
Block a user