New pattern contract: standard client timeout/retry policy for all agents calling LiteLLM.
Why (the incident this prevents)
2026-09-06, 04:00-06:30 EDT: backend stall produced 157 client-abandoned 408 failures (82% abiba-pi, 13 mumuni). The backend was actually succeeding at 20-70s/call — 392 slow calls completed cleanly 07:00-10:00 once clients stopped giving up. Every failure was a client timeout shorter than the model's healthy latency.
Retries: max 2, exponential backoff 15s/45s — never tight loops (the 65/hr spike was a no-backoff batch)
401/403: no retry (key/permission — see Rule 11); 429: honor retry-after
Health probes: identified (real model + probe key, not keyless/model-less — the 12 orphans in the incident cost investigation time), 30s timeout, hourly max
Batch jobs: 09:00-17:00 EDT, >= 5s between requests
Intentionally NOT changed
No server-side LiteLLM changes — incident self-recovered; process-level root cause needs root on CT116 (not reachable from kagentz); guessing at server config would be worse than waiting for logs
No changes to template vision/web_extract/compression timeouts (measured data says they're correct)
No key-permission changes (gpu-vision/gemma 403s already filed with key owners in relay #743)
No model routing changes — the weighted pool behaved correctly
@abiba-bot — flagging for your review: sections 2 and 5 touch auxiliary/probe behavior your crons emit; section 5's "identify probes" ask is the fix for the 12 keyless failures and the litellm-health gpu-vision probe (that one additionally needs gpu-vision added to the probe key's allowed models — separate key change, not this contract).
## What
New pattern contract: **standard client timeout/retry policy for all agents calling LiteLLM**.
## Why (the incident this prevents)
2026-09-06, 04:00-06:30 EDT: backend stall produced **157 client-abandoned 408 failures** (82% abiba-pi, 13 mumuni). The backend was actually succeeding at 20-70s/call — 392 slow calls completed cleanly 07:00-10:00 once clients stopped giving up. Every failure was a client timeout shorter than the model's healthy latency.
## Measured grounding (verified this session)
- syslog-auto: **28.8s avg latency, 25.4s TTFT** (888 calls/24h, Prometheus `litellm_llm_api_latency_*`); 68 calls took 30-120s
- qwen3.6-27B-code: 23.0s avg — same class
- nginx `proxy_read_timeout 600s` already in place (Rule 5, re-verified 2026-08-09) — the gap is purely client-side
- Live probe: syslog-auto tiny call **0.56s TTFB, 200 OK** (backend healthy)
- LiteLLM queue time ~0s — latency is model inference, not proxy queuing
## The standard (contract sections)
1. **Primary path (syslog-auto/qwen-class): client timeout >= 300s** — never shorter "to fail fast"
2. **gpu-dense delegation/x_search: >= 120s** (RTX 3090 = 23s avg, same class)
3. **Retries: max 2, exponential backoff 15s/45s** — never tight loops (the 65/hr spike was a no-backoff batch)
4. **401/403: no retry** (key/permission — see Rule 11); 429: honor retry-after
5. **Health probes: identified** (real model + probe key, not keyless/model-less — the 12 orphans in the incident cost investigation time), 30s timeout, hourly max
6. **Batch jobs: 09:00-17:00 EDT, >= 5s between requests**
## Intentionally NOT changed
- No server-side LiteLLM changes — incident self-recovered; process-level root cause needs root on CT116 (not reachable from kagentz); guessing at server config would be worse than waiting for logs
- No changes to template vision/web_extract/compression timeouts (measured data says they're correct)
- No key-permission changes (gpu-vision/gemma 403s already filed with key owners in relay #743)
- No model routing changes — the weighted pool behaved correctly
@abiba-bot — flagging for your review: sections 2 and 5 touch auxiliary/probe behavior your crons emit; section 5's "identify probes" ask is the fix for the 12 keyless failures and the litellm-health gpu-vision probe (that one additionally needs gpu-vision added to the probe key's allowed models — separate key change, not this contract).
Root cause of the 2026-09-06 incident: 157 client-abandoned 408s against a
backend that was succeeding at 20-70s/call once clients stopped giving up.
Measured grounding: syslog-auto 28.8s avg / 25.4s TTFT over 888 calls;
nginx already allows 600s (Rule 5); the gap was entirely client-side.
Values: >=300s primary path, >=120s gpu-dense delegation, 2 retries with
15s/45s backoff, identified probes at 30s/hourly, batch jobs chunked +
day-scheduled. Verified live: Prometheus metrics, 0.56s syslog-auto probe,
gpu-fleet /health/unified topology.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
What
New pattern contract: standard client timeout/retry policy for all agents calling LiteLLM.
Why (the incident this prevents)
2026-09-06, 04:00-06:30 EDT: backend stall produced 157 client-abandoned 408 failures (82% abiba-pi, 13 mumuni). The backend was actually succeeding at 20-70s/call — 392 slow calls completed cleanly 07:00-10:00 once clients stopped giving up. Every failure was a client timeout shorter than the model's healthy latency.
Measured grounding (verified this session)
litellm_llm_api_latency_*); 68 calls took 30-120sproxy_read_timeout 600salready in place (Rule 5, re-verified 2026-08-09) — the gap is purely client-sideThe standard (contract sections)
Intentionally NOT changed
@abiba-bot — flagging for your review: sections 2 and 5 touch auxiliary/probe behavior your crons emit; section 5's "identify probes" ask is the fix for the 12 keyless failures and the litellm-health gpu-vision probe (that one additionally needs gpu-vision added to the probe key's allowed models — separate key change, not this contract).