PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
Root cause of the 2026-09-06 incident: 157 client-abandoned 408s against a backend that was succeeding at 20-70s/call once clients stopped giving up. Measured grounding: syslog-auto 28.8s avg / 25.4s TTFT over 888 calls; nginx already allows 600s (Rule 5); the gap was entirely client-side. Values: >=300s primary path, >=120s gpu-dense delegation, 2 retries with 15s/45s backoff, identified probes at 30s/hourly, batch jobs chunked + day-scheduled. Verified live: Prometheus metrics, 0.56s syslog-auto probe, gpu-fleet /health/unified topology.
118 lines
5.9 KiB
Markdown
118 lines
5.9 KiB
Markdown
---
|
|
kind: pattern
|
|
name: litellm-client-timeouts
|
|
description: >
|
|
Standard client timeout and retry policy for ALL agents calling LiteLLM
|
|
(CT 116, http://192.168.68.116). Created 2026-09-08 after the Sep 6 incident:
|
|
a backend stall 04:00-06:30 EDT produced 157 client-abandoned 408 failures
|
|
(82% from abiba-pi, 13 from mumuni) against a backend that was actually
|
|
succeeding at 20-70s per call once clients stopped giving up. Grounded in
|
|
measured data: syslog-auto (LiteLLM virtual model, no dedicated GPU — routes
|
|
to backends) averages 28.8s/request with 25.4s TTFT over 888 calls/24h;
|
|
nginx already allows 600s (Rule 5, verified 2026-08-09); the gap is entirely
|
|
client-side. Blast radius if wrong: agents fall back to DeepSeek silently
|
|
(key/timeout failures present as model degradation, not errors) or abandon
|
|
healthy-but-slow reasoning calls, fragmenting long tasks.
|
|
---
|
|
|
|
## Maintains
|
|
|
|
- client_timeout_standard: "litellm-client-timeouts v1.0 (2026-09-08)"
|
|
- applies_to: ALL agents and scripts calling http://192.168.68.116 (main model path, auxiliary tasks, health probes, benchmark jobs)
|
|
- verified_against: Prometheus litellm_* metrics 24h window ending 2026-09-08 ~10:00 EDT; live probe (syslog-auto tiny call 0.56s TTFB, 200 OK); nginx 600s proxy_read_timeout (Rule 5)
|
|
|
|
## The measured numbers these values come from
|
|
|
|
| Model | avg latency | avg TTFT | p-profile (24h) |
|
|
|---|---|---|---|
|
|
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
|
|
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
|
|
| strix-moe | 7.5s | — | Strix Halo, healthy |
|
|
| gemma-4-12b | 2.6s | — | RTX 5070, healthy |
|
|
|
|
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
|
|
backend), full recovery 07:00-08:00 with ZERO client failures once requests
|
|
tolerated 20-70s — 392 successful slow calls in the three hours after recovery.
|
|
LiteLLM's internal queue time is ~0s; the latency is model inference, not
|
|
proxy queuing.
|
|
|
|
## Parameters
|
|
|
|
### 1. Primary model path (model.default / custom_providers) — NO client timeout below 300s
|
|
|
|
- The default Hermes HTTP timeout (~60s) is TOO SHORT for syslog-auto's healthy
|
|
28.8s average + 120-300s tail. Every 408 in the incident was a client
|
|
abandoning a request the backend would have answered.
|
|
- If the transport exposes a timeout setting for the main model, set it to
|
|
**300s or more**. If it does not (current Hermes custom-provider path has no
|
|
timeout knob), that is acceptable ONLY because nginx holds the request for
|
|
600s — but any wrapper, script, or direct API call you write MUST set its own
|
|
timeout >= 300s for syslog-auto/qwen-class calls.
|
|
- Never hardcode a shorter timeout "to fail fast" on this path — failing fast
|
|
here is what caused the incident.
|
|
|
|
### 2. Auxiliary tasks — keep template timeouts, one correction
|
|
|
|
- vision: 60s (keep), web_extract: 30s (keep) — gemma-4-12b averages 2.6s;
|
|
these are fine.
|
|
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
|
|
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
|
|
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
|
|
syslog-auto; delegation defaults that assume fast responses will 408 the
|
|
same way.
|
|
|
|
### 3. Retry policy — backoff, not repetition
|
|
|
|
- On timeout (408) or 5xx: retry up to **2 times** with exponential backoff
|
|
(**15s, 45s**) before giving up.
|
|
- Do NOT retry in a tight loop. The Sep 6 spike shape (65 failures in one hour
|
|
from one key) was a batch job retrying without backoff while the backend was
|
|
down — it multiplied load during recovery.
|
|
- On 401/403: do NOT retry — that is a key/permission problem (see
|
|
litellm-api-keys.prose.md and Rule 11); retrying just spams the log.
|
|
- On 429: honor the retry-after header if present, else back off 60s.
|
|
|
|
### 4. Health probes — identify yourself and time out sanely
|
|
|
|
- Probes MUST NOT appear as keyless, model-less failures in the metrics (12
|
|
such orphans appeared in the incident window and cost investigation time).
|
|
Send a real model name and use a real (probe-designated) key.
|
|
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert —
|
|
report "backend slow (>30s)" rather than hanging.
|
|
- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the
|
|
standard; sub-hourly synthetic traffic distorts latency baselines.
|
|
|
|
### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
|
|
|
|
- The incident window showed bulk clients amplifying a backend stall 5:1.
|
|
- Batch jobs that can tolerate delay: schedule 09:00-17:00 EDT.
|
|
- Any batch loop over N requests MUST sleep >= 5s between requests and honor
|
|
the retry policy in section 3.
|
|
|
|
## Returns
|
|
|
|
- A single standard any agent or script can cite: timeouts >= 300s on the
|
|
syslog-auto path, >= 120s on gpu-dense delegation, backoff retries (2x,
|
|
15s/45s), identified probes at 30s/hourly, batch jobs chunked and
|
|
day-scheduled.
|
|
- Failure signature recognition: bulk 408s from multiple keys in one window =
|
|
backend event (check gpu_utilization_percent: 0% = wedged, ~100% = saturated);
|
|
single-key 408s = that client's timeout is too short.
|
|
- Cross-references: hermes-config-template.prose.md (Rule 5 nginx 600s,
|
|
auxiliary timeouts), gpu-fleet.prose.md (compression 300s precedent,
|
|
stable aliases), litellm-api-keys.prose.md (key/permission failures).
|
|
|
|
## Intentionally NOT changed
|
|
|
|
- No server-side LiteLLM timeout/cooldown changes proposed — the incident
|
|
self-recovered and the server is healthy (0.56s live probe); changing
|
|
server behavior without process-level root cause (CT116 requires root;
|
|
not reachable from kagentz) would be guessing.
|
|
- No change to the template's vision/web_extract/compression timeouts —
|
|
measured data says they are correct.
|
|
- No per-agent key permission changes — those are litellm-api-keys.prose.md
|
|
territory (and the open gpu-vision/gemma 403 items are already filed with
|
|
the key owners).
|
|
- No model routing changes — syslog-auto's weighted pool behaved correctly
|
|
throughout the incident.
|