Files
prose-contracts/litellm-client-timeouts.prose.md
T
mumuni-bot 7bbf148778
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 10s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s
feat: litellm-client-timeouts contract — standard client timeout/retry policy
Root cause of the 2026-09-06 incident: 157 client-abandoned 408s against a
backend that was succeeding at 20-70s/call once clients stopped giving up.
Measured grounding: syslog-auto 28.8s avg / 25.4s TTFT over 888 calls;
nginx already allows 600s (Rule 5); the gap was entirely client-side.

Values: >=300s primary path, >=120s gpu-dense delegation, 2 retries with
15s/45s backoff, identified probes at 30s/hourly, batch jobs chunked +
day-scheduled. Verified live: Prometheus metrics, 0.56s syslog-auto probe,
gpu-fleet /health/unified topology.
2026-09-08 06:07:43 +00:00

118 lines
5.9 KiB
Markdown

---
kind: pattern
name: litellm-client-timeouts
description: >
Standard client timeout and retry policy for ALL agents calling LiteLLM
(CT 116, http://192.168.68.116). Created 2026-09-08 after the Sep 6 incident:
a backend stall 04:00-06:30 EDT produced 157 client-abandoned 408 failures
(82% from abiba-pi, 13 from mumuni) against a backend that was actually
succeeding at 20-70s per call once clients stopped giving up. Grounded in
measured data: syslog-auto (LiteLLM virtual model, no dedicated GPU — routes
to backends) averages 28.8s/request with 25.4s TTFT over 888 calls/24h;
nginx already allows 600s (Rule 5, verified 2026-08-09); the gap is entirely
client-side. Blast radius if wrong: agents fall back to DeepSeek silently
(key/timeout failures present as model degradation, not errors) or abandon
healthy-but-slow reasoning calls, fragmenting long tasks.
---
## Maintains
- client_timeout_standard: "litellm-client-timeouts v1.0 (2026-09-08)"
- applies_to: ALL agents and scripts calling http://192.168.68.116 (main model path, auxiliary tasks, health probes, benchmark jobs)
- verified_against: Prometheus litellm_* metrics 24h window ending 2026-09-08 ~10:00 EDT; live probe (syslog-auto tiny call 0.56s TTFB, 200 OK); nginx 600s proxy_read_timeout (Rule 5)
## The measured numbers these values come from
| Model | avg latency | avg TTFT | p-profile (24h) |
|---|---|---|---|
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
| strix-moe | 7.5s | — | Strix Halo, healthy |
| gemma-4-12b | 2.6s | — | RTX 5070, healthy |
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
backend), full recovery 07:00-08:00 with ZERO client failures once requests
tolerated 20-70s — 392 successful slow calls in the three hours after recovery.
LiteLLM's internal queue time is ~0s; the latency is model inference, not
proxy queuing.
## Parameters
### 1. Primary model path (model.default / custom_providers) — NO client timeout below 300s
- The default Hermes HTTP timeout (~60s) is TOO SHORT for syslog-auto's healthy
28.8s average + 120-300s tail. Every 408 in the incident was a client
abandoning a request the backend would have answered.
- If the transport exposes a timeout setting for the main model, set it to
**300s or more**. If it does not (current Hermes custom-provider path has no
timeout knob), that is acceptable ONLY because nginx holds the request for
600s — but any wrapper, script, or direct API call you write MUST set its own
timeout >= 300s for syslog-auto/qwen-class calls.
- Never hardcode a shorter timeout "to fail fast" on this path — failing fast
here is what caused the incident.
### 2. Auxiliary tasks — keep template timeouts, one correction
- vision: 60s (keep), web_extract: 30s (keep) — gemma-4-12b averages 2.6s;
these are fine.
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
syslog-auto; delegation defaults that assume fast responses will 408 the
same way.
### 3. Retry policy — backoff, not repetition
- On timeout (408) or 5xx: retry up to **2 times** with exponential backoff
(**15s, 45s**) before giving up.
- Do NOT retry in a tight loop. The Sep 6 spike shape (65 failures in one hour
from one key) was a batch job retrying without backoff while the backend was
down — it multiplied load during recovery.
- On 401/403: do NOT retry — that is a key/permission problem (see
litellm-api-keys.prose.md and Rule 11); retrying just spams the log.
- On 429: honor the retry-after header if present, else back off 60s.
### 4. Health probes — identify yourself and time out sanely
- Probes MUST NOT appear as keyless, model-less failures in the metrics (12
such orphans appeared in the incident window and cost investigation time).
Send a real model name and use a real (probe-designated) key.
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert —
report "backend slow (>30s)" rather than hanging.
- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the
standard; sub-hourly synthetic traffic distorts latency baselines.
### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
- The incident window showed bulk clients amplifying a backend stall 5:1.
- Batch jobs that can tolerate delay: schedule 09:00-17:00 EDT.
- Any batch loop over N requests MUST sleep >= 5s between requests and honor
the retry policy in section 3.
## Returns
- A single standard any agent or script can cite: timeouts >= 300s on the
syslog-auto path, >= 120s on gpu-dense delegation, backoff retries (2x,
15s/45s), identified probes at 30s/hourly, batch jobs chunked and
day-scheduled.
- Failure signature recognition: bulk 408s from multiple keys in one window =
backend event (check gpu_utilization_percent: 0% = wedged, ~100% = saturated);
single-key 408s = that client's timeout is too short.
- Cross-references: hermes-config-template.prose.md (Rule 5 nginx 600s,
auxiliary timeouts), gpu-fleet.prose.md (compression 300s precedent,
stable aliases), litellm-api-keys.prose.md (key/permission failures).
## Intentionally NOT changed
- No server-side LiteLLM timeout/cooldown changes proposed — the incident
self-recovered and the server is healthy (0.56s live probe); changing
server behavior without process-level root cause (CT116 requires root;
not reachable from kagentz) would be guessing.
- No change to the template's vision/web_extract/compression timeouts —
measured data says they are correct.
- No per-agent key permission changes — those are litellm-api-keys.prose.md
territory (and the open gpu-vision/gemma 403 items are already filed with
the key owners).
- No model routing changes — syslog-auto's weighted pool behaved correctly
throughout the incident.