From 7bbf148778229f891cfa1627019e3d5480fce857 Mon Sep 17 00:00:00 2001 From: mumuni-bot Date: Tue, 8 Sep 2026 06:07:43 +0000 Subject: [PATCH] =?UTF-8?q?feat:=20litellm-client-timeouts=20contract=20?= =?UTF-8?q?=E2=80=94=20standard=20client=20timeout/retry=20policy?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Root cause of the 2026-09-06 incident: 157 client-abandoned 408s against a backend that was succeeding at 20-70s/call once clients stopped giving up. Measured grounding: syslog-auto 28.8s avg / 25.4s TTFT over 888 calls; nginx already allows 600s (Rule 5); the gap was entirely client-side. Values: >=300s primary path, >=120s gpu-dense delegation, 2 retries with 15s/45s backoff, identified probes at 30s/hourly, batch jobs chunked + day-scheduled. Verified live: Prometheus metrics, 0.56s syslog-auto probe, gpu-fleet /health/unified topology. --- litellm-client-timeouts.prose.md | 117 +++++++++++++++++++++++++++++++ 1 file changed, 117 insertions(+) create mode 100644 litellm-client-timeouts.prose.md diff --git a/litellm-client-timeouts.prose.md b/litellm-client-timeouts.prose.md new file mode 100644 index 0000000..106fe37 --- /dev/null +++ b/litellm-client-timeouts.prose.md @@ -0,0 +1,117 @@ +--- +kind: pattern +name: litellm-client-timeouts +description: > + Standard client timeout and retry policy for ALL agents calling LiteLLM + (CT 116, http://192.168.68.116). Created 2026-09-08 after the Sep 6 incident: + a backend stall 04:00-06:30 EDT produced 157 client-abandoned 408 failures + (82% from abiba-pi, 13 from mumuni) against a backend that was actually + succeeding at 20-70s per call once clients stopped giving up. Grounded in + measured data: syslog-auto (LiteLLM virtual model, no dedicated GPU — routes + to backends) averages 28.8s/request with 25.4s TTFT over 888 calls/24h; + nginx already allows 600s (Rule 5, verified 2026-08-09); the gap is entirely + client-side. Blast radius if wrong: agents fall back to DeepSeek silently + (key/timeout failures present as model degradation, not errors) or abandon + healthy-but-slow reasoning calls, fragmenting long tasks. +--- + +## Maintains + +- client_timeout_standard: "litellm-client-timeouts v1.0 (2026-09-08)" +- applies_to: ALL agents and scripts calling http://192.168.68.116 (main model path, auxiliary tasks, health probes, benchmark jobs) +- verified_against: Prometheus litellm_* metrics 24h window ending 2026-09-08 ~10:00 EDT; live probe (syslog-auto tiny call 0.56s TTFB, 200 OK); nginx 600s proxy_read_timeout (Rule 5) + +## The measured numbers these values come from + +| Model | avg latency | avg TTFT | p-profile (24h) | +|---|---|---|---| +| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load | +| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto | +| strix-moe | 7.5s | — | Strix Halo, healthy | +| gemma-4-12b | 2.6s | — | RTX 5070, healthy | + +Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged +backend), full recovery 07:00-08:00 with ZERO client failures once requests +tolerated 20-70s — 392 successful slow calls in the three hours after recovery. +LiteLLM's internal queue time is ~0s; the latency is model inference, not +proxy queuing. + +## Parameters + +### 1. Primary model path (model.default / custom_providers) — NO client timeout below 300s + +- The default Hermes HTTP timeout (~60s) is TOO SHORT for syslog-auto's healthy + 28.8s average + 120-300s tail. Every 408 in the incident was a client + abandoning a request the backend would have answered. +- If the transport exposes a timeout setting for the main model, set it to + **300s or more**. If it does not (current Hermes custom-provider path has no + timeout knob), that is acceptable ONLY because nginx holds the request for + 600s — but any wrapper, script, or direct API call you write MUST set its own + timeout >= 300s for syslog-auto/qwen-class calls. +- Never hardcode a shorter timeout "to fail fast" on this path — failing fast + here is what caused the incident. + +### 2. Auxiliary tasks — keep template timeouts, one correction + +- vision: 60s (keep), web_extract: 30s (keep) — gemma-4-12b averages 2.6s; + these are fine. +- compression: 300s (keep — this was already raised from 60 per gpu-fleet). +- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090 + (qwen3.6-27B-code backend, 23.0s avg) is the same speed class as + syslog-auto; delegation defaults that assume fast responses will 408 the + same way. + +### 3. Retry policy — backoff, not repetition + +- On timeout (408) or 5xx: retry up to **2 times** with exponential backoff + (**15s, 45s**) before giving up. +- Do NOT retry in a tight loop. The Sep 6 spike shape (65 failures in one hour + from one key) was a batch job retrying without backoff while the backend was + down — it multiplied load during recovery. +- On 401/403: do NOT retry — that is a key/permission problem (see + litellm-api-keys.prose.md and Rule 11); retrying just spams the log. +- On 429: honor the retry-after header if present, else back off 60s. + +### 4. Health probes — identify yourself and time out sanely + +- Probes MUST NOT appear as keyless, model-less failures in the metrics (12 + such orphans appeared in the incident window and cost investigation time). + Send a real model name and use a real (probe-designated) key. +- Probe timeout: 30s. A probe that takes longer than 30s IS the alert — + report "backend slow (>30s)" rather than hanging. +- Probe cadence: at most hourly. The 6h litellm-health cron cadence is the + standard; sub-hourly synthetic traffic distorts latency baselines. + +### 5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk + +- The incident window showed bulk clients amplifying a backend stall 5:1. +- Batch jobs that can tolerate delay: schedule 09:00-17:00 EDT. +- Any batch loop over N requests MUST sleep >= 5s between requests and honor + the retry policy in section 3. + +## Returns + +- A single standard any agent or script can cite: timeouts >= 300s on the + syslog-auto path, >= 120s on gpu-dense delegation, backoff retries (2x, + 15s/45s), identified probes at 30s/hourly, batch jobs chunked and + day-scheduled. +- Failure signature recognition: bulk 408s from multiple keys in one window = + backend event (check gpu_utilization_percent: 0% = wedged, ~100% = saturated); + single-key 408s = that client's timeout is too short. +- Cross-references: hermes-config-template.prose.md (Rule 5 nginx 600s, + auxiliary timeouts), gpu-fleet.prose.md (compression 300s precedent, + stable aliases), litellm-api-keys.prose.md (key/permission failures). + +## Intentionally NOT changed + +- No server-side LiteLLM timeout/cooldown changes proposed — the incident + self-recovered and the server is healthy (0.56s live probe); changing + server behavior without process-level root cause (CT116 requires root; + not reachable from kagentz) would be guessing. +- No change to the template's vision/web_extract/compression timeouts — + measured data says they are correct. +- No per-agent key permission changes — those are litellm-api-keys.prose.md + territory (and the open gpu-vision/gemma 403 items are already filed with + the key owners). +- No model routing changes — syslog-auto's weighted pool behaved correctly + throughout the incident.