6.1 KiB
6.1 KiB
kind, name, description
| kind | name | description |
|---|---|---|
| pattern | litellm-client-timeouts | Standard client timeout and retry policy for ALL agents calling LiteLLM (CT 116, http://192.168.68.116). Created 2026-09-08 after the Sep 6 incident: a backend stall 04:00-06:30 EDT produced 157 client-abandoned 408 failures (82% from abiba-pi, 13 from mumuni) against a backend that was actually succeeding at 20-70s per call once clients stopped giving up. Grounded in measured data: syslog-auto (LiteLLM virtual model, no dedicated GPU — routes to backends) averages 28.8s/request with 25.4s TTFT over 888 calls/24h; nginx already allows 600s (Rule 5, verified 2026-08-09); the gap is entirely client-side. Blast radius if wrong: agents fall back to DeepSeek silently (key/timeout failures present as model degradation, not errors) or abandon healthy-but-slow reasoning calls, fragmenting long tasks. |
Maintains
- client_timeout_standard: "litellm-client-timeouts v1.0 (2026-09-08)"
- applies_to: ALL agents and scripts calling http://192.168.68.116 (main model path, auxiliary tasks, health probes, benchmark jobs)
- verified_against: Prometheus litellm_* metrics 24h window ending 2026-09-08 ~10:00 EDT; live probe (syslog-auto tiny call 0.56s TTFB, 200 OK); nginx 600s proxy_read_timeout (Rule 5)
The measured numbers these values come from
| Model | avg latency | avg TTFT | p-profile (24h) |
|---|---|---|---|
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
| strix-moe | 7.5s | — | Strix Halo, healthy |
gemma-4-12b (retired 2026-09-12; RTX 5070 now gpu-vision) |
2.6s | — | RTX 5070, healthy |
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged backend), full recovery 07:00-08:00 with ZERO client failures once requests tolerated 20-70s — 392 successful slow calls in the three hours after recovery. LiteLLM's internal queue time is ~0s; the latency is model inference, not proxy queuing.
Parameters
1. Primary model path (model.default / custom_providers) — NO client timeout below 300s
- The default Hermes HTTP timeout (~60s) is TOO SHORT for syslog-auto's healthy 28.8s average + 120-300s tail. Every 408 in the incident was a client abandoning a request the backend would have answered.
- If the transport exposes a timeout setting for the main model, set it to 300s or more. If it does not (current Hermes custom-provider path has no timeout knob), that is acceptable ONLY because nginx holds the request for 600s — but any wrapper, script, or direct API call you write MUST set its own timeout >= 300s for syslog-auto/qwen-class calls.
- Never hardcode a shorter timeout "to fail fast" on this path — failing fast here is what caused the incident.
2. Auxiliary tasks — keep template timeouts, one correction
- vision: 60s (keep), web_extract: 30s (keep) — the 2.6s average was measured on
gemma-4-12b(retired 2026-09-12); the live RTX 5070 alias isgpu-vision. - compression: 300s (keep — this was already raised from 60 per gpu-fleet).
- gpu-dense delegation/x_search: set timeout >= 120s. The RTX 3090 (qwen3.6-27B-code backend, 23.0s avg) is the same speed class as syslog-auto; delegation defaults that assume fast responses will 408 the same way.
3. Retry policy — backoff, not repetition
- On timeout (408) or 5xx: retry up to 2 times with exponential backoff (15s, 45s) before giving up.
- Do NOT retry in a tight loop. The Sep 6 spike shape (65 failures in one hour from one key) was a batch job retrying without backoff while the backend was down — it multiplied load during recovery.
- On 401/403: do NOT retry — that is a key/permission problem (see litellm-api-keys.prose.md and Rule 11); retrying just spams the log.
- On 429: honor the retry-after header if present, else back off 60s.
4. Health probes — identify yourself and time out sanely
- Probes MUST NOT appear as keyless, model-less failures in the metrics (12 such orphans appeared in the incident window and cost investigation time). Send a real model name and use a real (probe-designated) key.
- Probe timeout: 30s. A probe that takes longer than 30s IS the alert — report "backend slow (>30s)" rather than hanging.
- Probe cadence: at most hourly. The litellm-health cron cadence (authoritative trigger in contract-registry.yaml) is the standard; sub-hourly synthetic traffic distorts latency baselines.
5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk
- The incident window showed bulk clients amplifying a backend stall 5:1.
- Batch jobs that can tolerate delay: schedule 09:00-17:00 EDT.
- Any batch loop over N requests MUST sleep >= 5s between requests and honor the retry policy in section 3.
Returns
- A single standard any agent or script can cite: timeouts >= 300s on the syslog-auto path, >= 120s on gpu-dense delegation, backoff retries (2x, 15s/45s), identified probes at 30s/hourly, batch jobs chunked and day-scheduled.
- Failure signature recognition: bulk 408s from multiple keys in one window = backend event (check gpu_utilization_percent: 0% = wedged, ~100% = saturated); single-key 408s = that client's timeout is too short.
- Cross-references: hermes-config-template.prose.md (Rule 5 nginx 600s, auxiliary timeouts), gpu-fleet.prose.md (compression 300s precedent, stable aliases), litellm-api-keys.prose.md (key/permission failures).
Intentionally NOT changed
- No server-side LiteLLM timeout/cooldown changes proposed — the incident self-recovered and the server is healthy (0.56s live probe); changing server behavior without process-level root cause (CT116 requires root; not reachable from kagentz) would be guessing.
- No change to the template's vision/web_extract/compression timeouts — measured data says they are correct.
- No per-agent key permission changes — those are litellm-api-keys.prose.md territory (and the open gpu-vision/gemma 403 items are already filed with the key owners).
- No model routing changes — syslog-auto's weighted pool behaved correctly throughout the incident.