Files
prose-contracts/litellm-client-timeouts.prose.md
T
root 8210fd905c
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
fix: align contracts to 4-name LiteLLM registry (2026-09-12)
- Remove retired names (qwen3.6-27B-code, qwen3.6-35B-udq4) from live alias claims
- Update Strix Halo model to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe, 256K ctx)
- Fix litellm-health step 7 probe to gpu-vision (monitor key scoped)
- Move qwen3.6-27B-code/35B-udq4 from raw-but-live to non-resolving in audit
- Fold in pm2-self-heal: remove spoton-service (live PM2 set is 4/4)
- Update hermes templates, key enforcement, timeout tables to live names
2026-09-12 21:39:54 +00:00

6.0 KiB

kind, name, description
kind name description
pattern litellm-client-timeouts Standard client timeout and retry policy for ALL agents calling LiteLLM (CT 116, http://192.168.68.116). Created 2026-09-08 after the Sep 6 incident: a backend stall 04:00-06:30 EDT produced 157 client-abandoned 408 failures (82% from abiba-pi, 13 from mumuni) against a backend that was actually succeeding at 20-70s per call once clients stopped giving up. Grounded in measured data: syslog-auto (LiteLLM virtual model, no dedicated GPU — routes to backends) averages 28.8s/request with 25.4s TTFT over 888 calls/24h; nginx already allows 600s (Rule 5, verified 2026-08-09); the gap is entirely client-side. Blast radius if wrong: agents fall back to DeepSeek silently (key/timeout failures present as model degradation, not errors) or abandon healthy-but-slow reasoning calls, fragmenting long tasks.

Maintains

  • client_timeout_standard: "litellm-client-timeouts v1.0 (2026-09-08)"
  • applies_to: ALL agents and scripts calling http://192.168.68.116 (main model path, auxiliary tasks, health probes, benchmark jobs)
  • verified_against: Prometheus litellm_* metrics 24h window ending 2026-09-08 ~10:00 EDT; live probe (syslog-auto tiny call 0.56s TTFB, 200 OK); nginx 600s proxy_read_timeout (Rule 5)

The measured numbers these values come from

Model avg latency avg TTFT p-profile (24h)
syslog-auto 28.8s 25.4s 68 calls took 30-120s; tail to ~300s under load
strix-moe 7.5s — Strix Halo, healthy
gpu-vision (retired gemma-4-12b, RTX 5070) 2.6s — RTX 5070, healthy

Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged backend), full recovery 07:00-08:00 with ZERO client failures once requests tolerated 20-70s — 392 successful slow calls in the three hours after recovery. LiteLLM's internal queue time is ~0s; the latency is model inference, not proxy queuing.

Parameters

1. Primary model path (model.default / custom_providers) — NO client timeout below 300s

  • The default Hermes HTTP timeout (~60s) is TOO SHORT for syslog-auto's healthy 28.8s average + 120-300s tail. Every 408 in the incident was a client abandoning a request the backend would have answered.
  • If the transport exposes a timeout setting for the main model, set it to 300s or more. If it does not (current Hermes custom-provider path has no timeout knob), that is acceptable ONLY because nginx holds the request for 600s — but any wrapper, script, or direct API call you write MUST set its own timeout >= 300s for syslog-auto/qwen-class calls.
  • Never hardcode a shorter timeout "to fail fast" on this path — failing fast here is what caused the incident.

2. Auxiliary tasks — keep template timeouts, one correction

  • vision: 60s (keep), web_extract: 30s (keep) — the 2.6s average was measured on gemma-4-12b (retired 2026-09-12); the live RTX 5070 alias is gpu-vision.
  • compression: 300s (keep — this was already raised from 60 per gpu-fleet).
  • gpu-dense delegation/x_search: set timeout >= 120s. The RTX 3090 (gpu-dense backend) is the same speed class as syslog-auto; delegation defaults that assume fast responses will 408 the same way.

3. Retry policy — backoff, not repetition

  • On timeout (408) or 5xx: retry up to 2 times with exponential backoff (15s, 45s) before giving up.
  • Do NOT retry in a tight loop. The Sep 6 spike shape (65 failures in one hour from one key) was a batch job retrying without backoff while the backend was down — it multiplied load during recovery.
  • On 401/403: do NOT retry — that is a key/permission problem (see litellm-api-keys.prose.md and Rule 11); retrying just spams the log.
  • On 429: honor the retry-after header if present, else back off 60s.

4. Health probes — identify yourself and time out sanely

  • Probes MUST NOT appear as keyless, model-less failures in the metrics (12 such orphans appeared in the incident window and cost investigation time). Send a real model name and use a real (probe-designated) key.
  • Probe timeout: 30s. A probe that takes longer than 30s IS the alert — report "backend slow (>30s)" rather than hanging.
  • Probe cadence: at most hourly. The litellm-health cron cadence (authoritative trigger in contract-registry.yaml) is the standard; sub-hourly synthetic traffic distorts latency baselines.

5. Batch/benchmark jobs — schedule away from 04:00-07:00 EDT and chunk

  • The incident window showed bulk clients amplifying a backend stall 5:1.
  • Batch jobs that can tolerate delay: schedule 09:00-17:00 EDT.
  • Any batch loop over N requests MUST sleep >= 5s between requests and honor the retry policy in section 3.

Returns

  • A single standard any agent or script can cite: timeouts >= 300s on the syslog-auto path, >= 120s on gpu-dense delegation, backoff retries (2x, 15s/45s), identified probes at 30s/hourly, batch jobs chunked and day-scheduled.
  • Failure signature recognition: bulk 408s from multiple keys in one window = backend event (check gpu_utilization_percent: 0% = wedged, ~100% = saturated); single-key 408s = that client's timeout is too short.
  • Cross-references: hermes-config-template.prose.md (Rule 5 nginx 600s, auxiliary timeouts), gpu-fleet.prose.md (compression 300s precedent, stable aliases), litellm-api-keys.prose.md (key/permission failures).

Intentionally NOT changed

  • No server-side LiteLLM timeout/cooldown changes proposed — the incident self-recovered and the server is healthy (0.56s live probe); changing server behavior without process-level root cause (CT116 requires root; not reachable from kagentz) would be guessing.
  • No change to the template's vision/web_extract/compression timeouts — measured data says they are correct.
  • No per-agent key permission changes — those are litellm-api-keys.prose.md territory (and the open gpu-vision/gemma 403 items are already filed with the key owners).
  • No model routing changes — syslog-auto's weighted pool behaved correctly throughout the incident.