fix: align contracts to 4-name LiteLLM registry (2026-09-12)
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 1s
- Remove retired names (qwen3.6-27B-code, qwen3.6-35B-udq4) from live alias claims - Update Strix Halo model to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (strix-moe, 256K ctx) - Fix litellm-health step 7 probe to gpu-vision (monitor key scoped) - Move qwen3.6-27B-code/35B-udq4 from raw-but-live to non-resolving in audit - Fold in pm2-self-heal: remove spoton-service (live PM2 set is 4/4) - Update hermes templates, key enforcement, timeout tables to live names
This commit is contained in:
@@ -26,9 +26,8 @@ description: >
|
||||
| Model | avg latency | avg TTFT | p-profile (24h) |
|
||||
|---|---|---|---|
|
||||
| syslog-auto | 28.8s | 25.4s | 68 calls took 30-120s; tail to ~300s under load |
|
||||
| qwen3.6-27B-code | 23.0s | — | same backend class as syslog-auto |
|
||||
| strix-moe | 7.5s | — | Strix Halo, healthy |
|
||||
| gemma-4-12b (retired 2026-09-12; RTX 5070 now `gpu-vision`) | 2.6s | — | RTX 5070, healthy |
|
||||
| gpu-vision (retired gemma-4-12b, RTX 5070) | 2.6s | — | RTX 5070, healthy |
|
||||
|
||||
Sep 6 incident timeline: failures 04:00-07:00 EDT (0% GPU util = wedged
|
||||
backend), full recovery 07:00-08:00 with ZERO client failures once requests
|
||||
@@ -56,9 +55,8 @@ proxy queuing.
|
||||
- vision: 60s (keep), web_extract: 30s (keep) — the 2.6s average was measured on `gemma-4-12b` (retired 2026-09-12); the live RTX 5070 alias is `gpu-vision`.
|
||||
- compression: 300s (keep — this was already raised from 60 per gpu-fleet).
|
||||
- **gpu-dense delegation/x_search: set timeout >= 120s.** The RTX 3090
|
||||
(qwen3.6-27B-code backend, 23.0s avg) is the same speed class as
|
||||
syslog-auto; delegation defaults that assume fast responses will 408 the
|
||||
same way.
|
||||
(gpu-dense backend) is the same speed class as syslog-auto; delegation
|
||||
defaults that assume fast responses will 408 the same way.
|
||||
|
||||
### 3. Retry policy — backoff, not repetition
|
||||
|
||||
|
||||
Reference in New Issue
Block a user