gpu-fleet: July 8 optimization sweep — parallel 2 fleet-wide, 128K ctx on NVIDIA, LiteLLM timeout fixes
- All GPUs: --parallel 1 → 2 (6 concurrent slots, was 3) - .8 RTX 3090: ctx 256K→128K, VRAM 96%→83%, turbo4 KV cache - .110 RTX 5070: ctx 256K→128K, ubatch 4096→512 (was inverted), VRAM 90%→77%, q4_0 KV - .15 Strix Halo: parallel 1→2, 256K ctx (41GB free), q8_0 KV, AMD metrics via /sys/class/drm - LiteLLM: gemma timeout 25→120s, qwen timeout 40→90s, syslog-auto (qwen) 40→90s - Agent configs: context_length 262144 for syslog-auto, 131072 for direct qwen/gemma - Updated health-check operation, agent config implications, benchmark table (fixed model↔GPU mapping)
This commit is contained in:
@@ -73,6 +73,25 @@ description: >
|
||||
- **Fix**: If edit fails, send the response as a new message instead
|
||||
- **Applies to**: pi extension (fixed), Hermes adapter (verify edit fallback exists)
|
||||
|
||||
### 11. Model-ID Mismatch Causes Silent Worker Failure (pi-specific)
|
||||
- **Symptom**: User sends DM, sees placeholder, but never gets response. Worker stays
|
||||
"busy" forever, accumulating pending replies (10+). Health endpoint shows green but
|
||||
`messages_processed` stalls.
|
||||
- **Root cause**: Agent's model config (`models.json` or `settings.json`) references a
|
||||
model ID that doesn't exist in LiteLLM's authorized model list. Example: `qwen3.6-35B-A3B`
|
||||
configured but LiteLLM only exposes `ornith-1.0-35b` under that key. pi's session
|
||||
workers emit 403 on first prompt, then never recover because the error doesn't trigger
|
||||
`agent_end` — worker stays `busy` and all subsequent messages pile up in the steer queue.
|
||||
- **Detection**: Compare `~/.pi/agent/models.json` model IDs against `curl -H "Authorization: Bearer <KEY>" http://192.168.68.116/v1/models` output. A stuck worker shows
|
||||
`workers=[<id>:busy:N]` with growing N in extension logs.
|
||||
- **Fix**: (1) Update `models.json` to only include models from the authorized list.
|
||||
(2) Set `defaultModel` to `syslog-auto` (safe routing model). (3) Delete stale session
|
||||
JSONL files from `~/.pi/agent/sessions/zulip/`. (4) Restart PM2 process.
|
||||
- **Prevention**: Use `syslog-auto` as default model for all agents — it handles model
|
||||
routing and fallback automatically. Direct model IDs (`ornith-1.0-35b`, etc.) should
|
||||
only be used when explicitly requested. Validate model IDs at agent setup time.
|
||||
- **Applies to**: pi extension (Tdunna CT111, fixed 2026-07-08), any agent using `syslog-harness` provider
|
||||
|
||||
## Deployment Checklist
|
||||
|
||||
When deploying a new Zulip adapter, verify:
|
||||
@@ -86,3 +105,4 @@ When deploying a new Zulip adapter, verify:
|
||||
- [ ] `@all-bots` detected via configurable user_id
|
||||
- [ ] Poll uses long-poll (not `dont_block=true` polling)
|
||||
- [ ] Stuck/idle detection accounts for quiet periods
|
||||
- [ ] Model IDs in config validated against `GET /v1/models` with actual API key
|
||||
|
||||
Reference in New Issue
Block a user