Abiba
f2f8e8c921
feat: add request queuing to router (replaces hard 503 on saturation)
...
When all GPUs are saturated, requests now enter a queue loop (poll every 500ms)
instead of immediately returning 503. Configurable via QUEUE_TIMEOUT env var
(default 30s) or X-Queue-Timeout header per-request.
This prevents agent failures from cluster saturation — agents wait for a slot
instead of crashing on fallback.
2026-05-19 15:55:05 +00:00
Abiba
9c31b5d622
May 19, 2026: Full harness update
...
- Model migration: gemma-4-E4B → qwen3.5-9b-vlm
- Dashboard reorder: Usage Over Time + GPU Metrics to top
- Router counter leak fix (gpu_decr in except handler)
- VLM slot upgrade 1→2
- Redis stale key cleanup
- Automated maintenance cron job
- LiteLLM config update
- GPU router config update
- README update
2026-05-19 15:03:34 +00:00
Abiba (pi)
4f032b035c
Mumuni review action items: health checks for all containers, version pinning, 503+Retry-After on all-GPU saturation
2026-05-17 09:05:27 +00:00
Abiba (pi)
8f3b0c6647
Router: health check verifies actual llama.cpp endpoint, gpu_decr negative guard, AMD sidecar fixed (sysfs fallback)
2026-05-17 01:52:28 +00:00
Abiba (pi)
808c9d3d13
Router: 300s timeout, gpu_decr bugfix. Dashboard: Bootstrap 5 modern redesign with KPI stats, equal-height cards, queue ring. Nginx: 600s timeout.
2026-05-16 22:12:21 +00:00
Abiba (pi)
654cdff718
Dashboard: GPU slot indicators show active/max concurrent requests. Koonimo API key added. Real-time queuing visibility.
2026-05-16 20:43:22 +00:00
Abiba (pi)
bf90e57c5f
Load-aware routing: tracks active GPU requests in Redis, distributes overflow when MoE saturated. 6 concurrent requests now spread across all 3 GPUs instead of queuing on one.
2026-05-16 20:23:32 +00:00
Abiba (pi)
ec0f9fac63
Fix: clean_unicode now uses chr()-based replacements + ASCII strip to prevent bash heredoc corruption. Emoji and all non-ASCII now fully stripped.
2026-05-16 19:12:58 +00:00
Abiba (pi)
7b6c6aabe1
Initial commit: CT 116 inference harness — nginx, LiteLLM, router, dashboard, Redis
...
- Complexity-based routing (MoE default, Dense heavy, Gemma light)
- Per-agent API keys with metrics tracking
- Time-series usage graphs (24h/7d/30d)
- Streaming support (SSE passthrough)
- Unicode cleanup (ASCII-only output)
- Vision support (gemma-4-E4B)
- Tier enforcement (starter/professional/enterprise)
- GPU health monitoring via sidecar polling
- Unified dashboard with line graph
2026-05-16 18:51:50 +00:00