Abiba Bot
|
1e24bc3b9b
|
docs: update model table with context windows, capabilities, GPU labels
|
2026-06-05 23:49:03 +00:00 |
|
Abiba Bot
|
e6a6c30211
|
router: v3 — 5-tier routing with vision guard, MoE spillover, code hint
GUARD: multimodal -> VLM only
TIER 1 Tiny: VLM->Dense->MoE
TIER 2 Light: VLM->Dense->MoE
TIER 3 Medium: MoE<->Dense (40% spill)->VLM
TIER 4 Heavy: Dense->MoE->VLM (quality first)
TIER 5 Default: MoE<->Dense (40% spill)->VLM
HINTS: speed->VLM, quality->MoE, code->Dense
+ moe_spillover(): 60/40 MoE/Dense load distribution
|
2026-06-05 23:48:45 +00:00 |
|
Abiba Bot
|
c0be4c1699
|
router: reference - new 5-tier routing with vision guard (deployed CT116)
Vision guard: multimodal -> VLM only
TIER 1 Tiny: VLM->Dense->MoE
TIER 2 Light: Dense->VLM->MoE
TIER 3 Medium: MoE->Dense->VLM
TIER 4 Heavy: MoE->Dense->VLM
TIER 5 Default: MoE->Dense->VLM
Hints: speed->VLM, quality->MoE, code->Dense
|
2026-06-05 23:32:27 +00:00 |
|
Abiba Bot
|
424d943b12
|
router: fix Dense context window 98K→262K (matches actual n_ctx)
|
2026-06-05 23:15:18 +00:00 |
|
Abiba Bot
|
b6abd84062
|
dashboard: rename VLM GPU label 'New Backend' → 'RTX 5070' for consistency
|
2026-06-05 21:53:07 +00:00 |
|
Abiba Bot
|
dace488f93
|
VLM migration: qwen3.5-9b-vlm → gemma-4-12b across entire harness
- router/router.py: model ID, context window (131K→262K), VRAM comment
- litellm_config.yaml: model name updated
- gpu-router.conf, gpu-router-docker.conf: pool names, comments
- dashboard/dashboard.py: MODELS array refactor for single-point config
- README.md: architecture diagram, model table
|
2026-06-05 21:37:18 +00:00 |
|
Abiba Bot
|
370f5546dd
|
dashboard: migrate VLM model qwen3.5-9b-vlm → gemma-4-12b (new backend 192.168.68.110:8080/v1)
|
2026-06-05 21:26:51 +00:00 |
|