diff --git a/README.md b/README.md index 5e48169..19ab52e 100644 --- a/README.md +++ b/README.md @@ -6,10 +6,10 @@ CT 116 Docker stack for routing local GPU models through a unified OpenAI-compat ``` nginx :80 → router :9000 → GPU backends - ├─ qwen3.6-35B-A3B (MoE) @ 192.168.68.15:8080 [2 slots, 262K ctx] - ├─ qwen3.6-27B-code (Dense) @ 192.168.68.8:8080 [2 slots, 262K ctx] - └─ gpu-vision (VLM) @ 192.168.68.110:8080 [2 slots, 262K ctx] - Total: 6 concurrent slots + ├─ strix-moe (Qwen3.6-MoE-35B-A3B) @ 192.168.68.15:8080 [2 slots, 262K ctx] + ├─ gpu-dense (Qwen3.8-27B-U) @ 192.168.68.8:8080 [1 slot, 131K ctx] + └─ gpu-vision (Qwen3.5-9B VLM) @ 192.168.68.110:8080 [1 slot, 131K ctx] + Total: 4 concurrent slots LiteLLM :8081 (fallback) | Dashboard :3000 | Redis :6379 (local) ``` @@ -62,7 +62,7 @@ When all GPUs are saturated, requests enter a polling queue (500ms intervals) in | GPU | Model | VRAM | Slots | Context | Best For | |-----|-------|------|-------| | Strix Halo | qwen3.6-35B-A3B (MoE) | 65GB | 2 | 262K | General quality | -| RTX 3090 | qwen3.6-27B-code (Dense) | 24GB | 2 | 262K | Code, reasoning | +| RTX 3090 | gpu-dense (Qwen3.8-27B-Uncensored) | 24GB | 1 | 131K | Dense, general | | RTX 5070 | gpu-vision (VLM) | 12GB | 2 | 262K | Speed, vision | ## Maintenance diff --git a/dashboard/dashboard.html b/dashboard/dashboard.html index fdee4c7..c4bc16d 100644 --- a/dashboard/dashboard.html +++ b/dashboard/dashboard.html @@ -70,7 +70,7 @@ body{background:var(--bg);color:var(--text);font-family:-apple-system,BlinkMacSy