Align harness repo with verified live state; retire model-version names from the client surface #2
@@ -6,10 +6,10 @@ CT 116 Docker stack for routing local GPU models through a unified OpenAI-compat
|
||||
|
||||
```
|
||||
nginx :80 → router :9000 → GPU backends
|
||||
├─ qwen3.6-35B-A3B (MoE) @ 192.168.68.15:8080 [2 slots, 262K ctx]
|
||||
├─ strix-moe (Qwen3.6-MoE-35B-A3B) @ 192.168.68.15:8080 [2 slots, 262K ctx]
|
||||
├─ gpu-dense (Qwen3.8-27B-U) @ 192.168.68.8:8080 [1 slot, 131K ctx]
|
||||
└─ gpu-vision (VLM) @ 192.168.68.110:8080 [2 slots, 262K ctx]
|
||||
Total: 6 concurrent slots
|
||||
└─ gpu-vision (Qwen3.5-9B VLM) @ 192.168.68.110:8080 [1 slot, 131K ctx]
|
||||
Total: 4 concurrent slots
|
||||
|
||||
LiteLLM :8081 (fallback) | Dashboard :3000 | Redis :6379 (local)
|
||||
```
|
||||
|
||||
Reference in New Issue
Block a user