no-mistakes(review): align sibling contracts on aliases, crew cap, and monitor key

This commit is contained in:
root
2026-09-12 15:30:14 +00:00
parent 3f07b9bbcc
commit d9eb18c024
3 changed files with 13 additions and 8 deletions
+6 -3
View File
@@ -6,9 +6,10 @@ description: >
registration, health checks, LiteLLM sync, agent key management, GPU
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
gpu-dense, gpu-light. These never change — only the underlying model does.
gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
@@ -100,7 +101,9 @@ is retired and returns 400 `Invalid model name`.
| Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| `qwen3.6-27B-code` | RTX 3090 (.8:8080) | **0.70** | 500 | **300s** |
| `strix-moe` | Strix Halo (.15:8080) | **0.20** | 60 | **300s** |
| `gpu-vision` | RTX 5070 (.110:8080) | **0.10** | 200 | **300s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.