no-mistakes(review): delete duplicated config tables, point to CT 116 authority

This commit is contained in:
root
2026-09-12 15:39:15 +00:00
parent d9eb18c024
commit 2f65c38213
3 changed files with 53 additions and 90 deletions
+18 -47
View File
@@ -80,56 +80,29 @@ triggers:
Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names.
When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched.
| Alias | GPU | Current Model | Will Route To |
|-------|-----|---------------|---------------|
| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo |
| `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 |
| `gpu-vision` | RTX 5070 (.110) | gpu-vision | Whatever runs on RTX 5070 |
| Alias | Serves | Where | Kind |
|-------|--------|-------|------|
| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias |
| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member |
| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias |
| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router |
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there.
**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) still work
but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b`
is retired and returns 400 `Invalid model name`.
## Current Model Assignments (2026-07-15)
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|-------|-----|------|------|-----|----------|----------|-------------|--------|
## Routing Configuration (LiteLLM — July 2026)
### syslog-auto Weighted Pool (Direct GPU — bypasses router)
| Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------|
| `qwen3.6-27B-code` | RTX 3090 (.8:8080) | **0.70** | 500 | **300s** |
| `strix-moe` | Strix Halo (.15:8080) | **0.20** | 60 | **300s** |
| `gpu-vision` | RTX 5070 (.110:8080) | **0.10** | 200 | **300s** |
Single source of truth for models, aliases, rpm caps, weights and fallback chains:
CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in
contracts — read them there.
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
### Direct Model Endpoints
| Model | RPM Cap | Notes |
|-------|---------|-------|
### Stable Aliases (for agent configs — never change)
| Alias | RPM Cap | Routes To | Purpose |
|-------|---------|-----------|---------|
| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) |
| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning |
| `gpu-vision` | 500 | RTX 5070 | Vision, web extract, light tasks |
### Fallback Chains (live `router_settings.fallbacks`, request_timeout 300s throughout)
- syslog-auto → qwen3.6-27B-code → strix-moe → gpu-vision
- qwen3.6-27B-code → strix-moe
- strix-moe → qwen3.6-27B-code → gpu-vision
### Why Strix Halo RPM Is Capped
- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks
- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously
- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C
## Operations
### add-model
@@ -179,9 +152,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act
2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
- Qwen3.5-9B: 120s, qwen3.6-27B-code: 300s, Carnice-Qwen3.6-MoE-35B-A3B/strix-moe: 300s (strix-moe alias retained, legacy name qwen3.6-35B-udq4 deprecated)
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` — read the live values from the authority config; do not assert them from this contract.
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
7. Check port conflicts: verify only one llama-server on :8080 per host
8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`)
@@ -268,9 +239,9 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window.
### Stable Aliases — CRITICAL
All agent configs MUST use stable role-based aliases, never model-specific names:
- `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`)
- `auxiliary.vision.model: gpu-vision` (NOT `gemma-4-12b`)
- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`)
- `compression.model: strix-moe`
- `auxiliary.vision.model: gpu-vision`
- `delegation.model: gpu-dense`
- `auxiliary.web_extract.model: gpu-vision`
When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched.
@@ -288,7 +259,7 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the
| Setting | Value | Notes |
|---------|-------|-------|
| `model.default` | `syslog-auto` | Weighted pool (70% qwen, 20% strix, 10% gpu-vision) |
| `model.default` | `syslog-auto` | Balanced default (pool router) |
| `model.provider` | `custom:litellm` | LiteLLM on CT116 |
| `compression.model` | `strix-moe` | Stable alias — survives model swaps |
| `aux.compression.model` | `strix-moe` | Compression auxiliary model |