no-mistakes(review): align sibling contracts on aliases, crew cap, and monitor key

This commit is contained in:
root
2026-09-12 15:30:14 +00:00
parent 3f07b9bbcc
commit d9eb18c024
3 changed files with 13 additions and 8 deletions
+6 -3
View File
@@ -6,9 +6,10 @@ description: >
registration, health checks, LiteLLM sync, agent key management, GPU registration, health checks, LiteLLM sync, agent key management, GPU
saturation watchdog, Prometheus/Grafana monitoring, and self-healing. saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe, UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
gpu-dense, gpu-light. These never change — only the underlying model does. gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster). RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability. UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias). Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
@@ -100,7 +101,9 @@ is retired and returns 400 `Invalid model name`.
| Model | GPU | Weight | RPM Cap | Timeout | | Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------| |-------|-----|--------|---------|---------|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | | `qwen3.6-27B-code` | RTX 3090 (.8:8080) | **0.70** | 500 | **300s** |
| `strix-moe` | Strix Halo (.15:8080) | **0.20** | 60 | **300s** |
| `gpu-vision` | RTX 5070 (.110:8080) | **0.10** | 200 | **300s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path. Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
+3 -1
View File
@@ -75,7 +75,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- SSH key access to backend_host for container checks - SSH key access to backend_host for container checks
- Network access to public_url, auth_host, and gpu_dashboard_url - Network access to public_url, auth_host, and gpu_dashboard_url
- LiteLLM master key for key management endpoints - LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
model inference checks — the master key must never be used for inference
## GPU Fleet Topology ## GPU Fleet Topology
+4 -4
View File
@@ -70,17 +70,17 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
`model_name`s served (live registry, verified 2026-09-12 with the master key): `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. `gemma-4-12b`, `gpu-light`, and `crew-auto` are retired and absent from the registry. `model_name`s served (live registry, verified 2026-09-12 with the master key): `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. `gemma-4-12b`, `gpu-light`, and `crew-auto` are retired and absent from the registry.
### Context Cap Split (2026-08-20) ### Context Cap Split (2026-08-20; crew cap RETIRED)
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads - **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped - **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit - **Crewmates** (ops, tune, verify, auth-keys, build): the 64K cap was retired together with the `crew-auto` alias; NO context cap is currently in force.
Preferred implementation: uncap shared pool, add capped alias for crew-only. > The 64K crew cap was retired with `crew-auto` (2026-09-12). No limit is currently in force; reinstating one would need per-key model limits as a separate, deliberately-scoped change.
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.70, rpm 500, api_base .8) + strix-moe (0.20, rpm 60, api_base .15) + gpu-vision (0.10, rpm 200, api_base .110). - `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.70, rpm 500, api_base .8) + strix-moe (0.20, rpm 60, api_base .15) + gpu-vision (0.10, rpm 200, api_base .110).
- `gpu-dense` is the high-rpm alias (rpm 500) onto qwen3.6-27B-code; `gpu-light` is retired — the RTX 5070 alias is now `gpu-vision`. - `gpu-dense` is the high-rpm alias (rpm 500) onto qwen3.6-27B-code; `gpu-light` is retired — the RTX 5070 alias is now `gpu-vision`.
- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','strix-moe','gpu-dense']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard set only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them. `/v1/models` is key-scoped, so a model list is only meaningful with the key it was read with. - Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed. A representative agent key (read 2026-09-12) returns `deepseek-v4-pro`, `gpu-dense`, `gpu-vision`, `qwen3.6-27B-code`, `qwen3.6-35B-udq4`, `strix-moe`, `syslog-auto` — `gpu-vision` present, `gpu-light` and `gemma-4-12b` absent. Agents should still use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them.
- **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`. - **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`.
## Model Fallback Chains (LiteLLM) ## Model Fallback Chains (LiteLLM)