no-mistakes(review): align sibling contracts on aliases, crew cap, and monitor key

This commit is contained in:
root
2026-09-12 15:30:14 +00:00
parent 3f07b9bbcc
commit d9eb18c024
3 changed files with 13 additions and 8 deletions
+6 -3
View File
@@ -6,9 +6,10 @@ description: >
registration, health checks, LiteLLM sync, agent key management, GPU
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe,
gpu-dense, gpu-light. These never change — only the underlying model does.
gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12).
These never change — only the underlying model does.
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster).
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
@@ -100,7 +101,9 @@ is retired and returns 400 `Invalid model name`.
| Model | GPU | Weight | RPM Cap | Timeout |
|-------|-----|--------|---------|---------|
| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** |
| `qwen3.6-27B-code` | RTX 3090 (.8:8080) | **0.70** | 500 | **300s** |
| `strix-moe` | Strix Halo (.15:8080) | **0.20** | 60 | **300s** |
| `gpu-vision` | RTX 5070 (.110:8080) | **0.10** | 200 | **300s** |
Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path.
+3 -1
View File
@@ -75,7 +75,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
- SSH key access to backend_host for container checks
- Network access to public_url, auth_host, and gpu_dashboard_url
- LiteLLM master key for key management endpoints
- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`)
- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for
model inference checks — the master key must never be used for inference
## GPU Fleet Topology
+4 -4
View File
@@ -70,17 +70,17 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2)
`model_name`s served (live registry, verified 2026-09-12 with the master key): `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. `gemma-4-12b`, `gpu-light`, and `crew-auto` are retired and absent from the registry.
### Context Cap Split (2026-08-20)
### Context Cap Split (2026-08-20; crew cap RETIRED)
- **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads
- **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped
- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit
- **Crewmates** (ops, tune, verify, auth-keys, build): the 64K cap was retired together with the `crew-auto` alias; NO context cap is currently in force.
Preferred implementation: uncap shared pool, add capped alias for crew-only.
> The 64K crew cap was retired with `crew-auto` (2026-09-12). No limit is currently in force; reinstating one would need per-key model limits as a separate, deliberately-scoped change.
- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.70, rpm 500, api_base .8) + strix-moe (0.20, rpm 60, api_base .15) + gpu-vision (0.10, rpm 200, api_base .110).
- `gpu-dense` is the high-rpm alias (rpm 500) onto qwen3.6-27B-code; `gpu-light` is retired — the RTX 5070 alias is now `gpu-vision`.
- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','strix-moe','gpu-dense']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard set only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them. `/v1/models` is key-scoped, so a model list is only meaningful with the key it was read with.
- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed. A representative agent key (read 2026-09-12) returns `deepseek-v4-pro`, `gpu-dense`, `gpu-vision`, `qwen3.6-27B-code`, `qwen3.6-35B-udq4`, `strix-moe`, `syslog-auto` — `gpu-vision` present, `gpu-light` and `gemma-4-12b` absent. Agents should still use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them.
- **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`.
## Model Fallback Chains (LiteLLM)