diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 6f2b80b..e952fea 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -6,9 +6,10 @@ description: > registration, health checks, LiteLLM sync, agent key management, GPU saturation watchdog, Prometheus/Grafana monitoring, and self-healing. UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe, - gpu-dense, gpu-light. These never change — only the underlying model does. + gpu-dense, gpu-vision (gpu-light was superseded by gpu-vision on 2026-09-12). + These never change — only the underlying model does. Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). - RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster). + RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster). UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability. Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias). Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. @@ -100,7 +101,9 @@ is retired and returns 400 `Invalid model name`. | Model | GPU | Weight | RPM Cap | Timeout | |-------|-----|--------|---------|---------| -| Qwen3.8-27B-Uncensored-Q4_K_M | RTX 3090 (.8:8080) | **0.55** | 500 | **300s** | +| `qwen3.6-27B-code` | RTX 3090 (.8:8080) | **0.70** | 500 | **300s** | +| `strix-moe` | Strix Halo (.15:8080) | **0.20** | 60 | **300s** | +| `gpu-vision` | RTX 5070 (.110:8080) | **0.10** | 200 | **300s** | Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path. diff --git a/litellm-health.prose.md b/litellm-health.prose.md index f0134a2..aca07e0 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -75,7 +75,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - SSH key access to backend_host for container checks - Network access to public_url, auth_host, and gpu_dashboard_url -- LiteLLM master key for key management endpoints +- LiteLLM master key for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`) +- The dedicated `monitor` agent key on CT 116 at `/etc/litellm-monitor.env` (root-only 0600) for + model inference checks — the master key must never be used for inference ## GPU Fleet Topology diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 3768941..6101259 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -70,17 +70,17 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) `model_name`s served (live registry, verified 2026-09-12 with the master key): `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. `gemma-4-12b`, `gpu-light`, and `crew-auto` are retired and absent from the registry. -### Context Cap Split (2026-08-20) +### Context Cap Split (2026-08-20; crew cap RETIRED) - **Abiba (firstmate)**: 128K uncapped — unlimited context for primary workloads - **Hermes agents** (mumuni, tanko, koby, koonimo): 128K uncapped -- **Crewmates** (ops, tune, verify, auth-keys, build): 64K capped — alias `crew-auto` enforces 64K limit +- **Crewmates** (ops, tune, verify, auth-keys, build): the 64K cap was retired together with the `crew-auto` alias; NO context cap is currently in force. -Preferred implementation: uncap shared pool, add capped alias for crew-only. +> The 64K crew cap was retired with `crew-auto` (2026-09-12). No limit is currently in force; reinstating one would need per-key model limits as a separate, deliberately-scoped change. - `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.70, rpm 500, api_base .8) + strix-moe (0.20, rpm 60, api_base .15) + gpu-vision (0.10, rpm 200, api_base .110). - `gpu-dense` is the high-rpm alias (rpm 500) onto qwen3.6-27B-code; `gpu-light` is retired — the RTX 5070 alias is now `gpu-vision`. -- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','strix-moe','gpu-dense']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard set only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them. `/v1/models` is key-scoped, so a model list is only meaningful with the key it was read with. +- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed. A representative agent key (read 2026-09-12) returns `deepseek-v4-pro`, `gpu-dense`, `gpu-vision`, `qwen3.6-27B-code`, `qwen3.6-35B-udq4`, `strix-moe`, `syslog-auto` — `gpu-vision` present, `gpu-light` and `gemma-4-12b` absent. Agents should still use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them. - **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`. ## Model Fallback Chains (LiteLLM)