diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index e952fea..cebe3b4 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -80,56 +80,29 @@ triggers: Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names. When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched. -| Alias | GPU | Current Model | Will Route To | -|-------|-----|---------------|---------------| -| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | -| `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 | -| `gpu-vision` | RTX 5070 (.110) | gpu-vision | Whatever runs on RTX 5070 | +| Alias | Serves | Where | Kind | +|-------|--------|-------|------| +| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias | +| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member | +| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias | +| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router | + +Single source of truth for models, aliases, rpm caps, weights and fallback chains: +CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in +contracts — read them there. **Backward compatibility**: Old model-specific names (qwen3.6-27B-code, qwen3.5-9b-it) still work but are deprecated for agent configs. Only the stable aliases survive model swaps. `gemma-4-12b` is retired and returns 400 `Invalid model name`. -## Current Model Assignments (2026-07-15) - -| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | -|-------|-----|------|------|-----|----------|----------|-------------|--------| - ## Routing Configuration (LiteLLM — July 2026) -### syslog-auto Weighted Pool (Direct GPU — bypasses router) - -| Model | GPU | Weight | RPM Cap | Timeout | -|-------|-----|--------|---------|---------| -| `qwen3.6-27B-code` | RTX 3090 (.8:8080) | **0.70** | 500 | **300s** | -| `strix-moe` | Strix Halo (.15:8080) | **0.20** | 60 | **300s** | -| `gpu-vision` | RTX 5070 (.110:8080) | **0.10** | 200 | **300s** | +Single source of truth for models, aliases, rpm caps, weights and fallback chains: +CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in +contracts — read them there. Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) was decommissioned 2026-09-11 and is NOT in the inference path. -### Direct Model Endpoints - -| Model | RPM Cap | Notes | -|-------|---------|-------| - -### Stable Aliases (for agent configs — never change) - -| Alias | RPM Cap | Routes To | Purpose | -|-------|---------|-----------|---------| -| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) | -| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning | -| `gpu-vision` | 500 | RTX 5070 | Vision, web extract, light tasks | - -### Fallback Chains (live `router_settings.fallbacks`, request_timeout 300s throughout) -- syslog-auto → qwen3.6-27B-code → strix-moe → gpu-vision -- qwen3.6-27B-code → strix-moe -- strix-moe → qwen3.6-27B-code → gpu-vision - -### Why Strix Halo RPM Is Capped -- Direct (strix-moe): 40 RPM (tight) — Strix Halo is shared with compression tasks -- Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously -- Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C - ## Operations ### add-model @@ -179,9 +152,7 @@ Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, act 2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts 3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!") 4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models` -5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` - - Qwen3.5-9B: 120s, qwen3.6-27B-code: 300s, Carnice-Qwen3.6-MoE-35B-A3B/strix-moe: 300s (strix-moe alias retained, legacy name qwen3.6-35B-udq4 deprecated) - - global request_timeout: 300s, nginx proxy_read_timeout: 600s +5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` — read the live values from the authority config; do not assert them from this contract. 6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power) 7. Check port conflicts: verify only one llama-server on :8080 per host 8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`) @@ -268,9 +239,9 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window. ### Stable Aliases — CRITICAL All agent configs MUST use stable role-based aliases, never model-specific names: -- `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`) -- `auxiliary.vision.model: gpu-vision` (NOT `gemma-4-12b`) -- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`) +- `compression.model: strix-moe` +- `auxiliary.vision.model: gpu-vision` +- `delegation.model: gpu-dense` - `auxiliary.web_extract.model: gpu-vision` When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched. @@ -288,7 +259,7 @@ Mumuni (kagentz CT105, 192.168.68.14 — migrated from CT100 2026-08-29) is the | Setting | Value | Notes | |---------|-------|-------| -| `model.default` | `syslog-auto` | Weighted pool (70% qwen, 20% strix, 10% gpu-vision) | +| `model.default` | `syslog-auto` | Balanced default (pool router) | | `model.provider` | `custom:litellm` | LiteLLM on CT116 | | `compression.model` | `strix-moe` | Stable alias — survives model swaps | | `aux.compression.model` | `strix-moe` | Compression auxiliary model | diff --git a/litellm-health.prose.md b/litellm-health.prose.md index aca07e0..a28b555 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -81,23 +81,16 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) ## GPU Fleet Topology -| Host | IP | Hardware | Models Served | Engine | Context | Parallel | -|------|-----|----------|---------------|--------|---------|----------| -| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 | -| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gpu-vision | llama-server systemd | **128K** | 2 | -| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 | +| Host | IP | Hardware | Role | +|------|-----|----------|------| +| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) | +| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) | +| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) | -## Model Fallback Chains (LiteLLM) - -| Primary | Timeout | Fallback chain (live `router_settings.fallbacks`) | Timeout | -|---------|---------|--------------------------------------------------|---------| -| syslog-auto (balanced) | 300s | qwen3.6-27B-code → strix-moe → gpu-vision | 300s | -| qwen3.6-27B-code | 300s | strix-moe | 300s | -| strix-moe | 300s | qwen3.6-27B-code → gpu-vision | 300s | -| gpu-vision | 300s | — (leaf) | — | - -> Global: request_timeout=300s, nginx proxy_read_timeout=600s. `gemma-4-12b` was retired -> and is NOT in the registry — do not re-add it to this table. +Single source of truth for models, aliases, rpm caps, weights and fallback chains: +CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in +contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`, +`crew-auto`). ## Containers on CT 116 @@ -163,11 +156,11 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to this list. The RTX 5070 host now serves `gpu-vision`. - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so - an agent key can return a different set than the master key. Always state which key a - model list was read with. Verified 2026-09-12 on the backend surface - (`http://{{backend_host}}/litellm/v1/models`) with the `monitor` key: - `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, - `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. + the set depends on the key. Always state which key a model list was read with. This + probe uses the `monitor` key; on the backend surface + (`http://{{backend_host}}/litellm/v1/models`) that key returns `gpu-vision`, + `qwen3.6-27B-code`, `strix-moe`, `syslog-auto` (verified 2026-09-12). The master key + sees a larger registry — read that from the authority config, not from this probe. 8. **Check agent keys**: - GET http://{{backend_host}}/litellm/key/list with master key (admin endpoint) → verify all 6 agents have keys diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index 6101259..f3f8347 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -58,17 +58,29 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) ## GPU Fleet Topology -| Host | IP | Hardware | Models Served | Engine | Context | Parallel | -|------|-----|----------|---------------|--------|---------|----------| -| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 | -| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gpu-vision | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 | -| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 | +| Host | IP | Hardware | Role | +|------|-----|----------|------| +| llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | heavy reasoning (`gpu-dense`) | +| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | vision / web extract / light tasks (`gpu-vision`) | +| amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | compression (`strix-moe`) | -> Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced. +> Verified on the ground 2026-07-16. The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced. -## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116) +## LiteLLM Model Surface -`model_name`s served (live registry, verified 2026-09-12 with the master key): `qwen3.6-27B-code`, `gpu-vision`, `gpu-dense`, `qwen3.6-35B-udq4`, `qwen3.8-27B-uncensored`, `strix-moe`, `syslog-auto`. `gemma-4-12b`, `gpu-light`, and `crew-auto` are retired and absent from the registry. +Single source of truth for models, aliases, rpm caps, weights and fallback chains: +CT 116 `/opt/inference-harness/litellm_config.yaml`. Do not duplicate those values in +contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-light`, +`crew-auto`). + +### Alias Surface + +| Alias | Serves | Where | Kind | +|-------|--------|-------|------| +| `gpu-dense` | heavy reasoning | RTX 3090 (192.168.68.8) | direct alias | +| `gpu-vision` | vision / web extract / light tasks | RTX 5070 (192.168.68.110) | direct alias AND `syslog-auto` pool member | +| `strix-moe` | compression (MoE) | Strix Halo (192.168.68.15) | direct alias | +| `syslog-auto` | balanced default | weighted pool across the three GPU hosts | pool router | ### Context Cap Split (2026-08-20; crew cap RETIRED) @@ -78,22 +90,9 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) > The 64K crew cap was retired with `crew-auto` (2026-09-12). No limit is currently in force; reinstating one would need per-key model limits as a separate, deliberately-scoped change. -- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.70, rpm 500, api_base .8) + strix-moe (0.20, rpm 60, api_base .15) + gpu-vision (0.10, rpm 200, api_base .110). -- `gpu-dense` is the high-rpm alias (rpm 500) onto qwen3.6-27B-code; `gpu-light` is retired — the RTX 5070 alias is now `gpu-vision`. -- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed. A representative agent key (read 2026-09-12) returns `deepseek-v4-pro`, `gpu-dense`, `gpu-vision`, `qwen3.6-27B-code`, `qwen3.6-35B-udq4`, `strix-moe`, `syslog-auto` — `gpu-vision` present, `gpu-light` and `gemma-4-12b` absent. Agents should still use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them. +- Key scoping: `/v1/models` is key-scoped, so the set a caller sees must be read with a named key rather than assumed — a monitor key, an agent key, and the master key can each return a different set. Agents should use the stable aliases (`strix-moe`, `gpu-vision`, `gpu-dense`) rather than raw model names, so model swaps don't break them. - **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`. -## Model Fallback Chains (LiteLLM) - -| Primary | Timeout | Fallback chain (live `router_settings.fallbacks`) | Timeout | -|---------|---------|--------------------------------------------------|---------| -| syslog-auto (balanced) | 300s | qwen3.6-27B-code → strix-moe → gpu-vision | 300s | -| qwen3.6-27B-code | 300s | strix-moe | 300s | -| strix-moe | 300s | qwen3.6-27B-code → gpu-vision | 300s | -| gpu-vision | 300s | — (leaf) | — | - -> Global: request_timeout=300s, nginx proxy_read_timeout=600s - ## Containers on CT 116 | Container | Image | Port | Health Check |