diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 390c36d..0704ff7 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -5,13 +5,12 @@ description: > Manages the GPU inference fleet across all hosts. Handles model deployment, registration, health checks, LiteLLM sync, agent key management, GPU saturation watchdog, Prometheus/Grafana monitoring, and self-healing. - UPDATED 2026-07-12: Architecture is DIRECT GPU — LiteLLM routes directly - to llama-server on each GPU host (no router in inference path). Router - (port 9000) is running but NOT in request path. All GPUs standardized on - api-key 'not-needed'. RTX 5070 had api-key mismatch (sk-loc...5678) that - caused cascading 401→timeout→401 fallback loops — fixed. - Context: RTX 3090 verified at 256K (was documented as 128K — WRONG). - Workload: Compression moved to Strix Halo, RTX 5070 → vision/web only. + UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe, + gpu-dense, gpu-light. These never change — only the underlying model does. + Strix Halo: ornith-1.0-35b → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). + RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster). + RTX 5070 context: 131K → 256K. VRAM: 88% (10.8/12.2GB). + Compression timeout: 300s (was 120s). Mumuni context: 128K (was 256K). agent: abiba triggers: - on model add/remove @@ -76,40 +75,64 @@ triggers: └──────────┘ ``` -## Current Model Assignments (2026-07-12) +## Stable Role-Based Aliases (Introduced 2026-07-15) + +Agent configs, cron jobs, and workflows MUST use these aliases, never model-specific names. +When a model is swapped on a GPU, ONLY the infrastructure layer changes — agent configs are untouched. + +| Alias | GPU | Current Model | Will Route To | +|-------|-----|---------------|---------------| +| `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | +| `gpu-dense` | RTX 3090 (.8) | qwen3.6-27B-code | Whatever runs on RTX 3090 | +| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 | + +**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.6-35B-udq4) still work +but are deprecated for agent configs. Only the stable aliases survive model swaps. + +## Current Model Assignments (2026-07-15) | Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | |-------|-----|------|------|-----|----------|----------|-------------|--------| -| qwen3.6-27B-code | RTX 3090 | .8 (llm-gpu) | 20.7/24GB (84%) | **256K** | turbo4 | 1 | default | ✅ healthy | +| qwen3.6-27B-code (MTP) | RTX 3090 | .8 (llm-gpu) | 22.2/24.6GB (90%) | **256K** 🚀 | turbo4 | 2 | default | ✅ 63 tok/s | | gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 10.0/12.2GB (82%) | 131K | q4_0 | 2 | 2048/1024 | ✅ healthy | -| ornith-1.0-35b | Strix Halo Vulkan | .15 (amdpve) | ~9GB/64GB | 256K | q8_0 | 2 | 2048/512 | ✅ healthy | +| qwen3.6-35B-udq4 | Strix Halo Vulkan | .15 (amdpve) | ~9GB/64GB | 256K | q8_0 | 2 | 2048/512 | ✅ healthy | ## Routing Configuration (LiteLLM — July 2026) -### syslog-auto Weighted Pool +### syslog-auto Weighted Pool (Direct GPU — bypasses router) -| Model | GPU | Weight | RPM Cap | Purpose | +| Model | GPU | Weight | RPM Cap | Timeout | |-------|-----|--------|---------|---------| -| qwen3.6-27B-code | RTX 3090 | 0.55 | 500 | Heavy reasoning, code, long context | -| ornith-1.0-35b | Strix Halo | 0.30 | **60** | Agentic workflows, tool calling | -| gemma-4-12b | RTX 5070 | 0.15 | 200 | Overflow + vision | +| qwen3.6-27B-code | RTX 3090 (.8:8080) | **0.55** | 500 | 90s | +| qwen3.6-35B-udq4 | Strix Halo (.15:8080) | **0.30** | 60 | **300s** | +| gemma-4-12b | RTX 5070 (.110:8080) | **0.15** | 200 | **300s** | + +Note: All syslog-auto entries route directly to GPUs with `api_key: not-needed`. The router (port 9000) is NOT in the inference path. ### Direct Model Endpoints | Model | RPM Cap | Notes | |-------|---------|-------| -| ornith-1.0-35b | 40 | Tight cap — prevents Strix overload | +| qwen3.6-35B-udq4 | 40 | Tight cap — prevents Strix overload | | qwen3.6-27B-code | 500 | High cap — primary workhorse | -| gemma-4-12b | 500 | High cap — fast 12B | +| gemma-4-12b | 500 | High cap — IQ4_NL+MTP, 122 tok/s | + +### Stable Aliases (for agent configs — never change) + +| Alias | RPM Cap | Routes To | Purpose | +|-------|---------|-----------|---------| +| `strix-moe` | 40 | Strix Halo | Compression tasks (MoE models) | +| `gpu-dense` | 500 | RTX 3090 | Heavy reasoning | +| `gpu-light` | 500 | RTX 5070 | Vision, web extract, light tasks | ### Fallback Chains - gemma → qwen - qwen → gemma -- ornith → qwen → gemma -- syslog-auto → qwen → gemma → ornith +- qwen3.6-35B-udq4 → qwen → gemma +- syslog-auto → qwen → gemma → qwen3.6-35B-udq4 -### Why ornith RPM Is Capped -- Direct: 40 RPM (tight) — Strix Halo is shared with compression tasks +### Why Strix Halo RPM Is Capped +- Direct (qwen3.6-35B-udq4): 40 RPM (tight) — Strix Halo is shared with compression tasks - Via syslog-auto: 60 RPM (moderate) — prevents flooding when multiple agents use syslog-auto simultaneously - Combined max: ~100 RPM across both paths — Strix Halo can sustain this at 80°C @@ -203,7 +226,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions. | `/root/scripts/gpu-saturation-watchdog.py` | pi (.24) | Auto-restart stuck llama-server | | `/root/dashboard/gpu-fleet.html` | pi (.24) | Live HTML dashboard | | `/etc/systemd/system/llama-server.service` | .8, .110 | llama-server daemons (Nvidia GPUs) | -| `/etc/systemd/system/ornith-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo). Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. | +| `/etc/systemd/system/strix-server.service` | .15 (amdpve) | llama-server daemon (Vulkan, Strix Halo) running unsloth/Qwen3.6-35B-A3B-MTP-GGUF. Note: `llama-server.service` and `llama-server@.service` are **masked** on .15 to prevent port 8080 collisions. | ## Prometheus & Grafana @@ -220,14 +243,14 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions. - **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first. Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster(). - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. -- **VRAM (2026-07-12)**: RTX 3090 at 20.7/24GB (84%), RTX 5070 at 10.0/12.2GB (82%). RTX 3090 context increased to 256K (was incorrectly documented as 128K — verified via /proc/PID/cmdline). -- **RTX 3090 runs `--parallel 1`** (verified 2026-07-12). RTX 5070 and Strix at parallel 2. -- **RTX 3090 config**: `-c 262144 -ctk turbo4 -ctv turbo4 --parallel 1 --flash-attn on --cont-batching`. Service: `/home/llmuser/llama-wrapper.sh`. -- **RTX 5070 config**: `--ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --batch-size 2048 --ubatch-size 1024 --parallel 2`. Api-key standardized to `not-needed` (was `sk-loc...5678` causing 401 loops). Service: `/home/llmuser/llama-wrapper.sh`. +- **VRAM (2026-07-15)**: RTX 3090 at 22.2/24.6GB (90%) with **256K context** (corrected from 131K). RTX 5070 at 10.8/12.2GB (88%) with 256K context + MTP. Strix Halo at ~9GB/64GB. +- **RTX 3090 runs `--parallel 2`** with MTP draft (spec-type draft-mtp, spec-draft-n-max 2). +- **RTX 3090 config**: `-c 262144 -ctk turbo4 -ctv turbo4 --parallel 2 --flash-attn on --cont-batching --spec-type draft-mtp`. Context corrected to 256K (2026-07-15). VRAM: 90%. Service: `/home/llmuser/llama-wrapper.sh`. +- **RTX 5070 config (2026-07-15)**: Switched to IQ4_NL + MTP draft (Q8_0) at 256K context. Gen speed: 122 tok/s (was 70). VRAM: 10.8/12.2GB (88%). No draft model pre-upgrade due to VRAM constraints. Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model gemma-4-12b-it-IQ4_NL.gguf --spec-draft-model gemma-4-12b-it-Q8_0-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 --ctx-size 262144`. - **LiteLLM timeout tuning (2026-07-08)**: gemma-4-12b 25→120s, qwen3.6-27B-code 40→90s, syslog-auto (qwen route) 40→90s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. Config at `/opt/inference-harness/litellm_config.yaml`. -- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `ornith-server.service` on port 8080, 256K context, flash-attn + q8 KV. +- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080 (was `ornith-server.service`), model changed to `unsloth/Qwen3.6-35B-A3B-MTP-GGUF` (UD-Q4_K_M), alias `qwen3.6-35B-udq4`, 256K context, flash-attn + q8 KV. MTP support enabled for 1.4-2.2x faster inference. - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). -- **Strix Halo thermal safeguard (2026-07-02)**: `ornith-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. +- **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. - **Port 8080 firewall**: amdpve iptables restricts 8080 to 192.168.68.116 (LiteLLM/router host) only. All inbound connections are from .116 (LiteLLM proxied via nginx). Localhost curls hang (SYN dropped). Always test from .116. - **Router sidecar fallback**: `router.py` `check_gpu_health()` now probes GPU `/health` directly when sidecar at :8090 is absent. Sidecar JSON exporters not deployed on any GPU host — router relies on GPU-direct fallback. - **Router GPU_MOE_URL bug (fixed 2026-07-01)**: docker-compose had `GPU_MOE_URL=.110:8080` (gemma host) instead of `.15:8080` (amdpve). Corrected. @@ -237,11 +260,14 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions. ## GPU Inference Benchmarks (Current) -| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Samples | +| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Context | |-----|-------|-----------|--------------|----------|---------| -| RTX 3090 (.8) | qwen3.6-27B-code | 75 | 305 | 74 | 6 | -| RTX 5070 (.110) | gemma-4-12b | 75 | 323 | 75 | 6 | -| Strix Halo (.15) | ornith-1.0-35b | 70 | 532 | 70 | 6 | +| RTX 3090 (.8) | qwen3.6-27B-code (MTP) | **63** | — | — | 131K | +| RTX 5070 (.110) | gemma-4-12b (IQ4_NL+MTP) | **191** | — | — | **256K** | +| Strix Halo (.15) | qwen3.6-35B-udq4 | **71** | — | — | **256K** | + +Benchmarks from 2026-07-15 verification run. RTX 5070 MTP provides 2.7x speedup over pre-upgrade 70 tok/s. +All 3 GPUs now at 256K context (2026-07-15). Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes. Degradation alerts fire at 30% (warning) and 50% (critical) below baseline. @@ -249,12 +275,55 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window. **Note (2026-07-01)**: Strix Halo prompt tok/s jumped 209→532 after Vulkan rebuild (cooperative-matrix fast path now active on GFX1151). Baseline may need re-calibration. -## Agent Config Implications (2026-07-12) +## Agent Config Implications (2026-07-15) -With RTX 3090 at 256K context (verified July 2026): -- Agents using `syslog-auto` (55/30/15 qwen+ornith+gemma): `context_length: 262144` — ornith and qwen both support it -- Agents using `qwen3.6-27B-code` directly: `context_length: 262144` (256K ctx verified) -- Agents using `gemma-4-12b` directly (auxiliary tasks): `context_length: 131072` -- Compression threshold at 0.65: fires at ~170K for 262K context window on Strix Halo -- All Hermes clients MUST set `max_tokens: 4096` — first line of defense before server-side `-n 8192` cap -- Port 8080 is used on all 3 GPU hosts (not 8090 as previously documented) +### Stable Aliases — CRITICAL + +All agent configs MUST use stable role-based aliases, never model-specific names: +- `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`) +- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`) +- `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`) +- `auxiliary.web_extract.model: gpu-light` + +When the underlying model is swapped, only the LiteLLM config changes — agent configs are untouched. + +### Context Windows +- RTX 3090: **256K** (was 131K, bumped 2026-07-15) | RTX 5070: **256K** (up from 131K) | Strix Halo: **256K** +- **Mumuni compression context**: 128K (down from 256K) — ensures compression model doesn't timeout +- Compression threshold 0.65: fires at ~85K for 128K context window +- Mumuni compression model alias: `strix-moe` with 300s timeout + +### Mumuni Agent Profile + +Mumuni (CT114, 192.168.68.123) is the primary business assistant. This profile is the reference for all agent configs: + +| Setting | Value | Notes | +|---------|-------|-------| +| `model.default` | `syslog-auto` | Weighted pool (55% qwen, 30% strix, 15% gemma) | +| `model.provider` | `custom:litellm` | LiteLLM on CT116 | +| `compression.model` | `strix-moe` | Stable alias — survives model swaps | +| `aux.compression.model` | `strix-moe` | Compression auxiliary model | +| `aux.vision.model` | `gpu-light` | Vision tasks (RTX 5070) | +| `aux.web_extract.model` | `gpu-light` | Web extraction | +| `delegation.model` | `gpu-dense` | Sub-agent reasoning (RTX 3090) | +| `context.max_context_window` | 131072 (128K) | Kept at 128K to avoid compression timeout | +| `compression.threshold` | 0.65 | Triggers at ~85K | +| `compression.target_ratio` | 0.3 | Compresses to ~38K | +| `compression.protect_last_n` | 40 | Preserves last 40 messages | +| `memory.memory_char_limit` | 800 | Brief memory entries | +| `personalities` | `creative` | Creative assistant personality | +| Platforms | cli, discord, homeassistant, signal, telegram, zulip | All Hermes platforms | +| Main model timeout | 300s | LiteLLM global timeout | +| Compression model timeout | 300s | ornith timeout increased from 120s | + +### Agent Update Status (2026-07-15) + +| Agent | Host | Status | +|-------|------|--------| +| **Mumuni** | CT114 (.123) | ✅ Updated to stable aliases | +| **Tanko** | CT112 (.122) | ✅ Updated to stable aliases | +| **Koby** | CT111 (.129) | ❌ SSH unreachable — needs Zulip DM | +| **Koonimo** | CT113 | ❌ SSH unreachable — needs Zulip DM | +| **Kagenz0** | CT105 | ❌ SSH unreachable — needs Zulip DM | + +All Hermes clients MUST set `max_tokens: 4096` — first line of defense before server-side `-n 8192` cap.