gpu-fleet: July 8 optimization sweep — parallel 2 fleet-wide, 128K ctx on NVIDIA, LiteLLM timeout fixes
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
PR Pipeline — Authorize → Validate → Review → Merge / auth (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / validate (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / lint (push) Successful in 2s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (push) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (push) Successful in 0s
- All GPUs: --parallel 1 → 2 (6 concurrent slots, was 3) - .8 RTX 3090: ctx 256K→128K, VRAM 96%→83%, turbo4 KV cache - .110 RTX 5070: ctx 256K→128K, ubatch 4096→512 (was inverted), VRAM 90%→77%, q4_0 KV - .15 Strix Halo: parallel 1→2, 256K ctx (41GB free), q8_0 KV, AMD metrics via /sys/class/drm - LiteLLM: gemma timeout 25→120s, qwen timeout 40→90s, syslog-auto (qwen) 40→90s - Agent configs: context_length 262144 for syslog-auto, 131072 for direct qwen/gemma - Updated health-check operation, agent config implications, benchmark table (fixed model↔GPU mapping)
This commit is contained in:
+42
-17
@@ -5,8 +5,9 @@ description: >
|
||||
Manages the GPU inference fleet across all hosts. Handles model deployment,
|
||||
registration, health checks, LiteLLM sync, agent key management, GPU
|
||||
saturation watchdog, Prometheus/Grafana monitoring, and self-healing.
|
||||
Current as of 2026-06-30 post-migration: nginx routes /v1 → LiteLLM,
|
||||
router runs behind LiteLLM on :9000 (internal-only).
|
||||
Current as of 2026-07-08: context reduced to 128K on NVIDIA GPUs, parallel 2
|
||||
on all GPUs, LiteLLM timeouts tuned (gemma 25→120s, qwen 40→90s), router fully
|
||||
deprecated — nginx routes /v1 → LiteLLM directly.
|
||||
agent: abiba
|
||||
triggers:
|
||||
- on model add/remove
|
||||
@@ -63,21 +64,21 @@ triggers:
|
||||
│ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │
|
||||
│ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │
|
||||
│ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │
|
||||
│ │ │ │ │ llama-srv │ │ Watchdog │
|
||||
│ 128K ctx │ │ 128K ctx │ │ 256K ctx │ │ Watchdog │
|
||||
│ qwen3.6 │ │ gemma-4-12b │ │ ornith35B │ │ Prometheus │
|
||||
│ 27B-code │ │ :8090 │ │ :8080 │ │ exporter │
|
||||
│ :8090 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
||||
│ 27B-code │ │ :8080 │ │ :8080 │ │ exporter │
|
||||
│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │
|
||||
│ :9400 │ └─────────────┘ └───────────┘ └──────────────┘
|
||||
└──────────┘
|
||||
```
|
||||
|
||||
## Current Model Assignments
|
||||
## Current Model Assignments (2026-07-08)
|
||||
|
||||
| Model | GPU | Host | VRAM | Status |
|
||||
|-------|-----|------|------|--------|
|
||||
| qwen3.6-27B-code | RTX 3090 | .8 (llm-gpu) | 23.4/24GB | ✅ healthy |
|
||||
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 11.2/12.2GB | ✅ healthy |
|
||||
| ornith-1.0-35b | Strix Halo Vulkan | .15 (amdpve) | 22.7/64GB | ✅ healthy |
|
||||
| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status |
|
||||
|-------|-----|------|------|-----|----------|----------|-------------|--------|
|
||||
| qwen3.6-27B-code | RTX 3090 | .8 (llm-gpu) | 20.3/24GB (83%) | 128K | turbo4 | 2 | 512/512 | ✅ healthy |
|
||||
| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 9.4/12.2GB (77%) | 128K | q4_0 | 2 | 2048/512 | ✅ healthy |
|
||||
| ornith-1.0-35b | Strix Halo Vulkan | .15 (amdpve) | 24.4/64GB (35%) | 256K | q8_0 | 2 | 2048/512 | ✅ healthy |
|
||||
|
||||
## Operations
|
||||
|
||||
@@ -121,7 +122,19 @@ triggers:
|
||||
7. Document keys in knowledge graph
|
||||
|
||||
### list
|
||||
Show full fleet status: GPUs, models, VRAM, active requests, circuit breakers, keys
|
||||
Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, active requests, circuit breakers, keys
|
||||
|
||||
### health-check
|
||||
1. Check GPU hardware: nvidia-smi (.8, .110) + amdgpu sysfs (.15 via /sys/class/drm/card0/device/)
|
||||
2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts
|
||||
3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!")
|
||||
4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models`
|
||||
5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml`
|
||||
- gemma-4-12b: 120s, qwen3.6-27B: 90s, ornith-1.0-35b: 120s
|
||||
- global request_timeout: 300s, nginx proxy_read_timeout: 600s
|
||||
6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power)
|
||||
7. Check port conflicts: verify only one llama-server on :8080 per host
|
||||
8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`)
|
||||
|
||||
## Agent Keys (LiteLLM DB — Current 2026-06-30)
|
||||
|
||||
@@ -168,9 +181,11 @@ If no SSH access, send Zulip DM via abiba-bot.
|
||||
- **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first.
|
||||
Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster().
|
||||
- **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround.
|
||||
- **VRAM**: RTX 3090 at ~9.6/24GB (60% free), RTX 5070 at ~6.5/12GB (46% free). Healthy.
|
||||
- **RTX 3090 optimization**: `--parallel 2 --batch-size 4096` — doubled concurrent throughput.
|
||||
- **RTX 5070 optimization**: Removed draft model freed ~2GB VRAM for vision tasks. `--parallel 2`.
|
||||
- **VRAM (2026-07-08)**: RTX 3090 at 20.3/24GB (83%), RTX 5070 at 9.4/12.2GB (77%), Strix Halo at 24.4/64GB (35%). Context reduced from 256K→128K on NVIDIA GPUs freed ~3.3GB (.8) and ~1.5GB (.110).
|
||||
- **All GPUs at `--parallel 2` (2026-07-08)**: Fleet serves 6 concurrent requests (was 3). 2× throughput.
|
||||
- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2`. No explicit batch flags (512/512 default). Service: `/home/llmuser/llama-wrapper.sh`.
|
||||
- **RTX 5070 config**: `--ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --batch-size 2048 --ubatch-size 512 --parallel 2`. Ubatch fixed 4096→512 (was inverted — ubatch > batch killed prompt throughput). Service: `/home/llmuser/llama-wrapper.sh`.
|
||||
- **LiteLLM timeout tuning (2026-07-08)**: gemma-4-12b 25→120s, qwen3.6-27B-code 40→90s, syslog-auto (qwen route) 40→90s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. Config at `/opt/inference-harness/litellm_config.yaml`.
|
||||
- **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `ornith-server.service` on port 8080, 256K context, flash-attn + q8 KV.
|
||||
- **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill).
|
||||
- **Strix Halo thermal safeguard (2026-07-02)**: `ornith-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`.
|
||||
@@ -185,8 +200,8 @@ If no SSH access, send Zulip DM via abiba-bot.
|
||||
|
||||
| GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Samples |
|
||||
|-----|-------|-----------|--------------|----------|---------|
|
||||
| RTX 3090 (.8) | gemma-4-12b | 75 | 305 | 74 | 6 |
|
||||
| RTX 5070 (.110) | qwen3.6-27B-code | 75 | 323 | 75 | 6 |
|
||||
| RTX 3090 (.8) | qwen3.6-27B-code | 75 | 305 | 74 | 6 |
|
||||
| RTX 5070 (.110) | gemma-4-12b | 75 | 323 | 75 | 6 |
|
||||
| Strix Halo (.15) | ornith-1.0-35b | 70 | 532 | 70 | 6 |
|
||||
|
||||
Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes.
|
||||
@@ -194,3 +209,13 @@ Degradation alerts fire at 30% (warning) and 50% (critical) below baseline.
|
||||
History stored at `/root/data/toks-history.json` with 7-day rolling window.
|
||||
|
||||
**Note (2026-07-01)**: Strix Halo prompt tok/s jumped 209→532 after Vulkan rebuild (cooperative-matrix fast path now active on GFX1151). Baseline may need re-calibration.
|
||||
|
||||
## Agent Config Implications (2026-07-08)
|
||||
|
||||
With NVIDIA GPUs at 128K context:
|
||||
- Agents using `syslog-auto` (50/50 qwen+ornith): keep `context_length: 262144` — ornith supports it, Litellm fallbacks handle qwen overflow
|
||||
- Agents using `qwen3.6-27B-code` directly: set `context_length: 131072` and `max_tokens: 4096` per thermal safety rule
|
||||
- Agents using `gemma-4-12b` directly (auxiliary tasks): set `context_length: 131072`
|
||||
- Compression threshold at 0.65: fires at ~170K for syslog-auto (262K ctx), ~85K for direct qwen/gemma (128K ctx)
|
||||
- All Hermes clients MUST set `max_tokens: 4096` — first line of defense before server-side `-n 8192` cap
|
||||
- Port 8080 is used on all 3 GPU hosts (not 8090 as previously documented)
|
||||
|
||||
Reference in New Issue
Block a user