From 776df418ca02096295d0337c3d700222cdb2f9e8 Mon Sep 17 00:00:00 2001 From: root Date: Wed, 8 Jul 2026 17:25:42 +0000 Subject: [PATCH] =?UTF-8?q?gpu-fleet:=20July=208=20optimization=20sweep=20?= =?UTF-8?q?=E2=80=94=20parallel=202=20fleet-wide,=20128K=20ctx=20on=20NVID?= =?UTF-8?q?IA,=20LiteLLM=20timeout=20fixes?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - All GPUs: --parallel 1 → 2 (6 concurrent slots, was 3) - .8 RTX 3090: ctx 256K→128K, VRAM 96%→83%, turbo4 KV cache - .110 RTX 5070: ctx 256K→128K, ubatch 4096→512 (was inverted), VRAM 90%→77%, q4_0 KV - .15 Strix Halo: parallel 1→2, 256K ctx (41GB free), q8_0 KV, AMD metrics via /sys/class/drm - LiteLLM: gemma timeout 25→120s, qwen timeout 40→90s, syslog-auto (qwen) 40→90s - Agent configs: context_length 262144 for syslog-auto, 131072 for direct qwen/gemma - Updated health-check operation, agent config implications, benchmark table (fixed model↔GPU mapping) --- gpu-fleet.prose.md | 59 ++++++++++++++++++++++--------- hermes-agent-baseline.prose.md | 62 +++++++++++++++++++++++++++++---- hermes-config-template.prose.md | 33 +++++++++++++++--- infrastructure-update.prose.md | 4 ++- zulip-adapter-lessons.prose.md | 20 +++++++++++ 5 files changed, 149 insertions(+), 29 deletions(-) diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index f55cf9b..9e14cfd 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -5,8 +5,9 @@ description: > Manages the GPU inference fleet across all hosts. Handles model deployment, registration, health checks, LiteLLM sync, agent key management, GPU saturation watchdog, Prometheus/Grafana monitoring, and self-healing. - Current as of 2026-06-30 post-migration: nginx routes /v1 → LiteLLM, - router runs behind LiteLLM on :9000 (internal-only). + Current as of 2026-07-08: context reduced to 128K on NVIDIA GPUs, parallel 2 + on all GPUs, LiteLLM timeouts tuned (gemma 25→120s, qwen 40→90s), router fully + deprecated — nginx routes /v1 → LiteLLM directly. agent: abiba triggers: - on model add/remove @@ -63,21 +64,21 @@ triggers: │ CT 8 │ │ CT 110 │ │ CT 15 │ │ pi (.24) │ │ RTX 3090 │ │ RTX 5070 │ │ Strix Halo│ │ GPU Monitor │ │ 24GB │ │ 12GB │ │ 64GB UMA │ │ :9100 │ -│ │ │ │ │ llama-srv │ │ Watchdog │ +│ 128K ctx │ │ 128K ctx │ │ 256K ctx │ │ Watchdog │ │ qwen3.6 │ │ gemma-4-12b │ │ ornith35B │ │ Prometheus │ -│ 27B-code │ │ :8090 │ │ :8080 │ │ exporter │ -│ :8090 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │ +│ 27B-code │ │ :8080 │ │ :8080 │ │ exporter │ +│ :8080 │ │ :9400 (exp) │ │ :9400(exp)│ │ :9401 │ │ :9400 │ └─────────────┘ └───────────┘ └──────────────┘ └──────────┘ ``` -## Current Model Assignments +## Current Model Assignments (2026-07-08) -| Model | GPU | Host | VRAM | Status | -|-------|-----|------|------|--------| -| qwen3.6-27B-code | RTX 3090 | .8 (llm-gpu) | 23.4/24GB | ✅ healthy | -| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 11.2/12.2GB | ✅ healthy | -| ornith-1.0-35b | Strix Halo Vulkan | .15 (amdpve) | 22.7/64GB | ✅ healthy | +| Model | GPU | Host | VRAM | Ctx | KV Cache | Parallel | Batch/Ubatch | Status | +|-------|-----|------|------|-----|----------|----------|-------------|--------| +| qwen3.6-27B-code | RTX 3090 | .8 (llm-gpu) | 20.3/24GB (83%) | 128K | turbo4 | 2 | 512/512 | ✅ healthy | +| gemma-4-12b | RTX 5070 | .110 (ocu-llm) | 9.4/12.2GB (77%) | 128K | q4_0 | 2 | 2048/512 | ✅ healthy | +| ornith-1.0-35b | Strix Halo Vulkan | .15 (amdpve) | 24.4/64GB (35%) | 256K | q8_0 | 2 | 2048/512 | ✅ healthy | ## Operations @@ -121,7 +122,19 @@ triggers: 7. Document keys in knowledge graph ### list -Show full fleet status: GPUs, models, VRAM, active requests, circuit breakers, keys +Show full fleet status: GPUs, models, VRAM, context windows, parallel slots, active requests, circuit breakers, keys + +### health-check +1. Check GPU hardware: nvidia-smi (.8, .110) + amdgpu sysfs (.15 via /sys/class/drm/card0/device/) +2. Check llama-server processes: `ps aux | grep llama-server` on all 3 hosts +3. Check LiteLLM: `curl http://192.168.68.116/health` (expect "I'm alive!") +4. Check LiteLLM models: `curl -H "Authorization: Bearer $MASTER_KEY" http://192.168.68.116/v1/models` +5. Check LiteLLM timeouts: `grep -n 'timeout:' /opt/inference-harness/litellm_config.yaml` + - gemma-4-12b: 120s, qwen3.6-27B: 90s, ornith-1.0-35b: 120s + - global request_timeout: 300s, nginx proxy_read_timeout: 600s +6. Check AMD metrics: `curl http://192.168.68.15:9400/metrics` (Radeon 8060S, util%, VRAM, temp, power) +7. Check port conflicts: verify only one llama-server on :8080 per host +8. Verify agent keys: 9 keys in LiteLLM DB (`GET /key/list`) ## Agent Keys (LiteLLM DB — Current 2026-06-30) @@ -168,9 +181,11 @@ If no SSH access, send Zulip DM via abiba-bot. - **Router startup race**: Compose router.py doesn't call load_roster(). Reload thread sleeps 30s first. Fix: trigger roster reload via SSH after restart, or rebuild image with startup load_roster(). - **LiteLLM /metrics**: Requires auth. Prometheus uses `/health/liveliness` as workaround. -- **VRAM**: RTX 3090 at ~9.6/24GB (60% free), RTX 5070 at ~6.5/12GB (46% free). Healthy. -- **RTX 3090 optimization**: `--parallel 2 --batch-size 4096` — doubled concurrent throughput. -- **RTX 5070 optimization**: Removed draft model freed ~2GB VRAM for vision tasks. `--parallel 2`. +- **VRAM (2026-07-08)**: RTX 3090 at 20.3/24GB (83%), RTX 5070 at 9.4/12.2GB (77%), Strix Halo at 24.4/64GB (35%). Context reduced from 256K→128K on NVIDIA GPUs freed ~3.3GB (.8) and ~1.5GB (.110). +- **All GPUs at `--parallel 2` (2026-07-08)**: Fleet serves 6 concurrent requests (was 3). 2× throughput. +- **RTX 3090 config**: `-c 131072 -ctk turbo4 -ctv turbo4 --parallel 2`. No explicit batch flags (512/512 default). Service: `/home/llmuser/llama-wrapper.sh`. +- **RTX 5070 config**: `--ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --batch-size 2048 --ubatch-size 512 --parallel 2`. Ubatch fixed 4096→512 (was inverted — ubatch > batch killed prompt throughput). Service: `/home/llmuser/llama-wrapper.sh`. +- **LiteLLM timeout tuning (2026-07-08)**: gemma-4-12b 25→120s, qwen3.6-27B-code 40→90s, syslog-auto (qwen route) 40→90s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. Config at `/opt/inference-harness/litellm_config.yaml`. - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `ornith-server.service` on port 8080, 256K context, flash-attn + q8 KV. - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Strix Halo thermal safeguard (2026-07-02)**: `ornith-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (Mumuni) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. @@ -185,8 +200,8 @@ If no SSH access, send Zulip DM via abiba-bot. | GPU | Model | Gen tok/s | Prompt tok/s | Baseline | Samples | |-----|-------|-----------|--------------|----------|---------| -| RTX 3090 (.8) | gemma-4-12b | 75 | 305 | 74 | 6 | -| RTX 5070 (.110) | qwen3.6-27B-code | 75 | 323 | 75 | 6 | +| RTX 3090 (.8) | qwen3.6-27B-code | 75 | 305 | 74 | 6 | +| RTX 5070 (.110) | gemma-4-12b | 75 | 323 | 75 | 6 | | Strix Halo (.15) | ornith-1.0-35b | 70 | 532 | 70 | 6 | Benchmarks run through LiteLLM proxy (192.168.68.116:4001) every 5 minutes. @@ -194,3 +209,13 @@ Degradation alerts fire at 30% (warning) and 50% (critical) below baseline. History stored at `/root/data/toks-history.json` with 7-day rolling window. **Note (2026-07-01)**: Strix Halo prompt tok/s jumped 209→532 after Vulkan rebuild (cooperative-matrix fast path now active on GFX1151). Baseline may need re-calibration. + +## Agent Config Implications (2026-07-08) + +With NVIDIA GPUs at 128K context: +- Agents using `syslog-auto` (50/50 qwen+ornith): keep `context_length: 262144` — ornith supports it, Litellm fallbacks handle qwen overflow +- Agents using `qwen3.6-27B-code` directly: set `context_length: 131072` and `max_tokens: 4096` per thermal safety rule +- Agents using `gemma-4-12b` directly (auxiliary tasks): set `context_length: 131072` +- Compression threshold at 0.65: fires at ~170K for syslog-auto (262K ctx), ~85K for direct qwen/gemma (128K ctx) +- All Hermes clients MUST set `max_tokens: 4096` — first line of defense before server-side `-n 8192` cap +- Port 8080 is used on all 3 GPU hosts (not 8090 as previously documented) diff --git a/hermes-agent-baseline.prose.md b/hermes-agent-baseline.prose.md index fcda08d..95b9eb9 100644 --- a/hermes-agent-baseline.prose.md +++ b/hermes-agent-baseline.prose.md @@ -5,7 +5,7 @@ version: 1.0.0 description: > Canonical known-good baseline for all Syslog Hermes agents. Captures the exact configuration state, keys, workarounds, and audit procedure. When an agent's - configuration goes sideways, restore from this baseline. Last verified 2026-07-05. + configuration goes sideways, restore from this baseline. Last verified 2026-07-08. GPU context reduced to 128K on .8/.110, parallel 2 fleet-wide. author: Abiba (pi agent) --- @@ -22,12 +22,12 @@ done ## Agent Map -| Agent | CT | Node | IP | LiteLLM Key | LiteLLM Alias | -|-------|-----|------|-----|-------------|---------------| -| Tanko | 112 | amdpve | .122 | `sk-CggiHWlamQyShxWC3Hx6uw` | `tanko` | -| Mumuni | 114 | minipve | .123 | `sk-VrqCNlwUgzoNGOpikJ7nwQ` | `mumuni` | -| Tdunna | 111 | amdpve | ? | `sk-6sbCNjz2T6lTVDBdlNHXsA` | `tdunna` | -| Baggy | 113 | amdpve | ? | `sk-krnw_zGBwvvL5b7l2t-s-A` | `baggy` | +| Agent | CT | Node | IP | LiteLLM Key | LiteLLM Alias | Platform | +|-------|-----|------|-----|-------------|---------------|----------| +| Tanko | 112 | amdpve | .122 | `sk-CggiHWlamQyShxWC3Hx6uw` | `tanko` | Hermes | +| Mumuni | 114 | minipve | .123 | `sk-VrqCNlwUgzoNGOpikJ7nwQ` | `mumuni` | Hermes | +| Tdunna | 111 | amdpve | srv1079750 | `sk-Qvzi4uYQBhlSK_XstEhcyQ` | `tdunna` | **pi** | +| Baggy | 113 | amdpve | ? | `sk-krnw_zGBwvvL5b7l2t-s-A` | `baggy` | Hermes | Access: `pct-run ` — no IPs needed. GPU hosts (.8, .110, .15) use SSH. @@ -45,6 +45,8 @@ Agent (systemd) → LITELLM_API_KEY → LiteLLM (:116/v1) → Router (:9000) → ## Config Pattern — Mandatory Fields +### For Hermes Agents (Tanko, Mumuni, Baggy) + Every agent's `/root/.hermes/config.yaml` (or `/home/jerome/.hermes/config.yaml`) MUST have: ### 1. Main Model @@ -154,6 +156,51 @@ pct-run grep -A8 "vision:" /root/.hermes/config.yaml | grep api_key # Must show both api_key: sk-... and api_key_env: LITELLM_API_KEY ``` +### For pi Agents (Tdunna) + +Tdunna (CT111) runs pi 0.80.3 via PM2 with the Zulip extension (router-worker architecture). +Config files: `~/.pi/agent/models.json`, `~/.pi/agent/settings.json`. + +**models.json** — Must only list models authorized for the agent's LiteLLM key: +```json +{ + "providers": { + "syslog-harness": { + "baseUrl": "http://192.168.68.116/v1", + "api": "openai-completions", + "apiKey": "sk-...", + "models": [ + { "id": "syslog-auto" }, + { "id": "ornith-1.0-35b" }, + { "id": "qwen3.6-27B-code" }, + { "id": "gemma-4-12b" } + ] + } + } +} +``` + +**settings.json** — Always use `syslog-auto` as default: +```json +{ + "defaultProvider": "syslog-harness", + "defaultModel": "syslog-auto" +} +``` + +**Validation**: Verify models match LiteLLM's authorized list: +```bash +curl -s http://192.168.68.116:4000/v1/models \ + -H "Authorization: Bearer $(grep apiKey ~/.pi/agent/models.json | head -1 | cut -d'"' -f4)" \ + | jq '.data[].id' +``` + +**Stuck worker detection**: In PM2 logs, `workers=[:busy:N]` with growing N indicates +a stuck worker (model error, no `agent_end` emitted). Fix: correct models.json, delete +stale sessions from `~/.pi/agent/sessions/zulip/`, restart PM2. + +**Service**: `pm2 restart koby-zulip`, health at `:9201/health`. + ## Systemd Pattern For agents where the gateway runs as a user service: @@ -210,5 +257,6 @@ If they differ → ghost detected → kill ghost → start fresh. | Date | Change | |------|--------| +| 2026-07-08 | Tdunna: fixed model mismatch (qwen3.6-35B-A3B→syslog-auto), added pi-specific config section. Key updated to sk-Qvzi4uYQBhlSK_XstEhcyQ. Added Failure Mode #11 to zulip-adapter-lessons. | | 2026-07-06 | Port conflict detection added to all 3 GPU wrappers. Consolidated health check script deployed. Zulip streaming edit_message enabled for Tanko/Mumuni. | | 2026-07-05 | Baseline created. All 4 agents audited, master key removed, api_key workaround applied | diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index 25aef78..2edf9ce 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -5,7 +5,8 @@ description: > Standard Hermes configuration template for Syslog Solution LLC agents. Enforces shared infrastructure setup (Firecrawl, SearXNG, local models, RA-H OS MCP) while keeping agent-specific API keys and model choices. - Updated 2026-06-30: new LiteLLM keys, nginx routing, router back online. + Updated 2026-07-08: GPU context reduced to 128K on NVIDIA (.8, .110), + parallel 2 on all GPUs, LiteLLM timeouts tuned, context_length guidance added. --- ## Maintains @@ -92,7 +93,8 @@ model: base_url: http://192.168.68.116/v1 api_key_env: LITELLM_API_KEY # Set in /etc/environment AND ~/.hermes/.env max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation - context_length: 262144 # Must match model's max context window + context_length: 262144 # For syslog-auto (ornith route supports 256K). + # Set 131072 if using qwen3.6-27B-code or gemma-4-12b directly. fallback_providers: provider: deepseek @@ -123,8 +125,9 @@ compression: enabled: true model: gemma-4-12b # ⚠️ Must match auxiliary.compression.model provider: harness - max_context_window: 262144 # Must match model's actual capacity - threshold: 0.65 # Fires at ~170K for 262K window (not 0.25!) + max_context_window: 262144 # For syslog-auto (ornith supports 256K). + # Set 131072 if using qwen or gemma directly. + threshold: 0.65 # Fires at ~170K for 262K window, ~85K for 128K target_ratio: 0.30 protect_last_n: 40 hygiene_hard_message_limit: 350 @@ -236,6 +239,28 @@ The following MUST be identical across ALL profiles: - `max_context_window: 262144` MUST match the model's actual capacity - See `devops-hermes-compression` skill for full reference +### Rule 9: Default Model Must Be `syslog-auto` (All Agents) +- **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto` +- **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json` +- `syslog-auto` is the LiteLLM routing model — it load-balances between ornith-1.0-35b + and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against: + - Model name typos that cause 403 errors and silent worker failures + - Single GPU downtime (routing falls back automatically) + - Key/model authorization mismatches +- **Exception**: Sub-agent profiles (Mumuni's 6 profiles) may specify explicit models + for specialized tasks, but MUST validate those models exist in the key's authorized list + +### Rule 10: Validate Model IDs Before Deployment (pi Agents) +- After configuring a pi agent's `models.json`, verify every model ID: + ```bash + curl -s http://192.168.68.116:4000/v1/models \ + -H "Authorization: Bearer " | jq '.data[].id' + ``` +- All model IDs in `models.json` MUST appear in the LiteLLM response +- The agent's API key may have a SUBSET of the full model catalog — check per-key +- A non-existent model ID causes 403 errors that silently break the pi RPC worker + (no `agent_end` emitted, worker stays "busy", Zulip messages pile up unprocessed) + ## Execution 1. **Check current config** — Read the target agent's config.yaml diff --git a/infrastructure-update.prose.md b/infrastructure-update.prose.md index dcaff33..156b84c 100644 --- a/infrastructure-update.prose.md +++ b/infrastructure-update.prose.md @@ -64,7 +64,7 @@ Before ANY update wave: **Verify after Wave 2:** - All VMs/CTs running: check via Proxmox API - LiteLLM healthy: `curl localhost:4000/health/liveliness` (via CT 116) -- GPU servers responding: check :8080 on VM 101, VM 103; check ornith via router +- GPU servers responding: check :8080 on VM 101, VM 103; check ornith via router (http://192.168.68.116/health/unified — .15:8080 is firewalled to .116 only) - Zulip agents connected: check Mumuni/Tanko gateway state - Abiba PM2 processes online: `pm2 status` @@ -120,6 +120,8 @@ Before Wave 1, snapshot these files: /opt/home_stack/docker-compose.yml (VM 109 .7) /opt/audiobookshelf/docker-compose.yml (VM 109 .7) /root/.pi/agent/extensions/config.yaml (CT 100 .24) +/etc/systemd/system/ornith-server.service (amdpve .15) +/etc/systemd/system/llama-server.service (VM 101 .8, VM 103 .110) ``` Run: `mkdir -p /tmp/infra-update-backup-$(date +%Y%m%d) && rsync -av ...` diff --git a/zulip-adapter-lessons.prose.md b/zulip-adapter-lessons.prose.md index 1e0a874..7d4e72b 100644 --- a/zulip-adapter-lessons.prose.md +++ b/zulip-adapter-lessons.prose.md @@ -73,6 +73,25 @@ description: > - **Fix**: If edit fails, send the response as a new message instead - **Applies to**: pi extension (fixed), Hermes adapter (verify edit fallback exists) +### 11. Model-ID Mismatch Causes Silent Worker Failure (pi-specific) +- **Symptom**: User sends DM, sees placeholder, but never gets response. Worker stays + "busy" forever, accumulating pending replies (10+). Health endpoint shows green but + `messages_processed` stalls. +- **Root cause**: Agent's model config (`models.json` or `settings.json`) references a + model ID that doesn't exist in LiteLLM's authorized model list. Example: `qwen3.6-35B-A3B` + configured but LiteLLM only exposes `ornith-1.0-35b` under that key. pi's session + workers emit 403 on first prompt, then never recover because the error doesn't trigger + `agent_end` — worker stays `busy` and all subsequent messages pile up in the steer queue. +- **Detection**: Compare `~/.pi/agent/models.json` model IDs against `curl -H "Authorization: Bearer " http://192.168.68.116/v1/models` output. A stuck worker shows + `workers=[:busy:N]` with growing N in extension logs. +- **Fix**: (1) Update `models.json` to only include models from the authorized list. + (2) Set `defaultModel` to `syslog-auto` (safe routing model). (3) Delete stale session + JSONL files from `~/.pi/agent/sessions/zulip/`. (4) Restart PM2 process. +- **Prevention**: Use `syslog-auto` as default model for all agents — it handles model + routing and fallback automatically. Direct model IDs (`ornith-1.0-35b`, etc.) should + only be used when explicitly requested. Validate model IDs at agent setup time. +- **Applies to**: pi extension (Tdunna CT111, fixed 2026-07-08), any agent using `syslog-harness` provider + ## Deployment Checklist When deploying a new Zulip adapter, verify: @@ -86,3 +105,4 @@ When deploying a new Zulip adapter, verify: - [ ] `@all-bots` detected via configurable user_id - [ ] Poll uses long-poll (not `dont_block=true` polling) - [ ] Stuck/idle detection accounts for quiet periods +- [ ] Model IDs in config validated against `GET /v1/models` with actual API key