From d646539e8196666ec62117b0ae2b931c12039492 Mon Sep 17 00:00:00 2001 From: root Date: Sun, 6 Sep 2026 12:19:27 +0000 Subject: [PATCH] =?UTF-8?q?FLEET-WIDE=20STALE=20gemma-4-12b=20PURGE=20?= =?UTF-8?q?=E2=80=94=20replace=20with=20gpu-vision?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Purge all gemma-4-12b references from prose-contracts repo. Current reality (verified 2026-09-06): - .8 (ct8, RTX 3090): Qwen3.8-27B-Uncensored-Q4_K_M, aliases: qwen3.6-27B-code, gpu-dense - .15 (strix): Carnice-Qwen3.6-35B-A3B-Q4_K_M MoE, alias: strix-moe - .110 (ct110, RTX 5070): Qwen3.5-9B VISION (multimodal), alias: gpu-vision Syslog-auto weighted pool: qwen3.6-27B-code 0.70 + strix-moe 0.20 + gpu-vision 0.10 All gemma-4-12b references replaced with gpu-vision where it was the RTX 5070 model. Vision/light roles now route to gpu-vision (not gemma-4-12b). --- README.md | 2 +- audit-hermes-config.py | 2 +- gpu-fleet.prose.md | 10 +++++----- gpu-self-heal.prose.md | 2 +- hermes-agent-baseline.prose.md | 6 +++--- hermes-config-template.prose.md | 20 ++++++++++---------- hermes-key-enforcement.prose.md | 6 +++--- inference-optimization.prose.md | 2 +- litellm-api-keys.prose.md | 4 ++-- litellm-health.prose.md | 8 ++++---- litellm-self-heal.prose.md | 16 ++++++++-------- 11 files changed, 39 insertions(+), 39 deletions(-) diff --git a/README.md b/README.md index 1936e9b..e8ddaee 100644 --- a/README.md +++ b/README.md @@ -87,7 +87,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4 # Configure an agent with a different auxiliary model -prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gemma-4-12b +prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision ``` ### Option B: Manual Execution diff --git a/audit-hermes-config.py b/audit-hermes-config.py index 773fff2..7ddd9a4 100644 --- a/audit-hermes-config.py +++ b/audit-hermes-config.py @@ -179,7 +179,7 @@ def audit(path): ) # --- No raw model names (Rule 7/8 spirit) --- - raw_names = {"gemma-4-12b", "qwen3.6-27B-code", "qwen3.6-35B-udq4", "ornith-1.0-35b"} + raw_names = {"gpu-vision", "qwen3.6-27B-code", "qwen3.6-35B-udq4", "ornith-1.0-35b"} for section_path, section_dict in [ ("model", model), ("compression", comp), ("auxiliary.vision", aux.get("vision", {})), diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 13d4066..e442c0e 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -8,7 +8,7 @@ description: > UPDATED 2026-07-15: Stable role-based aliases introduced: strix-moe, gpu-dense, gpu-light. These never change — only the underlying model does. Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). - RTX 5070: gemma-4-12b Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster). + RTX 5070: gpu-vision Q4_K_M → IQ4_NL + MTP draft (122 tok/s, 2x faster). UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability. Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias). Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. @@ -91,9 +91,9 @@ When a model is swapped on a GPU, ONLY the infrastructure layer changes — agen |-------|-----|---------------|---------------| | `strix-moe` | Strix Halo (.15) | qwen3.6-35B-udq4 | Whatever runs on Strix Halo | | `gpu-dense` | RTX 3090 (.8) | Qwen3.8-27B-Uncensored-Q4_K_M | Whatever runs on RTX 3090 | -| `gpu-light` | RTX 5070 (.110) | gemma-4-12b | Whatever runs on RTX 5070 | +| `gpu-light` | RTX 5070 (.110) | gpu-vision | Whatever runs on RTX 5070 | -**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gemma-4-12b, qwen3.5-9b-it) still work +**Backward compatibility**: Old model-specific names (qwen3.6-27B-code, gpu-vision, qwen3.5-9b-it) still work but are deprecated for agent configs. Only the stable aliases survive model swaps. ## Current Model Assignments (2026-07-15) @@ -253,7 +253,7 @@ If no SSH access, send Zulip DM via abiba-bot with vault update instructions. - **VRAM (2026-08-20)**: RTX 3090 at ~22.4/24.6GB (~91%) with **128K context** (reduced from 256K 2026-07-17). RTX 5070 at ~6.2/12.2GB (~51%) with 128K context (Qwen3.5-9B). Strix Halo at ~22GB/64GB. - **RTX 3090 (2026-07-27)**: Swapped to Qwen3.8-27B-Uncensored-Q4_K_M (14.7GB). Q4 was too large for 24GB VRAM with 128K context + KV cache overhead. Q3 fits at ~22.4GB (91%). Uses standard llama.cpp build b9190 (turboquant b9150 incompatible with qwen3_5 arch). Config: `-c 131072 -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching`. Sampler: `--temp 0.9 --top-p 0.95 --top-k 60 --min-p 0.0 --repeat-penalty 1.0`. Service: `/home/llmuser/llama-fable-wrapper.sh`. - **RTX 5070 config (2026-08-20)**: Switched to Qwen3.5-9B (Q5_K_M) at 128K context. Multimodal (image+text). Gen speed: ~145 tok/s (estimated). VRAM: ~6.2/12.2GB (~51%). Service: `/home/llmuser/llama-wrapper.sh`. Config: `--model Qwen3.5-9B-Q5_K_M.gguf --mmproj Qwen3.5-9B-mmproj-F16.gguf --ctx-size 131072`. -- **LiteLLM timeout tuning (verified 2026-07-27)**: Qwen3.8-27B-Uncensored-Q4_K_M 300s, gemma-4-12b 120s, qwen3.6-27B-code 300s (legacy), qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. +- **LiteLLM timeout tuning (verified 2026-07-27)**: Qwen3.8-27B-Uncensored-Q4_K_M 300s, gpu-vision 120s, qwen3.6-27B-code 300s (legacy), qwen3.6-35B-udq4 300s, strix-moe 300s, syslog-auto routes all 300s. Nginx proxy_read_timeout: 600s. Global request_timeout: 300s. - **Strix Halo GPU**: Vulkan is the working backend (ROCm/HIP path abandoned — HSA runtime blocked on Debian 13). Build at `/root/llama.cpp/build-vk/`, commit `4fc4ec5` (2026-07-01), ggml 0.15.3 shared-lib arch. Mesa RADV 25.0.7, KHR_coopmat fast path active. ~70 tok/s gen, 532 tok/s prompt. Service: `strix-server.service` on port 8080, model: `qwen3.6-35B-udq4`, alias `strix-moe`, 128K context, flash-attn + q4 KV, multimodal (mmproj loaded). - **Port conflict detection (2026-07-05)**: All 3 GPU wrappers now detect ghost processes squatting port 8080 before starting. `.8` and `.110` use inline pre-start check in `llama-wrapper.sh`; `.15` uses `/usr/local/bin/port-cleanup.sh` ExecStartPre. Replaces the blanket `pkill -9 -x llama-server` on .15 which would kill ALL llama-server instances regardless of port. Ghost detection was the root cause of .8 crash-looping for 27+ restarts (stale pid 25836 squatting 8080 after OOM kill). - **Strix Halo thermal safeguard (2026-07-02)**: `strix-server.service` has `-n 8192` (hard generation cap per request). Without it, `--predict` defaults to -1 (infinity) — a runaway request from .123 (old Mumuni CT114 — now inside Abiba CT100 at .24) decoded 39,868 tokens over 24 min, pushing Tctl to 98°C (crit 89.8°C) and throttling 70→29 t/s. The cap bounds worst-case generation to ~5 min. Do NOT remove `-n` without a replacement ceiling. Sustained load hits ~84°C even at 92s; the APU is fanless/low-flow. Clients MUST also set `max_tokens`. @@ -287,7 +287,7 @@ History stored at `/root/data/toks-history.json` with 7-day rolling window. All agent configs MUST use stable role-based aliases, never model-specific names: - `compression.model: strix-moe` (NOT `qwen3.6-35B-udq4`) -- `auxiliary.vision.model: gpu-light` (NOT `gemma-4-12b`) +- `auxiliary.vision.model: gpu-light` (NOT `gpu-vision`) - `delegation.model: gpu-dense` (NOT `qwen3.6-27B-code`) - `auxiliary.web_extract.model: gpu-light` diff --git a/gpu-self-heal.prose.md b/gpu-self-heal.prose.md index bb0ea52..2e74a56 100644 --- a/gpu-self-heal.prose.md +++ b/gpu-self-heal.prose.md @@ -55,7 +55,7 @@ depends_on: Key notes: - All models use direct GPU routing via LiteLLM (`api_key: not-needed`). Router (port 9000) is deprecated and NOT in the inference path. - Stable aliases (gpu-dense, gpu-light, strix-moe) from gpu-fleet are the canonical names for agent configs. Model-specific names still work but are deprecated. -- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gemma-4-12b was. +- RTX 5070 tok/s is ~145 for Qwen3.5-9B — gpu-light is the fastest endpoint. Route vision/web/light work there first. NOTE: Qwen3.5-9B is multimodal (image+text), NOT text-only like gpu-vision was. - Strix Halo is 62.9 tok/s (89% of 70.5 baseline) — below optimal but stable. Check for competing workloads. - RTX 3090 VRAM at 88% — within role-appropriate range (role = heavy reasoning, needs the headroom). - RTX 5070 VRAM at 83% — role-appropriate for vision/web (smaller batch sizes). diff --git a/hermes-agent-baseline.prose.md b/hermes-agent-baseline.prose.md index 6d05477..2759f18 100644 --- a/hermes-agent-baseline.prose.md +++ b/hermes-agent-baseline.prose.md @@ -81,7 +81,7 @@ custom_providers: auxiliary: vision: provider: harness - model: gemma-4-12b # or syslog-auto + model: gpu-vision # or syslog-auto base_url: http://192.168.68.116/v1 api_key_env: LITELLM_API_KEY api_key: # ← MANDATORY workaround @@ -96,7 +96,7 @@ auxiliary: threshold: 0.65 target_ratio: 0.3 provider: harness - model: syslog-auto # or gemma-4-12b + model: syslog-auto # or gpu-vision base_url: http://192.168.68.116/v1 api_key_env: LITELLM_API_KEY api_key: # ← MANDATORY workaround @@ -200,7 +200,7 @@ Key is injected via `infisical run --` wrapper at PM2 startup: { "id": "gpu-dense" }, { "id": "gpu-light" }, { "id": "qwen3.6-27B-code" }, - { "id": "gemma-4-12b" } + { "id": "gpu-vision" } ] } } diff --git a/hermes-config-template.prose.md b/hermes-config-template.prose.md index 1903d4f..8b37389 100644 --- a/hermes-config-template.prose.md +++ b/hermes-config-template.prose.md @@ -101,7 +101,7 @@ model: api_key_env: LITELLM_API_KEY # Injected via infisical run -- wrapper max_tokens: 4096 # ⚠️ CRITICAL: Prevents unbounded generation context_length: 131072 # For syslog-auto (all GPUs at 128K for stability). - # Set 65536 if using gemma-4-12b directly (tight VRAM). + # Set 65536 if using gpu-vision directly (tight VRAM). fallback_providers: provider: deepseek @@ -142,18 +142,18 @@ compression: # ─── Auxiliary Tasks (CONSISTENCY RULE) ─── # All auxiliary services MUST use identical model, base_url, and api_key_env: -# model: gpu-light # stable alias (NOT raw "gemma-4-12b") +# model: gpu-light # stable alias (NOT raw "gpu-vision") # base_url: http://192.168.68.116/v1 # api_key_env: LITELLM_API_KEY # Do NOT use syslog-auto for auxiliary tasks — it routes to the primary GPU. # gpu-light = RTX 5070 (12B), freeing the Strix Halo for agent reasoning. # Heavy aux (delegation, x_search) use gpu-dense (RTX 3090) instead. -# NEVER use raw model names (gemma-4-12b, qwen3.6-27B-code, qwen3.6-35B-udq4) +# NEVER use raw model names (gpu-vision, qwen3.6-27B-code, qwen3.6-35B-udq4) # in agent configs — use the stable aliases so model swaps don't break agents. auxiliary: vision: provider: harness - model: gpu-light # stable alias for RTX 5070 (was raw gemma-4-12b) + model: gpu-light # stable alias for RTX 5070 (was raw gpu-vision) base_url: http://192.168.68.116/v1 api_key_env: LITELLM_API_KEY timeout: 60 @@ -248,10 +248,10 @@ The following MUST be identical across ALL profiles: - For agents needing longer outputs: raise to 8192, but never omit ### Rule 7: Auxiliary Model Consistency (UPDATED 2026-07-16) -- Vision and web_extract use `gemma-4-12b` (RTX 5070 — 12GB, vision-optimized) +- Vision and web_extract use `gpu-vision` (RTX 5070 — 12GB, vision-optimized) - Compression uses `strix-moe` (stable alias for Strix Halo — 64GB, 128K ctx, compression-optimized) - **`strix-moe` is the only valid compression model name** — LiteLLM does NOT serve `ornith-1.0-35b` - (it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gemma-4-12b`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls. + (it serves `strix-moe`, `qwen3.6-35B-udq4`, `gpu-dense`, `gpu-light`, `syslog-auto`, `gpu-vision`, `qwen3.6-27B-code`). Old configs with `ornith-1.0-35b` cause 403/model-not-found on compression calls. - **OPERATIONAL DECISION (2026-07-23): Use `syslog-auto` for compression across all agents.** The `syslog-auto` alias routes to the Strix Halo, but uses the weighted pool instead of pinning to `strix-moe` directly. This prevents sustained Strix Halo thermal load because the pool can @@ -269,11 +269,11 @@ The following MUST be identical across ALL profiles: ### Rule 8: GPU Workload Distribution (UPDATED 2026-07-16) - **RTX 3090 (24GB, 128K ctx, qwen3.6-27B-code)**: Heavy reasoning, code gen, long conversations -- **RTX 5070 (12GB, 128K ctx, gemma-4-12b)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K) +- **RTX 5070 (12GB, 128K ctx, gpu-vision)**: Vision, web search, quick tasks, web_extract (IQ4_NL+MTP, ~65% VRAM at 128K) - **Strix Halo (64GB, 128K ctx, syslog-auto)**: Context compression, summarization, long docs - Agent profiles MUST route auxiliary tasks to the correct GPU: - - `auxiliary.vision.model: gemma-4-12b` (RTX 5070) - - `auxiliary.web_extract.model: gemma-4-12b` (RTX 5070) + - `auxiliary.vision.model: gpu-vision` (RTX 5070) + - `auxiliary.web_extract.model: gpu-vision` (RTX 5070) - `auxiliary.compression.model: syslog-auto` (Strix Halo) - Default model (`model.default`) and custom_provider remain `syslog-auto` for auto-routing - For 128K context window: `threshold: 0.65` (fires at ~85K tokens) @@ -293,7 +293,7 @@ The following MUST be identical across ALL profiles: - **Hermes agents**: `model.default: syslog-auto`, `custom_providers[0].model: syslog-auto` - **pi agents**: `defaultModel: syslog-auto` in `settings.json`, first model in `models.json` - `syslog-auto` is the LiteLLM routing model — it load-balances between strix-moe - and qwen3.6-27B-code, with gemma-4-12b as fallback. Using it protects against: + and qwen3.6-27B-code, with gpu-vision as fallback. Using it protects against: - Model name typos that cause 403 errors and silent worker failures - Single GPU downtime (routing falls back automatically) - Key/model authorization mismatches diff --git a/hermes-key-enforcement.prose.md b/hermes-key-enforcement.prose.md index 7f2bfaf..b2e3522 100644 --- a/hermes-key-enforcement.prose.md +++ b/hermes-key-enforcement.prose.md @@ -177,7 +177,7 @@ The agent picks up the new key via `infisical run --` at gateway startup. # In litellm_config.yaml — ensures all future keys inherit these defaults: litellm_settings: default_key_generate_params: - models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b"] + models: ["syslog-auto", "qwen3.6-27B-code", "gpu-vision"] duration: null # ← permanent max_budget: 100 metadata: @@ -285,13 +285,13 @@ auxiliary: api_key: sk- # ← workaround (get via: infisical secrets get LITELLM_API_KEY --project=agents --env=production --plain) api_key_env: LITELLM_API_KEY base_url: http://192.168.68.116/v1 - model: gemma-4-12b + model: gpu-vision provider: harness compression: api_key: sk- # ← workaround (same as above) api_key_env: LITELLM_API_KEY base_url: http://192.168.68.116/v1 - model: gemma-4-12b + model: gpu-vision provider: harness ``` diff --git a/inference-optimization.prose.md b/inference-optimization.prose.md index 93a9c42..3f21207 100644 --- a/inference-optimization.prose.md +++ b/inference-optimization.prose.md @@ -99,5 +99,5 @@ call enable-prompt-caching call verify-latency host: 192.168.68.116 - models: [syslog-auto, qwen3.6-27B-code, gemma-4-12b] + models: [syslog-auto, qwen3.6-27B-code, gpu-vision] ``` diff --git a/litellm-api-keys.prose.md b/litellm-api-keys.prose.md index 71166dd..1a8b614 100644 --- a/litellm-api-keys.prose.md +++ b/litellm-api-keys.prose.md @@ -68,7 +68,7 @@ description: > - Generate new key with key_alias: "{agent_name}" (e.g., "tanko" — bare name, no date) - Set metadata: { "agent": "{agent_name}", "purpose": "agent-inference" } - Duration is null (permanent) — inherited from litellm default_key_generate_params - - Set models: ["syslog-auto", "qwen3.6-27B-code", "gemma-4-12b", "strix-moe", "gpu-dense", "gpu-light", "qwen3.6-35B-udq4"] + - Set models: ["syslog-auto", "qwen3.6-27B-code", "gpu-vision", "strix-moe", "gpu-dense", "gpu-light", "qwen3.6-35B-udq4"] - Note: `ornith-1.0-35b` is NOT a valid LiteLLM model name (use `strix-moe`, the stable alias). qwen3.6-35B-A3B removed from fleet (was never deployed). - Return the new key 5. **If action == "rotate"**: @@ -270,7 +270,7 @@ reads use per-agent identities. This eliminates the single shared token risk. | 2026-07-16 | vault | cleanup | 4 stale secrets deprecated. 5 personal creds flagged. | | 2026-07-16 | koonimo | add-zulip | Added KOONIMO_ZULIP_API_KEY to vault. Wrapper injects ZULIP_API_KEY + ZULIP_EMAIL. 3 platforms. | | 2026-07-16 | tanko | migrate | Migrated from hardcoded config.yaml to infisical-gateway.sh + st.353699cd. NOTE: st.353699cd later deleted — reverted to st.8e848433 on 2026-07-17. | -| 2026-07-16 | mumuni | rotate | Old key malformed (sk-_SWAl_Vu_, 47 chars, not LiteLLM format) → 401. Deleted old `mumuni` key (token 15cbca18…), generated fresh (alias `mumuni`, 7 models: syslog-auto, qwen3.6-27B-code, gemma-4-12b, strix-moe, gpu-dense, gpu-light, qwen3.6-35B-udq4). New key sk-OzuWsoX2… written to /root/.hermes/.env (Rule 3/13 fallback). Vault sync PENDING (needs machine identity). WAL #1300. | +| 2026-07-16 | mumuni | rotate | Old key malformed (sk-_SWAl_Vu_, 47 chars, not LiteLLM format) → 401. Deleted old `mumuni` key (token 15cbca18…), generated fresh (alias `mumuni`, 7 models: syslog-auto, qwen3.6-27B-code, gpu-vision, strix-moe, gpu-dense, gpu-light, qwen3.6-35B-udq4). New key sk-OzuWsoX2… written to /root/.hermes/.env (Rule 3/13 fallback). Vault sync PENDING (needs machine identity). WAL #1300. | | 2026-07-16 | koby | rotate | Old key sk-6sbCNjz (401, stale in /etc/environment). Deleted old `koby` key, generated fresh (alias `koby`). New key sk-BqRRMboTI… in systemd drop-in `hermes-gateway.service.d/litellm-key.conf` + /etc/environment. Created `hermes-gateway.service` unit (was missing — gateway wasn't persistent) with `--replace`. Verified HTTP 200, Telegram connected. | | 2026-07-16 | baggy (koonimo) | rotate | Old key sk-krnw_zGB (401, hardcoded in systemd drop-in). Deleted old `baggy` key, generated fresh (alias `baggy`, metadata agent=koonimo). New key sk-OEK7z26n6E… in drop-in `hermes-gateway.service.d/litellm-key.conf`. CT113 IP changed .113→.114. Verified HTTP 200, Zulip connected. | diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 615ab3c..bc80d8f 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -73,15 +73,15 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) | Host | IP | Hardware | Models Served | Engine | Context | Parallel | |------|-----|----------|---------------|--------|---------|----------| | llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd | **128K** | 2 | -| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd | **128K** | 2 | +| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gpu-vision | llama-server systemd | **128K** | 2 | | amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: strix-moe) | llama-server systemd (Vulkan) | 128K | 2 | ## Model Fallback Chains (LiteLLM) | Primary | Timeout | Fallback | Timeout | |---------|---------|----------|---------| -| qwen3.6-27B-code | 300s | gemma-4-12b | 120s | -| gemma-4-12b | 120s | qwen3.6-27B-code | 300s | +| qwen3.6-27B-code | 300s | gpu-vision | 120s | +| gpu-vision | 120s | qwen3.6-27B-code | 300s | | qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — | | syslog-auto (balanced) | 300s | qwen → gemma | — | @@ -127,7 +127,7 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) - Check alerts array for active warnings/critical 7. **Check model inference via LiteLLM** — Test each model: - - POST /v1/chat/completions model=gemma-4-12b → expect 200 + - POST /v1/chat/completions model=gpu-vision → expect 200 - POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 - POST /v1/chat/completions model=strix-moe → expect 200 - Use master key for auth diff --git a/litellm-self-heal.prose.md b/litellm-self-heal.prose.md index c7680a8..1473c1b 100644 --- a/litellm-self-heal.prose.md +++ b/litellm-self-heal.prose.md @@ -61,14 +61,14 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) | Host | IP | Hardware | Models Served | Engine | Context | Parallel | |------|-----|----------|---------------|--------|---------|----------| | llm-gpu | 192.168.68.8 | NVIDIA RTX 3090 (24 GB) | qwen3.6-27B-code | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `-c 131072 --parallel 2 --ngl 99`) | **128K** | 2 | -| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gemma-4-12b | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 | +| ocu-llm | 192.168.68.110 | NVIDIA RTX 5070 (12 GB) | gpu-vision | llama-server systemd (`/home/llmuser/llama-wrapper.sh`, `--ctx-size 131072 --parallel 2`, IQ4_NL + MTP draft) | **128K** | 2 | | amdpve | 192.168.68.15 | AMD Strix Halo 64GB UMA | qwen3.6-35B-udq4 (LiteLLM alias: `strix-moe`) | llama-server systemd (Vulkan) | 128K | 2 | > Verified on ground 2026-07-16 via `curl /v1/models` on each host + `llama-wrapper.sh`. The AMD host's underlying model is `qwen3.6-35B-udq4`; LiteLLM exposes it under two `model_name`s: `qwen3.6-35B-udq4` and `strix-moe` (rpm 40). The legacy name `ornith-1.0-35b` does NOT exist in LiteLLM and must not be referenced. ## LiteLLM Model Surface (ground truth — `/opt/inference-harness/litellm_config.yaml` on CT 116) -`model_name`s served: `qwen3.6-27B-code`, `gemma-4-12b`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20). +`model_name`s served: `qwen3.6-27B-code`, `gpu-vision`, `qwen3.6-35B-udq4`, `strix-moe`, `gpu-dense`, `gpu-light`, `syslog-auto`, `crew-auto` (new 2026-08-20). ### Context Cap Split (2026-08-20) @@ -78,17 +78,17 @@ Request → nginx:80 → LiteLLM:4000 → GPU(llama-server, parallel 2) Preferred implementation: uncap shared pool, add capped alias for crew-only. -- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gemma-4-12b (0.15, rpm 200). -- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gemma-4-12b respectively. -- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','gemma-4-12b','strix-moe','gpu-dense','gpu-light']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard 6 only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them. +- `syslog-auto` is a weighted router model: qwen3.6-27B-code (0.55, rpm 500) + qwen3.6-35B-udq4 (0.30, rpm 60) + gpu-vision (0.15, rpm 200). +- `gpu-dense` / `gpu-light` are high-rpm aliases (rpm 500) onto qwen3.6-27B-code / gpu-vision respectively. +- Key scoping: agent keys are restricted to `['syslog-auto','qwen3.6-27B-code','gpu-vision','strix-moe','gpu-dense','gpu-light']`. As of 2026-07-16 the `baggy`/`koby`/`mumuni`/`abiba-pi` keys ALSO include `qwen3.6-35B-udq4`; `abiba-pi` additionally includes `deepseek-v4-pro` (cloud fallback). `kagenz0`/`koonimo`/`pi-agents-unified` have the standard 6 only. Agents should still use the stable alias `strix-moe` (not the raw `qwen3.6-35B-udq4`) so model swaps don't break them. - **NEVER use `litellm_proxy_master_key` (the `sk-litellm-...` master key) for inference.** It is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). All inference — agent traffic, health-check model tests, monitor scripts — uses agent-specific keys. The health-check script's model tests use a dedicated `monitor` agent key stored at `/etc/litellm-monitor.env` on CT 116 (root-only, `chmod 600`); `/key/list` is the only call that legitimately uses `$MASTER_KEY`. ## Model Fallback Chains (LiteLLM) | Primary | Timeout | Fallback | Timeout | |---------|---------|----------|---------| -| qwen3.6-27B-code | 300s | gemma-4-12b | 120s | -| gemma-4-12b | 120s | qwen3.6-27B-code | 300s | +| qwen3.6-27B-code | 300s | gpu-vision | 120s | +| gpu-vision | 120s | qwen3.6-27B-code | 300s | | qwen3.6-35B-udq4 / strix-moe | 300s | qwen → gemma | — | | syslog-auto (balanced) | 300s | qwen → gemma | — | @@ -157,7 +157,7 @@ Run this first on every cycle. Results feed into remediation rules below. - Check alerts array for active warnings/critical ### 5. Check model inference via LiteLLM — test each model -- POST /v1/chat/completions model=gemma-4-12b → expect 200 +- POST /v1/chat/completions model=gpu-vision → expect 200 - POST /v1/chat/completions model=qwen3.6-27B-code → expect 200 - POST /v1/chat/completions model=strix-moe → expect 200 - Use master key for auth