From f99f7e1e34e243fb3520f37e6c9802c2858e704d Mon Sep 17 00:00:00 2001 From: root Date: Sat, 12 Sep 2026 21:56:24 +0000 Subject: [PATCH] fix: restore per-host probe coverage + sweep residual retired names MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A. litellm-health step 7: restore .8 → gpu-dense probe (was duplicated to gpu-vision after 8210fd9). Add note documenting monitor key scope gap for gpu-dense. B. Sweep remaining retired names presented as usable: - README.md:91: qwen3.6-27B-code → gpu-dense in runnable example - gpu-fleet.prose.md:14: qwen3.6-35B-udq4 → Carnice-Qwen3.6-MoE... - infrastructure-control.prose.md:226: qwen3.6-35B-udq4 → strix-moe - proxmox-monitor.prose.md:90: qwen3.6-35B-udq4 → strix-moe C. Audit test: 10/10 passed (retired raw names now hard-fail) Fix-forward from 8210fd9 (direct master push). --- README.md | 2 +- gpu-fleet.prose.md | 2 +- infrastructure-control.prose.md | 2 +- litellm-health.prose.md | 6 +++++- proxmox-monitor.prose.md | 2 +- 5 files changed, 9 insertions(+), 5 deletions(-) diff --git a/README.md b/README.md index c3860c2..c850817 100644 --- a/README.md +++ b/README.md @@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4 # Configure an agent with a different auxiliary model -prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision +prose run hermes-config-template agent_name=syslog-code default_model=gpu-dense auxiliary_model=gpu-vision ``` ### Option B: Manual Execution diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 490649d..83c97c1 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -11,7 +11,7 @@ description: > Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster). UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability. - Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias). + Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context). Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. For larger context needs → fall back to external providers (deepseek). VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index 5c40667..9734a1c 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -223,7 +223,7 @@ description: > **Prometheus targets**: - 192.168.68.8:9400 (RTX 3090 — qwen) - 192.168.68.110:9400 (RTX 5070 — gpu-vision) -- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4) +- 192.168.68.15:9400 (Strix Halo — strix-moe) - harness-litellm:4000 (LiteLLM health) ### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud) diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 2061843..386e718 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -144,12 +144,16 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- 7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health check runs on the **backend edge**, not the public edge, so these paths carry the `/litellm/` prefix: - - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) + - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8) - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) - POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) - Auth uses the dedicated `monitor` agent key, read on CT 116 from `/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference — the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). + - NOTE: the monitor key is scoped for gpu-vision, strix-moe, syslog-auto but NOT gpu-dense. + The gpu-dense probe above requires a key with gpu-dense access; if unavailable, the probe + should be run with a different key or the contract should note the gap rather than silently + dropping host coverage. - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to this list. The RTX 5070 host now serves `gpu-vision`. - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so diff --git a/proxmox-monitor.prose.md b/proxmox-monitor.prose.md index 39e81d6..a7904e1 100644 --- a/proxmox-monitor.prose.md +++ b/proxmox-monitor.prose.md @@ -87,7 +87,7 @@ agent: abiba | storepve | 192.168.68.6 | PVE | | acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) | | minipve | 192.168.68.12 | PVE | -| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) | +| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) | ## Operations