diff --git a/README.md b/README.md index c3860c2..c850817 100644 --- a/README.md +++ b/README.md @@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4 # Configure an agent with a different auxiliary model -prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision +prose run hermes-config-template agent_name=syslog-code default_model=gpu-dense auxiliary_model=gpu-vision ``` ### Option B: Manual Execution diff --git a/gpu-fleet.prose.md b/gpu-fleet.prose.md index 490649d..83c97c1 100644 --- a/gpu-fleet.prose.md +++ b/gpu-fleet.prose.md @@ -11,7 +11,7 @@ description: > Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB). RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster). UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability. - Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias). + Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context). Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling. For larger context needs → fall back to external providers (deepseek). VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%. diff --git a/infrastructure-control.prose.md b/infrastructure-control.prose.md index 5c40667..9734a1c 100644 --- a/infrastructure-control.prose.md +++ b/infrastructure-control.prose.md @@ -223,7 +223,7 @@ description: > **Prometheus targets**: - 192.168.68.8:9400 (RTX 3090 — qwen) - 192.168.68.110:9400 (RTX 5070 — gpu-vision) -- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4) +- 192.168.68.15:9400 (Strix Halo — strix-moe) - harness-litellm:4000 (LiteLLM health) ### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud) diff --git a/litellm-health.prose.md b/litellm-health.prose.md index 2061843..386e718 100644 --- a/litellm-health.prose.md +++ b/litellm-health.prose.md @@ -144,12 +144,16 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu- 7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health check runs on the **backend edge**, not the public edge, so these paths carry the `/litellm/` prefix: - - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) + - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8) - POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110) - POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15) - Auth uses the dedicated `monitor` agent key, read on CT 116 from `/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference — the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`). + - NOTE: the monitor key is scoped for gpu-vision, strix-moe, syslog-auto but NOT gpu-dense. + The gpu-dense probe above requires a key with gpu-dense access; if unavailable, the probe + should be run with a different key or the contract should note the gap rather than silently + dropping host coverage. - `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to this list. The RTX 5070 host now serves `gpu-vision`. - `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so diff --git a/proxmox-monitor.prose.md b/proxmox-monitor.prose.md index 39e81d6..a7904e1 100644 --- a/proxmox-monitor.prose.md +++ b/proxmox-monitor.prose.md @@ -87,7 +87,7 @@ agent: abiba | storepve | 192.168.68.6 | PVE | | acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) | | minipve | 192.168.68.12 | PVE | -| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) | +| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) | ## Operations