fix: restore per-host probe coverage + sweep residual retired names
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 1s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s

A. litellm-health step 7: restore .8 → gpu-dense probe (was duplicated
   to gpu-vision after 8210fd9). Add note documenting monitor key scope
   gap for gpu-dense.
B. Sweep remaining retired names presented as usable:
   - README.md:91: qwen3.6-27B-code → gpu-dense in runnable example
   - gpu-fleet.prose.md:14: qwen3.6-35B-udq4 → Carnice-Qwen3.6-MoE...
   - infrastructure-control.prose.md:226: qwen3.6-35B-udq4 → strix-moe
   - proxmox-monitor.prose.md:90: qwen3.6-35B-udq4 → strix-moe
C. Audit test: 10/10 passed (retired raw names now hard-fail)

Fix-forward from 8210fd9 (direct master push).
This commit is contained in:
root
2026-09-12 21:56:24 +00:00
parent 8210fd905c
commit f99f7e1e34
5 changed files with 9 additions and 5 deletions
+1 -1
View File
@@ -88,7 +88,7 @@ prose run memory-audit-maintenance memory_threshold=90 verify_configs=true
prose run hermes-config-template agent_name=syslog-devops default_model=claude-sonnet-4
# Configure an agent with a different auxiliary model
prose run hermes-config-template agent_name=syslog-code default_model=qwen3.6-27B-code auxiliary_model=gpu-vision
prose run hermes-config-template agent_name=syslog-code default_model=gpu-dense auxiliary_model=gpu-vision
```
### Option B: Manual Execution
+1 -1
View File
@@ -11,7 +11,7 @@ description: >
Strix Halo: strix-moe → unsloth/Qwen3.6-35B-A3B-MTP (UD-Q4_K_M, 22GB).
RTX 5070: gpu-vision — IQ4_NL + MTP draft (~122 tok/s, 2x faster).
UPDATED 2026-07-17: Context reduced fleet-wide from 256K to 128K for stability.
Strix Halo model swapped to qwen3.6-35B-udq4 (22GB, strix-moe alias).
Strix Halo model: Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf (alias strix-moe, 256K context).
Instability observed near 100K at 256K (now all GPUs at 128K). 128K is the stable ceiling.
For larger context needs → fall back to external providers (deepseek).
VRAM headroom improved: RTX 3090 ~70%, RTX 5070 ~65%.
+1 -1
View File
@@ -223,7 +223,7 @@ description: >
**Prometheus targets**:
- 192.168.68.8:9400 (RTX 3090 — qwen)
- 192.168.68.110:9400 (RTX 5070 — gpu-vision)
- 192.168.68.15:9400 (Strix Halo — qwen3.6-35B-udq4)
- 192.168.68.15:9400 (Strix Halo — strix-moe)
- harness-litellm:4000 (LiteLLM health)
### Ecosystem C: Netbird (72.61.0.17 — Hostinger srv1079750.hstgr.cloud)
+5 -1
View File
@@ -144,12 +144,16 @@ contracts — read them there. Do not re-add retired names (`gemma-4-12b`, `gpu-
7. **Check model inference via LiteLLM** — Test one model on each GPU host. The health
check runs on the **backend edge**, not the public edge, so these paths carry the
`/litellm/` prefix:
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-dense → expect 200 (RTX 3090, .8)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=gpu-vision → expect 200 (RTX 5070, .110)
- POST http://{{backend_host}}/litellm/v1/chat/completions model=strix-moe → expect 200 (Strix Halo, .15)
- Auth uses the dedicated `monitor` agent key, read on CT 116 from
`/etc/litellm-monitor.env` (root-only 0600). Do NOT use the master key for inference —
the master key is for admin endpoints only (`/key/list`, `/key/generate`, `/key/info`).
- NOTE: the monitor key is scoped for gpu-vision, strix-moe, syslog-auto but NOT gpu-dense.
The gpu-dense probe above requires a key with gpu-dense access; if unavailable, the probe
should be run with a different key or the contract should note the gap rather than silently
dropping host coverage.
- `gemma-4-12b` was retired and returns 400 `Invalid model name` — do not re-add it to
this list. The RTX 5070 host now serves `gpu-vision`.
- `/v1/models` is **key-scoped**: a model is only visible to keys allowed to use it, so
+1 -1
View File
@@ -87,7 +87,7 @@ agent: abiba
| storepve | 192.168.68.6 | PVE |
| acerpve | 192.168.68.9 | PVE (hosts llm-gpu qemu/101) |
| minipve | 192.168.68.12 | PVE |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (qwen3.6-35B-udq4, strix-moe) |
| amdpve | 192.168.68.15 | PVE + Strix Halo LLM (strix-moe) |
## Operations