litellm_config.yaml - client-visible model_list reduced to capability names: syslog-auto, gpu-dense, strix-moe, gpu-vision - retired qwen3.6-27B-code, qwen3.8-27B-uncensored, qwen3.6-35B-udq4 (and the already-dead gemma-4-12b) - fallbacks and model_cost re-keyed to the surviving names - upstream litellm_params.model set to the capability names; the backends ignore the model field (verified HTTP 200 on all three hosts), so this needs no llama.cpp relaunch and loses no KV warmth - verified live: /v1/models advertises exactly the four names, each returns 200 through nginx gpu_roster.yaml - launch args replaced with the VERIFIED live command lines, read from the running processes via the PVE guest agent (acerpve VM 101 = .8, ocupve VM 103 = .110) and on amdpve (.15) - .8: model_path corrected to Qwen3.8-27B-Uncensored-Q4_K_M.gguf, max_concurrent 2 -> 1, context 262144 -> 131072, full arg list recorded (incl. --parallel 1 and the new --slot-save-path / --metrics added 2026-09-12) - .15: model_path corrected to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf, full arg list recorded - keys renamed to the capability names; hosts.current_model aligned README.md / dashboard - README dense entry corrected (1 slot, 131K ctx, actual model file) - dashboard picker now uses capability names; fixed the "Gemma 4 12B" / "12B VLM" labels (the RTX 5070 serves a 9B Qwen3.5) and the gpu-vision -> gpu-light id mapping scripts/ - added the three operational monitors as tracked files (they were untracked): gpu-monitor.py, gpu-self-heal.py, litellm-health-check.sh - cleared their references to retired model names, which were causing failed calls every benchmark cycle (150 failed gemma-4-12b calls in the last 7 days); gpu-monitor.py's .110 entry also wrongly listed .8's model Intentionally NOT changed - LITELLM-MIGRATION-PLAN.md: historical planning document (June 14), not a live-state claim - backups/, graphify-out/, litellm_config.yaml.backup: historical artifacts - unrelated untracked files (router.py, docker-compose.yml.pre-1991-20260911, nginx/default.conf, dashboard/gpu-monitor.html, scripts/gitea-logger.sh): out of scope for this change
55 lines
2.4 KiB
YAML
55 lines
2.4 KiB
YAML
models:
|
|
gpu-vision:
|
|
gpu_url: http://192.168.68.110:8080/v1
|
|
gpu_host: 192.168.68.110
|
|
label: Qwen3.5-9B Vision (RTX 5070)
|
|
max_concurrent: 1
|
|
context: 131072
|
|
tiers: [starter, professional, enterprise]
|
|
capabilities: [completion, multimodal]
|
|
model_path: /home/llmuser/models/qwen3.5-9b/Qwen3.5-9B-Q5_K_M.gguf
|
|
args: --mmproj /home/llmuser/models/qwen3.5-9b/mmproj-F16.gguf --ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --batch-size 2048 --ubatch-size 1024 --n-gpu-layers 99 --flash-attn 1 --cont-batching --parallel 1 --image-min-tokens 2048 --image-max-tokens 16384 --alias gpu-light --reasoning off --api-key not-needed --predict 8192 --cache-prompt --port 8080 --host 0.0.0.0
|
|
|
|
gpu-dense:
|
|
gpu_url: http://192.168.68.8:8080/v1
|
|
sidecar_url: http://192.168.68.8:8090
|
|
gpu_host: 192.168.68.8
|
|
label: Qwen3.8-27B-Uncensored (RTX 3090)
|
|
max_concurrent: 1
|
|
context: 131072
|
|
tiers: [professional, enterprise]
|
|
capabilities: [completion]
|
|
model_path: /home/llmuser/models/Qwen3.8-27B-Uncensored-Q4_K_M.gguf
|
|
args: -m /home/llmuser/models/Qwen3.8-27B-Uncensored-Q4_K_M.gguf -c 131072 --host 0.0.0.0 --port 8080 -ngl 99 --alias qwen3.6-27B-code --api-key not-needed -ctk q4_0 -ctv q4_0 --flash-attn on --cont-batching --parallel 1 --slot-save-path /home/llmuser/.llama-slots --metrics --reasoning off --spec-type draft-mtp --spec-draft-n-max 2 -t 8
|
|
|
|
strix-moe:
|
|
gpu_url: http://192.168.68.15:8080/v1
|
|
sidecar_url: http://192.168.68.15:8090
|
|
gpu_host: 192.168.68.15
|
|
label: Carnice Qwen3.6 MoE 35B-A3B (Strix Halo)
|
|
max_concurrent: 1
|
|
context: 262144
|
|
tiers: [professional, enterprise]
|
|
capabilities: [completion, multimodal]
|
|
model_path: /models/Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf
|
|
args: -m /models/Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf --host 0.0.0.0 --port 8080 --alias strix-moe -c 262144 -ngl 99 -fa on --cache-type-k q4_0 --cache-type-v q4_0 --kv-unified --cache-prompt --cont-batching -b 4096 --ubatch-size 1024 --parallel 2 -t 16 --timeout 600 -n 8192 --reasoning off --no-warmup --metrics
|
|
|
|
hosts:
|
|
gpu-light:
|
|
address: 192.168.68.110
|
|
gpu_name: NVIDIA GeForce RTX 5070
|
|
vram_gb: 12
|
|
current_model: gpu-vision
|
|
|
|
gpu-dense:
|
|
address: 192.168.68.8
|
|
gpu_name: NVIDIA GeForce RTX 3090
|
|
vram_gb: 24
|
|
current_model: gpu-dense
|
|
|
|
gpu-moe:
|
|
address: 192.168.68.15
|
|
gpu_name: AMD Strix Halo (iGPU)
|
|
vram_gb: 64
|
|
current_model: strix-moe
|