From 80811511e7859c228cb6f891289fd17ffee29e06 Mon Sep 17 00:00:00 2001 From: mumuni-bot Date: Sat, 12 Sep 2026 21:03:48 +0000 Subject: [PATCH 1/7] fix(harness): align repo with verified live state; retire model-version names from the client surface litellm_config.yaml - client-visible model_list reduced to capability names: syslog-auto, gpu-dense, strix-moe, gpu-vision - retired qwen3.6-27B-code, qwen3.8-27B-uncensored, qwen3.6-35B-udq4 (and the already-dead gemma-4-12b) - fallbacks and model_cost re-keyed to the surviving names - upstream litellm_params.model set to the capability names; the backends ignore the model field (verified HTTP 200 on all three hosts), so this needs no llama.cpp relaunch and loses no KV warmth - verified live: /v1/models advertises exactly the four names, each returns 200 through nginx gpu_roster.yaml - launch args replaced with the VERIFIED live command lines, read from the running processes via the PVE guest agent (acerpve VM 101 = .8, ocupve VM 103 = .110) and on amdpve (.15) - .8: model_path corrected to Qwen3.8-27B-Uncensored-Q4_K_M.gguf, max_concurrent 2 -> 1, context 262144 -> 131072, full arg list recorded (incl. --parallel 1 and the new --slot-save-path / --metrics added 2026-09-12) - .15: model_path corrected to Carnice-Qwen3.6-MoE-35B-A3B-Q4_K_M.gguf, full arg list recorded - keys renamed to the capability names; hosts.current_model aligned README.md / dashboard - README dense entry corrected (1 slot, 131K ctx, actual model file) - dashboard picker now uses capability names; fixed the "Gemma 4 12B" / "12B VLM" labels (the RTX 5070 serves a 9B Qwen3.5) and the gpu-vision -> gpu-light id mapping scripts/ - added the three operational monitors as tracked files (they were untracked): gpu-monitor.py, gpu-self-heal.py, litellm-health-check.sh - cleared their references to retired model names, which were causing failed calls every benchmark cycle (150 failed gemma-4-12b calls in the last 7 days); gpu-monitor.py's .110 entry also wrongly listed .8's model Intentionally NOT changed - LITELLM-MIGRATION-PLAN.md: historical planning document (June 14), not a live-state claim - backups/, graphify-out/, litellm_config.yaml.backup: historical artifacts - unrelated untracked files (router.py, docker-compose.yml.pre-1991-20260911, nginx/default.conf, dashboard/gpu-monitor.html, scripts/gitea-logger.sh): out of scope for this change --- README.md | 4 +- dashboard/dashboard.html | 2 +- dashboard/dashboard.py | 26 +- gpu_roster.yaml | 26 +- litellm_config.yaml | 44 +-- scripts/gpu-monitor.py | 391 ++++++++++++++++++++++++++ scripts/gpu-self-heal.py | 471 ++++++++++++++++++++++++++++++++ scripts/litellm-health-check.sh | 359 ++++++++++++++++++++++++ 8 files changed, 1259 insertions(+), 64 deletions(-) create mode 100644 scripts/gpu-monitor.py create mode 100755 scripts/gpu-self-heal.py create mode 100755 scripts/litellm-health-check.sh diff --git a/README.md b/README.md index 5e48169..e61d6be 100644 --- a/README.md +++ b/README.md @@ -7,7 +7,7 @@ CT 116 Docker stack for routing local GPU models through a unified OpenAI-compat ``` nginx :80 → router :9000 → GPU backends ├─ qwen3.6-35B-A3B (MoE) @ 192.168.68.15:8080 [2 slots, 262K ctx] - ├─ qwen3.6-27B-code (Dense) @ 192.168.68.8:8080 [2 slots, 262K ctx] + ├─ gpu-dense (Qwen3.8-27B-U) @ 192.168.68.8:8080 [1 slot, 131K ctx] └─ gpu-vision (VLM) @ 192.168.68.110:8080 [2 slots, 262K ctx] Total: 6 concurrent slots @@ -62,7 +62,7 @@ When all GPUs are saturated, requests enter a polling queue (500ms intervals) in | GPU | Model | VRAM | Slots | Context | Best For | |-----|-------|------|-------| | Strix Halo | qwen3.6-35B-A3B (MoE) | 65GB | 2 | 262K | General quality | -| RTX 3090 | qwen3.6-27B-code (Dense) | 24GB | 2 | 262K | Code, reasoning | +| RTX 3090 | gpu-dense (Qwen3.8-27B-Uncensored) | 24GB | 1 | 131K | Dense, general | | RTX 5070 | gpu-vision (VLM) | 12GB | 2 | 262K | Speed, vision | ## Maintenance diff --git a/dashboard/dashboard.html b/dashboard/dashboard.html index fdee4c7..c4bc16d 100644 --- a/dashboard/dashboard.html +++ b/dashboard/dashboard.html @@ -70,7 +70,7 @@ body{background:var(--bg);color:var(--text);font-family:-apple-system,BlinkMacSy