fix: swap gpu-dense to SmartCode-Fable-5 (ThinkingCap + Fable 5 CoT)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 0s

gpu-dense on RTX 3090 swapped from qwen3.6-27B-code to SmartCode-Fable-5-CoT-Reasoning-QKVO-Qwen-3.6-27B-Distilled (UD-Q4_K_XL, ~17.9GB).\n\nImprovements:\n- ~50% fewer thinking tokens via ThinkingCap finetune\n- Fable 5 CoT distillation for improved coding reasoning\n- Same 27B base, fits RTX 3090 at ~73% VRAM\n- Recommended samplers: temp 0.9, top-p 0.95, top-k 60\n- Context: 128K (fleet standard)
This commit is contained in:
root
2026-07-27 16:43:29 +00:00
parent fb45007ace
commit 179529de71
2 changed files with 14 additions and 12 deletions
+4 -4
View File
@@ -13,10 +13,10 @@ description: >
against the live system. Policy fields are authoritative. See the
`verify-before-mutate` skill.
**Last verified:** 2026-07-26CT 114 (mumuni) destroyed. Mumuni moved inside
CT 100 (abiba) — Zulip connected on .24. Verification protocol run: CT 118 set
to static IP .20, llama-server on .110 restored. All PVE nodes, CT hostnames,
and public endpoints confirmed. See data/learnings.md.
**Last verified:** 2026-07-27gpu-dense swapped to SmartCode-Fable-5
(Qwen3.6-27B distilled, ~50% fewer thinking tokens, improved coding reasoning).
Model pricing reduced 100x across all models ($0.15/$0.60 per 1M tokens).
All LiteLLM models set to 128K max_model_tokens. See data/learnings.md.
---
# Infrastructure Control Pattern