fix: correct quant to UD-Q3_K_XL (Q4 too large for 24GB VRAM) #34

Merged
mumuni-bot merged 1 commits from fix/gpu-dense-q3-correction into master 2026-07-27 20:26:14 +00:00
Owner

SmartCode-Fable-5 UD-Q4_K_XL (17GB) did not fit on RTX 3090 with 128K context + KV cache. Switched to UD-Q3_K_XL (14.7GB) — fits at ~22.4GB/24.6GB (91%). Updated VRAM, config, and status.

SmartCode-Fable-5 UD-Q4_K_XL (17GB) did not fit on RTX 3090 with 128K context + KV cache. Switched to UD-Q3_K_XL (14.7GB) — fits at ~22.4GB/24.6GB (91%). Updated VRAM, config, and status.
mumuni-bot added 1 commit 2026-07-27 20:26:05 +00:00
fix: correct quant to UD-Q3_K_XL (Q4 too large for 24GB VRAM)
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Failing after 5s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Skipped
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Skipped
22b015e182
mumuni-bot merged commit 835647e241 into master 2026-07-27 20:26:14 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SyslogSolution/prose-contracts#34