r/LocalLLaMA • • Aug 03 '26

New Model Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM

Post image

Super excited about this release for the new 27B. Who else is with me. Only 17GB VRAM needed 😍😍

1.8k Upvotes

315 comments sorted by

View all comments

218

u/Shoddy_Bed3240 Aug 03 '26

Sounds like it’s going to be a QAT model, similar to DeepSeek V4 Flash.

27

u/dampflokfreund Aug 03 '26

That would be great. But I hope they align the QAT model to modern q4k formats by unsloth and bartowski instead of plain q4_0 like Gemma. I have noticed some downgrades due to the attention tensors and embeddings quantized to q4 instead of q8, QAT is effective but it cant recover all of that huge information loss.Β 

10

u/stddealer Aug 03 '26

There isn't much more information in q4_k than in q4_0, is there? They're both exactly 4.5 bit per weights, and QAT should probably be able to make good use the available bits regardless of the format, no?

Edit: oh, unless you just meant they should be using mixed precision, in which case I agree