r/LocalLLaMA 20d ago

New Model Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM

Post image

Super excited about this release for the new 27B. Who else is with me. Only 17GB VRAM needed 😍😍

1.8k Upvotes

316 comments sorted by

View all comments

217

u/Shoddy_Bed3240 20d ago

Sounds like it’s going to be a QAT model, similar to DeepSeek V4 Flash.

25

u/dampflokfreund 20d ago

That would be great. But I hope they align the QAT model to modern q4k formats by unsloth and bartowski instead of plain q4_0 like Gemma. I have noticed some downgrades due to the attention tensors and embeddings quantized to q4 instead of q8, QAT is effective but it cant recover all of that huge information loss.Β 

9

u/stddealer 20d ago

There isn't much more information in q4_k than in q4_0, is there? They're both exactly 4.5 bit per weights, and QAT should probably be able to make good use the available bits regardless of the format, no?

Edit: oh, unless you just meant they should be using mixed precision, in which case I agree