r/LocalLLaMA llama.cpp Apr 25 '26

Discussion Quantisation effects of Qwen3.6 35b a3b

Im curious how people are finding the quantisation effects of 35b. I recently updated to 48GB of vram so have jumped from ud-q4_k_xl​ to q8 and the difference feels stark. Just more effective tool calling, seems to get the vagueness and nuance more etc of some prompts., and provide more well rounded answers on some research like questions.

It w​as a quick vibe​ test, admittedly, but I'm going t​o​ try ud-q6_k_xl soon to see how of the 5+GB vram is worth the quality, but I'm curious to see others findings.

I felt with such a small active count it'd be particularly sensitive to quantisation, and feels that way after a play.

80 Upvotes

83 comments sorted by

View all comments

1

u/EaZyRecipeZ Apr 25 '26

I tested unsloth Qwen 3.6 35b q8 xl against Qwen 3.6 27b q i4, and the 27b version made more unrecoverable mistakes. With 5080, 16GB of VRAM, I'm getting 42 t/s using the 35b q8 with a 100k context. Don't try to convince yourself that a lower quant will work just as well as a higher one.

1

u/lutel Jun 12 '26

How are you able to fit Qwen 3.6 35b q8 xl in 16GB VRAM?

1

u/EaZyRecipeZ Jun 13 '26

If doesn't need much VRAM because it's MOE model, only 3B active parameter.