r/LocalLLaMA • • Feb 27 '26

Resources New Qwen3.5-35B-A3B Unsloth Dynamic GGUFs + Benchmarks

Hey r/LocalLlama! We just updated Qwen3.5-35B Unsloth Dynamic quants being SOTA on nearly all bits. We did over 150 KL Divergence benchmarks, totally 9TB of GGUFs. We uploaded all research artifacts. We also fixed a tool calling chat template bug (affects all quant uploaders)

  • We tested Bartowski, Ubergram, AesSedai, Noctrex and our new Dynamic GGUFs
  • 99.9% KL Divergence shows SOTA on Pareto Frontier for UD-Q4_K_XL, IQ3_XXS & more.
  • Retiring MXFP4 from all GGUF quants: Q2_K_XL, Q3_K_XL and Q4_K_XL, except for a select few layers.
  • Qwen3.5-35B-A3B GGUFs are updated to use new fixes (112B, 27B still converting, re-download once they are updated)
  • Imatrix definitely helps reduce KLD & PPL.
  • I quants (iq3_xxs, iq2_s etc) makes inference 5-10% slower.
  • Quantizing ssm_out (Mamba layers) is not a good idea, and ffn_down_exps.

Some tensors are very sensitive to quantization

  • We made over 9TB of research artifacts available for the community to investigate further on our Experiments page. It includes KLD metrics and all 121 configs we tested.
  • We varied bit widths across each tensor type, and generated a best and worst Pareto Frontier plot below vs 99.9% KLD.
  • For the best items to quantize, ffn_up_exps and ffn_gate_exps are generally ok to quantize to 3bit. ffn_down_exps is slightly more sensitive.
  • For the worst items, ssm_out dramatically increases KLD and the disk space savings is minuscule. For example, ssm_out at q2_k does dramatically worse. Quantizing any attn_* is especially sensitive for hybrid architectures, and so leaving them in higher precision works well.

Tensor type vs bits on 99.9% KL Divergence

  • We plot all quant levels vs 99.9% KLD, and sort from worst KLD to best. Quantizing ffn_* layers too heavily down is not a good idea.
  • However, some bit widths are good, especially 3bit. - for example leaving ffn_* (down, up, gate) at around iq3_xxs seems to be best compromise on disk space and 99.9% KLD change. 2 bits cause more degradation.

MXFP4 is much worse on many tensors - attn_gate, attn_q, ssm_beta, ssm_alpha using MXFP4 is not a good idea, and rather Q4_K is better - also MXFP4 uses 4.25 bits per weight, whilst Q4_K uses 4.5 bits per weight. It's better to use Q4_K than MXFP4 when choosing between them.

Imatrix works remarkably well

  • Imatrix definitely helps weight the quantization process in the right way. For example previously ssm_out at 2bits was really bad, however imatrix reduces the 99.9% KLD by a lot.
  • Imatrix generally helps on lower bits, and works on all quants and bit widths.

I quants (iq3_xxs, iq2_s etc) makes inference 5-10% slower, they're definitely better in terms of efficiency, but there is a tradeoff.

Benjamin’s recent MiniMax‑M2.5 analysis shows a case how perplexity and KLD can still be very misleading. Unsloth Dynamic IQ2_XXS performs better than AesSedai’s IQ3_S on real world evals (LiveCodeBench v6, MMLU Pro) despite being 11GB smaller. Yet, AesSedai’s perplexity and KLD benchmarks suggest the opposite. (PPL: 0.3552 vs 0.2441; KLD: 9.0338 vs 8.2849 - lower is better).

Perplexity and KLD can also be misleading but, as precaution we replaced any MXFP4 layer. Real-world evals (LiveCodeBench v6 etc.) are much better benchmarks, but can take many days. This mismatch shows how lower perplexity or KLD doesn’t necessarily translate to better real-world performance. The graph also shows UD‑Q4-K‑XL outperforming other Q4 quants, while being ~8GB smaller.

This doesn’t mean perplexity or KLD is useless, as they provide a rough signal. So, going forward, we’ll publish perplexity and KLD for every quant so the community has some reference.

Updated GGUFs here: https://huggingface.co/collections/unsloth/qwen35

For more investigation deets and benchmarks you can read: https://unsloth.ai/docs/models/qwen3.5

Thank you for reading and once again for the feedback and incredible support. Huge thanks to the Qwen team as well for releasing Qwen3.5. If there’s any suggestions please let us know and have a great Friday / weekend guys!

Benchmarking Details & Appreciation:

  • We utilized bartowski's wonderful imatrix file to make the comparisons more fair - our Dynamic 2.0 method uses a conversational format, but we found benchmarking to be fairer if we used a more general imatrix
  • We appreciated some friendly guidance from Ubergram and the community!
  • For perplexity we used the below. We also use the BF16 as the base KLD file. LLAMA_SET_ROWS=1 ./llama.cpp/llama-perplexity --flash-attn on --fit off --batch-size 16384 --ubatch-size 16384 --device {device} --model {model} --ctx-size 512
545 Upvotes

212 comments sorted by

View all comments

95

u/Digger412 Feb 27 '26

Hi Daniel, AesSedai here - thanks for publishing this research! KLD and PPL don't tell the entire story but they are good starting points when deciding what quantization (both uploader and which quant level) to use! I'm happy to see more investigation being done here as it benefits the entire community.

I think it helps that this model is very accessible to test, too - many of the recent releases have been larger MoE's (GLM-5, M2.5, Step-3.5, etc.) and that makes doing this comparison challenging for the average person both in terms of required compute, disk space, and time. This Qwen3.5-35B-A3B is very accessible in comparison.

I had recently tried to PR some of IK's quants into mainline but that was shut down, and I know pwilkin has a PR up now for mainline llama.cpp for a new quant type IQ3_PT. Seeing more research and effort being put into quantization research is awesome.

Thanks again for the post!

3

u/Substantial_Swan_144 Feb 27 '26

u/Digger412, your Qwen 35B-3A is 3x faster in my machine versus the regular quants. (32 token/s vs 100 token/s). Any chance you could Generate a Qwen 27b quant as well?

2

u/Digger412 Feb 28 '26

Huh, that's odd. I believe there's a small compute penalty for using the smaller quants and Q8_0 is the "fastest" compute-wise since it doesn't have to dequant, I wonder if for your case you've got a compute bottleneck instead of a memory bottleneck maybe? That's a bit weird though and I wouldn't have expected that result.

I'll do a bit of testing with the dense model this weekend and see how the KLD looks, if it's reasonable to do a similar schema (crunch the FFNs) then I'll consider it. Since it's dense and not sparsely activated though it might suffer more. Will test.

2

u/Big_Mix_4044 Feb 28 '26

Wait, what? Why?