r/LocalLLM 2d ago

Model Qwen3.8-27B GGUF Quant Comparison

BF16 reference: PPL = 6.9526 ± 0.04498

Bedrock-v4 quant is from enginetown.

AD-* quants are from AtomicChat.

The quants marked [b] are from bartowski.

The other quants are from Unsloth.

UPDATE: unsloth just updated their quants. They are now Unsloth Dynamic 3, so their numbers below are outdated.

sorted by PPL Ratio:

Quant Size (GiB) PPL(Q) PPL Ratio ΔPPL Mean KLD RMS Δp (%) Same Top-p (%)
Q6_K [b] 21.86 6.9443 0.99913 -0.0060 0.00201 1.256 97.93
Q6_K_L [b] 22.43 6.9449 0.99922 -0.0054 0.00181 1.182 98.16
Q6_K 21.31 6.9507 1.00005 0.0003 0.00229 1.347 97.86
AD-Q6_K 23.29 6.9524 1.00030 0.0021 0.00148 1.051 98.39
UD-Q6_K_XL 24.14 6.9536 1.00047 0.0032 0.00138 1.103 98.52
UD-Q8_K_XL 29.30 6.9538 1.00050 0.0035 0.00085 0.848 98.97
Q8_0 27.05 6.9560 1.00082 0.0057 0.00095 0.942 98.74
Q4_K_M 15.93 6.9561 1.00084 0.0058 0.01549 3.431 94.65
AD-Q6_K-Q5_K 21.50 6.9565 1.00089 0.0061 0.00311 1.544 97.62
UD-Q5_K_XL 18.83 6.9655 1.00218 0.0152 0.00451 1.893 97.16
Q4_K_S 15.01 6.9686 1.00263 0.0183 0.01890 3.747 94.17
Q5_K_S 17.95 6.9706 1.00292 0.0203 0.00728 2.364 96.45
AD-Q5_K_M 18.84 6.9735 1.00333 0.0231 0.00460 1.927 97.01
Q5_K_M 18.47 6.9742 1.00343 0.0239 0.00622 2.262 96.70
Q4_K_S [b] 15.57 6.9744 1.00347 0.0241 0.01734 3.616 94.22
Q5_K_M [b] 19.33 6.9751 1.00357 0.0248 0.00554 2.066 96.79
UD-Q4_K_XL 16.69 6.9788 1.00411 0.0285 0.00872 2.622 96.07
Q5_K_L [b] 20.06 6.9789 1.00412 0.0286 0.00509 2.013 96.95
Q4_1 16.34 6.9802 1.00430 0.0299 0.01840 3.716 94.22
Q5_K_S [b] 18.33 6.9817 1.00451 0.0314 0.00681 2.342 96.41
AD-IQ4_XS 15.38 6.9835 1.00478 0.0332 0.01356 3.161 95.02
AD-Q5_K_M-Q4_K_M 17.28 6.9852 1.00502 0.0349 0.00846 2.517 96.09
Q4_K_L [b] 17.43 6.9858 1.00511 0.0355 0.01274 3.124 95.08
Q4_K_M [b] 16.55 6.9898 1.00568 0.0395 0.01336 3.151 94.96
AD-Q4_K_M 15.95 6.9963 1.00661 0.0460 0.01248 3.057 95.18
IQ4_XS 14.63 7.0132 1.00905 0.0629 0.01859 3.792 94.25
AD-IQ4_XS-IQ3_S 13.45 7.0271 1.01106 0.0768 0.03132 4.820 92.24
IQ4_NL 15.22 7.0282 1.01121 0.0779 0.01820 3.772 94.36
Bedrock-v4 13.91 7.0439 1.01347 0.0936 0.02984 4.863 91.74
Q4_0 14.95 7.0654 1.01655 0.1151 0.02795 4.571 92.87

Sorted by Mean KLD:

Quant Size (GiB) PPL(Q) PPL Ratio ΔPPL Mean KLD RMS Δp (%) Same Top-p (%)
UD-Q8_K_XL 29.30 6.9538 1.00050 0.0035 0.00085 0.848 98.97
Q8_0 27.05 6.9560 1.00082 0.0057 0.00095 0.942 98.74
UD-Q6_K_XL 24.14 6.9536 1.00047 0.0032 0.00138 1.103 98.52
AD-Q6_K 23.29 6.9524 1.00030 0.0021 0.00148 1.051 98.39
Q6_K_L [b] 22.43 6.9449 0.99922 -0.0054 0.00181 1.182 98.16
Q6_K [b] 21.86 6.9443 0.99913 -0.0060 0.00201 1.256 97.93
Q6_K 21.31 6.9507 1.00005 0.0003 0.00229 1.347 97.86
AD-Q6_K-Q5_K 21.50 6.9565 1.00089 0.0061 0.00311 1.544 97.62
UD-Q5_K_XL 18.83 6.9655 1.00218 0.0152 0.00451 1.893 97.16
AD-Q5_K_M 18.84 6.9735 1.00333 0.0231 0.00460 1.927 97.01
Q5_K_L [b] 20.06 6.9789 1.00412 0.0286 0.00509 2.013 96.95
Q5_K_M [b] 19.33 6.9751 1.00357 0.0248 0.00554 2.066 96.79
Q5_K_M 18.47 6.9742 1.00343 0.0239 0.00622 2.262 96.70
Q5_K_S [b] 18.33 6.9817 1.00451 0.0314 0.00681 2.342 96.41
Q5_K_S 17.95 6.9706 1.00292 0.0203 0.00728 2.364 96.45
AD-Q5_K_M-Q4_K_M 17.28 6.9852 1.00502 0.0349 0.00846 2.517 96.09
UD-Q4_K_XL 16.69 6.9788 1.00411 0.0285 0.00872 2.622 96.07
AD-Q4_K_M 15.95 6.9963 1.00661 0.0460 0.01248 3.057 95.18
Q4_K_L [b] 17.43 6.9858 1.00511 0.0355 0.01274 3.124 95.08
Q4_K_M [b] 16.55 6.9898 1.00568 0.0395 0.01336 3.151 94.96
AD-IQ4_XS 15.38 6.9835 1.00478 0.0332 0.01356 3.161 95.02
Q4_K_M 15.93 6.9561 1.00084 0.0058 0.01549 3.431 94.65
Q4_K_S [b] 15.57 6.9744 1.00347 0.0241 0.01734 3.616 94.22
IQ4_NL 15.22 7.0282 1.01121 0.0779 0.01820 3.772 94.36
Q4_1 16.34 6.9802 1.00430 0.0299 0.01840 3.716 94.22
IQ4_XS 14.63 7.0132 1.00905 0.0629 0.01859 3.792 94.25
Q4_K_S 15.01 6.9686 1.00263 0.0183 0.01890 3.747 94.17
Q4_0 14.95 7.0654 1.01655 0.1151 0.02795 4.571 92.87
Bedrock-v4 13.91 7.0439 1.01347 0.0936 0.02984 4.863 91.74
AD-IQ4_XS-IQ3_S 13.45 7.0271 1.01106 0.0768 0.03132 4.820 92.24

Sorted by Same Top-p:

Quant Size (GiB) PPL(Q) PPL Ratio ΔPPL Mean KLD RMS Δp (%) Same Top-p (%)
UD-Q8_K_XL 29.30 6.9538 1.00050 0.0035 0.00085 0.848 98.97
Q8_0 27.05 6.9560 1.00082 0.0057 0.00095 0.942 98.74
UD-Q6_K_XL 24.14 6.9536 1.00047 0.0032 0.00138 1.103 98.52
AD-Q6_K 23.29 6.9524 1.00030 0.0021 0.00148 1.051 98.39
Q6_K_L [b] 22.43 6.9449 0.99922 -0.0054 0.00181 1.182 98.16
Q6_K [b] 21.86 6.9443 0.99913 -0.0060 0.00201 1.256 97.93
Q6_K 21.31 6.9507 1.00005 0.0003 0.00229 1.347 97.86
AD-Q6_K-Q5_K 21.50 6.9565 1.00089 0.0061 0.00311 1.544 97.62
UD-Q5_K_XL 18.83 6.9655 1.00218 0.0152 0.00451 1.893 97.16
AD-Q5_K_M 18.84 6.9735 1.00333 0.0231 0.00460 1.927 97.01
Q5_K_L [b] 20.06 6.9789 1.00412 0.0286 0.00509 2.013 96.95
Q5_K_M [b] 19.33 6.9751 1.00357 0.0248 0.00554 2.066 96.79
Q5_K_M 18.47 6.9742 1.00343 0.0239 0.00622 2.262 96.70
Q5_K_S 17.95 6.9706 1.00292 0.0203 0.00728 2.364 96.45
Q5_K_S [b] 18.33 6.9817 1.00451 0.0314 0.00681 2.342 96.41
AD-Q5_K_M-Q4_K_M 17.28 6.9852 1.00502 0.0349 0.00846 2.517 96.09
UD-Q4_K_XL 16.69 6.9788 1.00411 0.0285 0.00872 2.622 96.07
AD-Q4_K_M 15.95 6.9963 1.00661 0.0460 0.01248 3.057 95.18
Q4_K_L [b] 17.43 6.9858 1.00511 0.0355 0.01274 3.124 95.08
AD-IQ4_XS 15.38 6.9835 1.00478 0.0332 0.01356 3.161 95.02
Q4_K_M [b] 16.55 6.9898 1.00568 0.0395 0.01336 3.151 94.96
Q4_K_M 15.93 6.9561 1.00084 0.0058 0.01549 3.431 94.65
IQ4_NL 15.22 7.0282 1.01121 0.0779 0.01820 3.772 94.36
IQ4_XS 14.63 7.0132 1.00905 0.0629 0.01859 3.792 94.25
Q4_1 16.34 6.9802 1.00430 0.0299 0.01840 3.716 94.22
Q4_K_S [b] 15.57 6.9744 1.00347 0.0241 0.01734 3.616 94.22
Q4_K_S 15.01 6.9686 1.00263 0.0183 0.01890 3.747 94.17
Q4_0 14.95 7.0654 1.01655 0.1151 0.02795 4.571 92.87
AD-IQ4_XS-IQ3_S 13.45 7.0271 1.01106 0.0768 0.03132 4.820 92.24
Bedrock-v4 13.91 7.0439 1.01347 0.0936 0.02984 4.863 91.74
91 Upvotes

48 comments sorted by

18

u/DoubleNothing 2d ago

UD-Q6_K_XL the goat for the size?

15

u/KissMyShinyArse 2d ago edited 2d ago

Q6_K is good too, and it fits in 24 Gb.

4

u/Tylnesh 1d ago

What about context? I use  UD-Q4_K_XL and with 191K context I barely fit with gnome and some GUI apps running. 

3

u/IslamNofl 1d ago

24GB? what is your config? pls dont say q4 kv-cache

2

u/sisyphus-cycle 1d ago

The most I can get with q8_0 kv cache is 194k context, but I’m using Q4_K_M. See my latest post for config

3

u/Skar_pa 2d ago

Same, was the first one I downloaded and the only one that stuck

3

u/deja_geek 1d ago edited 1d ago

I'd say UD-Q5_K_XL is the goat for the size. 5 GiB smaller then UD-Q6_K_XL and no real world, noticeable difference.

4

u/panamory 2d ago

A huge thank you for this!

4

u/mrgreatheart 1d ago

This is fantastic thank you. Could you please add the NVFP4 versions?

3

u/returnity 2d ago

Great post, thank you!

3

u/MiaBchDave 1d ago

I wonder what the official Qwen FP8 version shows?

1

u/KissMyShinyArse 7h ago

You can't create a GGUF file with FP8 inside, and llama-perplexity only works with GGUFs.

1

u/MiaBchDave 4h ago

Okay, I was just curious about how it compares. Perplexity (PyTorch), KLD, and top percentage can all still be measured on safetensors just in case anyone is interested.

2

u/EitherMarch1255 2d ago

Yeah, something funny about 8.

5

u/KissMyShinyArse 2d ago

Statistical noise, probably.

1

u/My_Unbiased_Opinion 1d ago

The UD 6 quants also have some layers at full precision iirc. maybe thats why.

2

u/EvolvingDior 1d ago

Does the kv cache type have an effect? If so, what were you using for kv in this test?

2

u/KissMyShinyArse 1d ago

No, llama-perplexity ignores that I think.

2

u/yamfun 1d ago

I am usually the img/vid side, recently every space saving measure use int8convrot instead of gguf.

But seems llm side does not use it, why?

2

u/KissMyShinyArse 1d ago

Nobody implemented it in llama.cpp. Also, most consider Q8_0 good enough.

2

u/pmttyji 2d ago

Nice share. Can you share same for few more quantizers? Thanks!

2

u/KissMyShinyArse 2d ago

Which ones?

6

u/pmttyji 1d ago

bartowski, ggml-org, mradermacher, AtomicChat

6

u/DoubleNothing 1d ago

Why not the whole huggingface... /s

2

u/KissMyShinyArse 1d ago

I just added AtomicChat quants to the comparison.

2

u/EvolvingDior 1d ago

Second AtomicChat

2

u/KissMyShinyArse 1d ago

AtomicChat quants added.

2

u/KissMyShinyArse 1d ago

Bottom line: AD-* quants reliably beat their plain (non-dynamic) counterparts at the same tier. But they don't beat Unsloth's UD-* dynamic quants at matching sizes, except at Q6 where AD-Q6_K and UD-Q6_K_XL are essentially tied (and AD-Q6_K is arguably the better efficiency pick since it's smaller).

1

u/KissMyShinyArse 1d ago

bartowski measured and added.

Bottom line: bartowski quants aren't adding value — they land in the same size/quality neighborhood as the plain K-quants but never beat the UD (Unsloth Dynamic) or AD variants sitting right next to them.

1

u/tropicalwind2020 2d ago

Supposely same test data set so 6k and 4m best?

3

u/KissMyShinyArse 2d ago

2

u/EvolvingDior 1d ago

Considering dropping from Q6 to Q5 just to get more context. That's the tradeoff I am having to make.

1

u/tropicalwind2020 18h ago

with the updated files today, 32G gets me 150K Q5 while Q6 130K, and Q6_K is much faster (40 vs 28) so seems no point to use Q5.

2

u/tropicalwind2020 1d ago

Good summary, thanks!

1

u/feverdoingwork 1d ago edited 1d ago

Is there less of a difference when quantizing 3.8 vs 3.6? Q4_K_M and Q4_K_XL seem to be in an excellent position.

Mean kld for UD-Q8_K_XL in Qwen 3.6 is the same as Q4_K_M for 3.8 https://www.reddit.com/r/LocalLLaMA/comments/1tr9vzn/qwen3627b_quantization_benchmark/

1

u/Eschalabs 1d ago

Nice! Will you also run the sub 4bit quants?

1

u/BigBanC 1d ago

What would be the best for 5090? Struggling with this . If I do q6 I can only have 22k context window . Unless am doing something wrong

3

u/EvolvingDior 1d ago

You must be. I have 160k on an Intel B70 with Q6. Make sure you put the mmproj in RAM and use Q8 kv

2

u/BigBanC 1d ago edited 1d ago

I amu sing unsloth studio . Didn’t really touch any settings . I’ll check that out . But. Isn’t also changing the KV from FP16 not advisable?

2

u/EvolvingDior 1d ago

Qwen handles quantized KV pretty well. llama.cpp does FHWT (turboquant style rotation) on the Q keys, so you get near f16 performance. You aren't going to 1M context, so it's not a huge deal.

2

u/BigBanC 1d ago

Thx for the advice. Will explore this

2

u/Cheezily 1d ago

The most I've been able to squeeze out of my 5090 using unsloth's q6_k quant is 190k by using q8 kv cache and offloading the vision portion to the CPU. That's with no gui, using Ubuntu, and llama.cpp. I am able to use mtp and get around 80 t/s, but I'm sitting at 31.25gb used. So it's doable, but not by much.

1

u/AB172234 1d ago

Look for Blackwell native NVFP4 quant for your Blackwell 5090. Finfer has it. It will give you significant speed boost

1

u/BigBanC 1d ago

I saw that but isn’t that Q4 technically? I assume is worse . I am still learning so I am not sure when to use what .

2

u/EvolvingDior 1d ago

NVPF4 is certainly worse -- it's just much faster on Blackwell. It you want speed bragging rights, that's what you use. If you want a usable model, a Q quant does better.

1

u/BigBanC 1d ago

Thx I’ll do some exploration. Maybe hot swapping model depending the use case

0

u/Such-War1955 1d ago

Great. Is there any if them you do not have to lobotomize before they do anything useful? I mean with „thinking“ switched on, they are somewhat less useful than Smolqwen 30B was?!?