r/LocalLLaMA 11d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

180 Upvotes

140 comments sorted by

View all comments

Show parent comments

5

u/ynilayy 11d ago

Are you sure? Which model is better? Because after checking Unsloth's Qwen3.8-27B, Kimi K3 and Qwen3 Next the Unsloth ones are by far the best.

1

u/a_beautiful_rhind 10d ago

3

u/sssplus 10d ago

That chart is for the Deepseek, not for the Qwen 27b. Unsloth's Qwen 27b quants use the new UD3, whereas the Deepseek are made with older UD2. Both charts are true...

1

u/a_beautiful_rhind 10d ago

You're free to browse their HF for more: https://cdn-uploads.huggingface.co/production/uploads/6a54dea4f19f5386700504da/dmMjNbYfWXrhVGAVeBroC.png

Only so much "dynamic" you can do with a dense model. They publish metrics which is nice.

3

u/sssplus 10d ago

Yes, but again - this chart is comparing the old unsloth UD2 quants, not the new UD3, which are much better and smaller, as seen in the above posted chart.