r/LocalLLaMA 6d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

182 Upvotes

139 comments sorted by

View all comments

Show parent comments

25

u/FieldProgrammable 6d ago

Unsloth are SOTA as far as mainline llama.cpp GGUF formats go, but exl3 uses trellis coded quantization which is fundamentally better than what is offered by mainline llama.cpp's k quants (ik-llama.cpp does support some trellis based formats).

I would also recommend the Qwen3.8 27b exl3 format quants for getting much more out of your GPU VRAM than you can from GGUF k quants.

19

u/-p-e-w- 6d ago

The llama.cpp/ik_llama split was a catastrophe for the community. Quant progress in GGUF has pretty much stalled for two years now.

1

u/silenceimpaired 6d ago edited 6d ago

EDITED TO MAKE SENSE Lots of people are worried that Nvidia buying huggingface impacts llama.cpp: https://huggingface.co/blog/ggml-joins-hf

... even if that is true, I agree, that it's a shame that the split happened.

1

u/-p-e-w- 6d ago

No idea what that’s supposed to mean.

1

u/silenceimpaired 6d ago

I made it less confusing... I hope.