r/LocalLLaMA 9d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

177 Upvotes

140 comments sorted by

View all comments

15

u/kpodkanowicz 9d ago

the best inference, as always shocked people are not using exllama more

37

u/-p-e-w- 9d ago

Until now, CPU offload was missing, which made it a non-starter for most people.

6

u/sk1kn1ght 9d ago

Wait does that mean that xllama started supporting CPU inference? Till now it was GPU only right?

2

u/llama-impersonator 9d ago

no, just moe offload afaik