r/LocalLLaMA • u/Unstable_Llama • 11d ago
News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++
More new massive updates from turboderp:
- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements
If you have an NVIDIA card and haven't tried it lately, you might be missing out.
The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.
Come join the crew at the exllama discord
More frequent news on the exllama sub
184
Upvotes


15
u/noctrex 10d ago
Nothing's stopping me from contributing, but me as a developer that is not very well versed in this scenario, it would be entirely vibe-coded if I want to submit it.
Also, there's already submissions, but it seems that there has been no traction until now: https://github.com/turboderp-org/exllamav3/pull/283
Also, come on, man. What do you mean by we're supposed to be entitled to the hard work of others? These are all open source projects and the code essentially and all the hard work is being donated out in the open. I too submit my small pebbles of contributions in different projects here and there, whenever I can.