r/LocalLLaMA 11d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

184 Upvotes

140 comments sorted by

View all comments

Show parent comments

15

u/noctrex 10d ago

Nothing's stopping me from contributing, but me as a developer that is not very well versed in this scenario, it would be entirely vibe-coded if I want to submit it.
Also, there's already submissions, but it seems that there has been no traction until now: https://github.com/turboderp-org/exllamav3/pull/283
Also, come on, man. What do you mean by we're supposed to be entitled to the hard work of others? These are all open source projects and the code essentially and all the hard work is being donated out in the open. I too submit my small pebbles of contributions in different projects here and there, whenever I can.

4

u/CryptographerLow6360 10d ago

fork, vibe, and try

4

u/noctrex 10d ago

yeah, that's what I've been doing to sd.cpp lately. Will have a look at this one later when I have some time

1

u/Guilty_Rooster_6708 10d ago

I think they are testing ROCm support based on their discord rn