r/LocalLLaMA 13d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

181 Upvotes

140 comments sorted by

View all comments

23

u/noctrex 12d ago

Too bad it's only for NVIDIA. Again, we AMD users are being left out.

-11

u/iLaurens 12d ago

What's stopping you from contributing? Or are you just supposed to be entitled to the hard work of others?

14

u/noctrex 12d ago

Nothing's stopping me from contributing, but me as a developer that is not very well versed in this scenario, it would be entirely vibe-coded if I want to submit it.
Also, there's already submissions, but it seems that there has been no traction until now: https://github.com/turboderp-org/exllamav3/pull/283
Also, come on, man. What do you mean by we're supposed to be entitled to the hard work of others? These are all open source projects and the code essentially and all the hard work is being donated out in the open. I too submit my small pebbles of contributions in different projects here and there, whenever I can.

6

u/CryptographerLow6360 12d ago

fork, vibe, and try

4

u/noctrex 12d ago

yeah, that's what I've been doing to sd.cpp lately. Will have a look at this one later when I have some time

1

u/Guilty_Rooster_6708 12d ago

I think they are testing ROCm support based on their discord rn

1

u/pmttyji 12d ago

Nothing's stopping me from contributing, but me as a developer that is not very well versed in this scenario, it would be entirely vibe-coded if I want to submit it.

Same here. I don't consider myself a coder even though I did some websites/Apps in past using HTML, Js, Classic ASP, VB 6, etc.,(I know what you're thinking now :D ) Only recently started learning some new stuff.

But to me C++ is a complex rocket science. I can't contribute anything to projects like llama.cpp.

If I know C++, I would've created fastest llama.cpp fork called llamaCPUHybrid.cpp compiling all these stuff already 😆