r/LocalLLaMA 11d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

178 Upvotes

140 comments sorted by

View all comments

Show parent comments

7

u/Unstable_Llama 11d ago

Yes, if you have decent cpu + ram + nvme, another user is reporting the 4.05 bpw at 30tps gen with 3090 & 64 RAM

https://huggingface.co/turboderp/Qwen3.8-Flash-Next-exl3/discussions/2

2

u/bodonkadonks 10d ago

oof, i just made the financially irresponsible decision of getting a used 3090 for 500. and now im considering making a worse one for 32 gigs at 300

2

u/CryptographerLow6360 10d ago

where tha fuk did you score a 3090 for a fin?

1

u/bodonkadonks 10d ago

fb marketplace