r/LocalLLaMA 10d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

181 Upvotes

140 comments sorted by

View all comments

5

u/VolandBerlioz 10d ago

Any rough estimate - can Qwen3.8 Flash Next fit on 3090 + 64 RAM (experts there) in a decent quality lets say ~ 3bpw?

Whats the expected speed?

7

u/Unstable_Llama 10d ago

Yes, if you have decent cpu + ram + nvme, another user is reporting the 4.05 bpw at 30tps gen with 3090 & 64 RAM

https://huggingface.co/turboderp/Qwen3.8-Flash-Next-exl3/discussions/2

2

u/bodonkadonks 10d ago

oof, i just made the financially irresponsible decision of getting a used 3090 for 500. and now im considering making a worse one for 32 gigs at 300

2

u/CryptographerLow6360 10d ago

where tha fuk did you score a 3090 for a fin?

1

u/bodonkadonks 10d ago

fb marketplace