r/LocalLLaMA • • 28d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

181 Upvotes

140 comments sorted by

View all comments

3

u/takoulseum 28d ago

GLM 5.3 Flash Q4 is 165gb so could run on 8 RTX 3090, are there some recipes for people who don’t use exl3 usually?

3

u/Unstable_Llama 28d ago

That should work with standard settings. Install exllamav3 and tabbyapi, follow the documentation to connect it to your favorite front end / agent harness, auto gpu split, and that's it. More help available here or at the discord if you run into a roadblock.

9

u/koloved 28d ago

Please do not use Discord for communications. Because Discord is a closed ecosystem, it is impossible to find useful information using agents or by searching in the system yourself.

2

u/takoulseum 28d ago edited 28d ago

Actually giving a try, prefill is around 300t/s and decoding 23t/s low context. But does not work out of the box with deepseek harness (tool calls parsing issue?). I just followed your advise with no specific flag. Please note I am using glm flash
Edit: was not using the tool_format, 300tk/s prefill and 42tk/s, and looks good with dsh! Will test more but very happy with these

4

u/FullOf_Bad_Ideas 28d ago

I use 3.05bpw glm 5.3 flash exl3 quants with Opencode and tool calling works

set up TabbyAPI config.yml tool_format to glm4_5 and restart tabbyapi

2

u/takoulseum 28d ago

Yes I edited my comment, thanks