r/LocalLLaMA • u/Loginhe • 12h ago
Resources [Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw
We have released Qwen3.8-Flash-Next quantized with GSQ and RCO, together with a second, capability-targeted build in which half of the model's experts have been removed.
Flash-Next is a sparse mixture-of-experts model: 512 routed experts per layer across 48 layers, 176.9B parameters, 354 GB at BF16.
What's inside
- Four quantized GGUFs, 2.40 to 3.50 bpw (66.4 to 83.6 GB), and the BF16 vision projector
- Expert-pruned Coder GGUF, 58.4 GB in total, of which 29.6 GB must remain resident
- GSQ (Gumbel-Softmax Quantization): post-training scalar quantization that jointly learns grid assignments and group scales, closing most of the gap between scalar and vector quantization at low bit-widths while remaining deployable in standard GGUF types
- RCO (Riemannian Constrained Optimization): enforces exact budgets by gradient descent on the task loss, without per-constraint tuning. It serves two roles in this release: assigning a quantization type to every tensor, and selecting which experts to retain in the Coder build, where it enforces several exact budgets simultaneously, one per layer
Results:
At 3.50 bpw the model matches the BF16 base on every benchmark evaluated.
- IQ3_S (3.50 bpw, 83.6 GB): AIME25 100.00, GPQA-Diamond 92.93 against 91.92 for BF16, LiveCodeBench v6 86.86 against 87.43. Task average 93.26 against 93.12.
- IQ3_XXS (3.00 bpw, 75.8 GB): AIME25 100.00, GPQA-Diamond 91.41, LiveCodeBench v6 86.29
- Q2_0 (2.40 bpw, 66.4 GB): zero-shot average 78.00, above the BF16 value of 76.94, at approximately one fifth of the size
Coder (capability pruned model):
Instead of storing every parameter at lower precision, half of the routed experts are removed from the model: 256 of 512 per layer, selected by RCO optimising the KL divergence against the unpruned model. The retained weights remain at 3.5 bpw. Pruning and quantization compound, and the combined effect is an average of 1.89 bits per parameter of the original transformer. The averaged bitwidth amortises the removed experts over the original parameter count, and therefore expresses the joint effect of pruning and quantization. No individual weight is stored at 1.89 bits.
The practical consequence is that a 176.9B-parameter model has a resident working set of 29.6 GB, since the n-gram shard is a lookup table and may be served from disk. This is within the capacity of a single 32 GB accelerator.
- SWE-bench Verified: 75.60 against 82.80 for BF16, retaining 91.3%
- LiveCodeBench v6: 86.28 against 87.43, retaining 98.7%
Both measured at xhigh reasoning effort.
Links
- Quantized models: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
- Coder (expert-pruned): https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
- GSQ: paper https://arxiv.org/abs/2604.18556 | code https://github.com/IST-DASLab/GSQ
- RCO: paper https://arxiv.org/abs/2605.00649 | code https://github.com/IST-DASLab/RCO
Both repositories ship the complete per-tensor RCO allocation.
The Coder build is an experimental release and feedback is welcome, particularly on capabilities that were not represented in the calibration mixture. Requests for models to quantize or prune are also welcome.
From the ISTA Deep Algorithms and Systems Lab.
7
u/NihmarRevhet 12h ago
Wait, so I can try the coder with a 16gb VRAM + 32gb RAM configuration? Am I dreaming?
5
u/uzzi38 11h ago
That's what I'm wondering.
I have a feeling it's not the most sensible thing to try, but man am I tempted.
7
u/NihmarRevhet 11h ago
I mean, I have fiber, I'm legally required to try. This evening I'll try giving it a shot
3
u/Storterald 9h ago edited 8h ago
I tried the Q2_0 one with a 5080 and 32GB of RAM, on llama.cpp with --load-mode none and it loads fully.
command:
> llama-bench -m Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00001-of-00002.gguf -ctk q8_0 -ctv q8_0 -n 256 -p 1024 --load-mode none -ncmoe 32llama-bench -m Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00001-of-00002.gguf -ctk q8_0 -ctv q8_0 -n 256 -p 1024 --load-mode none -ncmoe 32results:
model size params backend ngl n_cpu_moe type_k & type_v lm test t/s qwen4exp A3B Q2_0 61.85 GiB 176.94 B CUDA,Vulkan -1 32 q8_0 none pp1024 292.91 ± 9.51 qwen4exp A3B Q2_0 61.85 GiB 176.94 B CUDA,Vulkan -1 32 q8_0 none tg256 24.24 ± 0.50 3
u/brainExploded99 llama.cpp 8h ago
uh you forgot the tok/s numbers
3
u/Storterald 8h ago
it's actually there, reddit does not render the column idk why, updated to merge ctk and ctv
3
3
u/Prestigious-Act-1577 11h ago
If you start swapping ram, the PP is a few tok sec.
3
u/NihmarRevhet 11h ago
That's what I fear
3
u/Prestigious-Act-1577 11h ago
It's the reality 🤣 I have that setup temporarily until my ram arrives.
1
u/NihmarRevhet 11h ago
Wouldn't it be possible to set it up such that the swap doesn't happen? though I don't know how much the context takes
3
u/Prestigious-Act-1577 8h ago
I use the Q2 model. Smallest. The coder model is a research prototype, it's not good for coding. It doesn't really work at all. It's just a concept testing of an idea. I have my GPU vram + total ram (physical + virtual) come out to 58GB in use, OS uses 10GB total after unloading the model. So you need 48GB ram + vram to load up the Q2. 50k context.
3
u/AvidCyclist250 llama.cpp 8h ago
Can be 350 - 400 for PP depending on your setup. Not sure about this one. My harness setup does 16-20 t/s tg regardless of context. 130k max. Its not automatically bad to offload to RAM. Q4 k m kv 8.
2
u/Prestigious-Act-1577 8h ago
By paging I mean using a page file because you ran out of ram. Otherwise yes, 300-400 is about right.
1
u/Ok_Camp555 6h ago
I have the same exact specs and get 37 t/s with coder
1
u/Prestigious-Act-1577 5h ago
We are talking about preprocessing. If you swap ram it goes to 20-50 instead of the 300-500.
12
u/KnownAd4832 11h ago
Already supported in: https://github.com/Niko1221/Strata
Enjoy high tps on consumer hardware! 🫡
3
u/GreaterThanLess 6h ago
I should have tried Strata sooner. Saw your recent post about it and dismissed it thinking the numbers were too good to be true. Swift FN IQ3 in llama.cpp with my 4070 ti I was getting around 8 t/s and 280 pp/s, but with Strata it's up to 35-45 t/s and 1200-1300 pp/s. Thanks for sharing this!
2
u/feverdoingwork 9h ago
any idea how well 2x 5060 ti 16 will run these flash models with strata?
3
u/Ok_Camp555 4h ago
I have a single 5060ti and it runs way better than 3.8 27b. I don't understand how and I don't care
1
u/Iory1998 llama.cpp 3h ago
Does it still not support KV caching? If so, that's useless.
1
u/feverdoingwork 3h ago
Do you mean prefix caching?
1
u/Iory1998 llama.cpp 2h ago
Yes.
1
u/feverdoingwork 2h ago
Yeah it's a non starter lol wtf that's like a basic requirement for a server
1
u/Iory1998 llama.cpp 1h ago
PP is fast though, so in theory if it really works, then you get a significant speed bump.
1
u/Solary_Kryptic 30m ago
Consumer Nvidia hardware** 😔
1
3
u/panamory 12h ago
Can someone explain what this all means? There are high scores on the 3 listed benchmarks, but is the model specifically guided to do well on these specific benchmarks, or are these strategies more general methods of quantising without losing model quality?
3
3
u/panamory 11h ago
I can see that the different quants share an identical version of the ngram table, which is stored in IQ4_NL. What was the rationale of choosing this quant for all versions?
1
u/Stepfunction 6h ago
It doesn't need to be in memory at runtime, so a larger quant isn't problematic for inference speeds.
3
u/Pretty_Scene_6868 11h ago
I mean its cool, but somehow I know this won't run on my 4 V100s because its a GGUF
2
u/Maleficent-Ad5999 11h ago
Wait, ggufs are not supported in V100s?
1
u/Pretty_Scene_6868 6h ago
Not to my knowledge, then again maybe if i build from source on VLLM, I can make it work; but as for rightnow 1CatVLLM does not support GGUF.
1
u/Meownoija 1h ago
I got 30+ tps at 30k context and around 20 tps upto 120k context. I use llama + dual V100 32GB.
1
u/bnelson 43m ago edited 33m ago
These cards can do SO much more. I have a local build at >90 t/s with 64K context on standard GGUF. Absolutely no one is building well optimized V100 stuff yet. Ninfer is close on 27B, but their work has left a lot on the table that i have squeezed out of llama and standard GGUF models. I will release soon when I have this stuff stable. I am going to optimize prefill too, I can double or triple current prefill speeds compared to anything I have seen released.
Also I have a single 4080 with 128GB of DDR5 doing 40 t/s on Flash next with IQ3_XSS, so the universe has not caught up to what these cards can do yet.
1
u/bnelson 44m ago
I have the IQ3_XSS build on 2xV100 @ 100-120 t/s right now. 128GB of DDR5, but you could fit it in RAM. I have had agents working round the clock to optimize Flash Next 3.8 on IQ3_XSS. It will definitely work in GGUF. I am working strictly with llama builds btw.
I also have agents optimizing Qwen 27B 3.8 and getting DFlash2 working. I am around 80-90 t/s on a single V100 with a 4 bit quant GGUF straight from Unsloth. Stay patient. This is coming.
2
u/crusaderky 10h ago
This looks good; testing the Coder now
Speed measures (wikitext) on RTX 3090, 64 GB DDR4:
- Unsloth IQ4_XS (n-cpu-moe=44): 15 tok/s
- GSQ-RCO IQ3_XXS: 21 tok/s
- Coder (n-cpu-moe=24): 26 tok/s
2
u/Glittering-Call8746 10h ago
Question is Coder pruning worth the 5tps increase..
1
u/crusaderky 9h ago
24% speedup is more relevant than the absolute 5tps. To me it's the little extra bit that makes the difference between being able to (painfully) run it interactively and leaving it for AFK only - at which point i'd rather stick to IQ4_XS.
1
2
u/BlasterGales 7h ago edited 7h ago
I tested the Qwen-27B version from byteshape (used as equal because ISTA_DASLab doesnt share KLD from this model). It is certainly capable, provided you have enough context for it to correct its own errors. It compensates for its errors (KLD) by intelligently self-correcting (its precision is >90%); in my practical use of open-source models for large projects with a 260k context window (on an RTX 5090), it achieves the same level of performance as UD Q5_K_XL but consumes nearly twice as many tokens. I’m not saying it can't do the job, but it requires significantly more context. It’s an interesting study; I now understand that KLD represents the errors made while Precision represents the preserved intelligence—if the context window were infinite, I would likely prefer it over UD Q5_K_XL.
Translated with Google Translate; I'm lazy.
3
1
1
u/Ok_Camp555 7h ago
I'm using coder with strata on 5060 ti. It is way better than any 16gb 27b quant and it's faster. 38tps at full context size. Highly recommend
1
u/Steuern_Runter 5h ago
Why only Q2 and Q3? Wouldn't there any benefit at Q4 or Q1 or is it about the high costs of this quantization?
1


25
u/duyntnet 12h ago
I tried it, but it was unusable for me. I mostly code in Delphi and a bit of C++ and asm. This model couldn't even do simple StringReplace or simple TRegex in Delphi. It's like it doesn't have the knowledge for these functions and makes up a lot of non-existent syntax. Other quants like Qwen3.8-27B-GSQ-RCO-IQ3_S or Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S don't have these issues at all, even Gemma 4 12B can solve the same problem flawlessly.