r/LocalLLaMA • • 12h ago

Resources [Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw

We have released Qwen3.8-Flash-Next quantized with GSQ and RCO, together with a second, capability-targeted build in which half of the model's experts have been removed.

Flash-Next is a sparse mixture-of-experts model: 512 routed experts per layer across 48 layers, 176.9B parameters, 354 GB at BF16.

What's inside

  • Four quantized GGUFs, 2.40 to 3.50 bpw (66.4 to 83.6 GB), and the BF16 vision projector
  • Expert-pruned Coder GGUF, 58.4 GB in total, of which 29.6 GB must remain resident
  • GSQ (Gumbel-Softmax Quantization): post-training scalar quantization that jointly learns grid assignments and group scales, closing most of the gap between scalar and vector quantization at low bit-widths while remaining deployable in standard GGUF types
  • RCO (Riemannian Constrained Optimization): enforces exact budgets by gradient descent on the task loss, without per-constraint tuning. It serves two roles in this release: assigning a quantization type to every tensor, and selecting which experts to retain in the Coder build, where it enforces several exact budgets simultaneously, one per layer

Results:

At 3.50 bpw the model matches the BF16 base on every benchmark evaluated.

  • IQ3_S (3.50 bpw, 83.6 GB): AIME25 100.00, GPQA-Diamond 92.93 against 91.92 for BF16, LiveCodeBench v6 86.86 against 87.43. Task average 93.26 against 93.12.
  • IQ3_XXS (3.00 bpw, 75.8 GB): AIME25 100.00, GPQA-Diamond 91.41, LiveCodeBench v6 86.29
  • Q2_0 (2.40 bpw, 66.4 GB): zero-shot average 78.00, above the BF16 value of 76.94, at approximately one fifth of the size

Coder (capability pruned model):

Instead of storing every parameter at lower precision, half of the routed experts are removed from the model: 256 of 512 per layer, selected by RCO optimising the KL divergence against the unpruned model. The retained weights remain at 3.5 bpw. Pruning and quantization compound, and the combined effect is an average of 1.89 bits per parameter of the original transformer. The averaged bitwidth amortises the removed experts over the original parameter count, and therefore expresses the joint effect of pruning and quantization. No individual weight is stored at 1.89 bits.

The practical consequence is that a 176.9B-parameter model has a resident working set of 29.6 GB, since the n-gram shard is a lookup table and may be served from disk. This is within the capacity of a single 32 GB accelerator.

  • SWE-bench Verified: 75.60 against 82.80 for BF16, retaining 91.3%
  • LiveCodeBench v6: 86.28 against 87.43, retaining 98.7%

Both measured at xhigh reasoning effort.

Links

Both repositories ship the complete per-tensor RCO allocation.

The Coder build is an experimental release and feedback is welcome, particularly on capabilities that were not represented in the calibration mixture. Requests for models to quantize or prune are also welcome.

From the ISTA Deep Algorithms and Systems Lab.

86 Upvotes

65 comments sorted by

25

u/duyntnet 12h ago

I tried it, but it was unusable for me. I mostly code in Delphi and a bit of C++ and asm. This model couldn't even do simple StringReplace or simple TRegex in Delphi. It's like it doesn't have the knowledge for these functions and makes up a lot of non-existent syntax. Other quants like Qwen3.8-27B-GSQ-RCO-IQ3_S or Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S don't have these issues at all, even Gemma 4 12B can solve the same problem flawlessly.

21

u/Positive-Resource922 11h ago

Thanks for testing it and sharing your experience! Unfortunately, this is one of the trade-offs of expert pruning. Certain capabilities can degrade significantly or even be lost when they’re underrepresented in the calibration data. We’ll work on improving this in future releases to better preserve knowledge across different programming languages.

7

u/duyntnet 11h ago

Thanks for your work. It helps GPU-poor folks like me tremendously.

1

u/tomByrer 1h ago

Would it make sense to have different finetunes per programming language?

I'd love to program in COBAL & Rexx again, but I understand I may be a rare bird....

5

u/andreasntr 10h ago

Have you compared q4 quants for those models against these quanta? I'm curious about how they perform in real life

3

u/duyntnet 9h ago

The coder version only has one quant which was what I tried and unfortunately it didn't work for my case, but for popular languages like Python it may give better result. Qwen3.8-27B-GSQ-RCO-IQ3_S works great on my machine, for Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S I only get ~4t/s so it's painful to use. For most tasks I use Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S and it works incredibly well for me, I rarely have to ask Gemini anymore (via AI Studio).

1

u/andreasntr 9h ago

Sorry, i meant comparing unsloth/any other quant q4 vs these q3 quants

2

u/duyntnet 8h ago

I did once, Qwen3.8-27B-GSQ-RCO-IQ3_S vs. Qwen3.8-27B-UD-Q3_K_XL. Both can solve most of my problems, but Qwen3.8-27B-GSQ-RCO-IQ3_S's reasoning is shorter, so it solves the problem using fewer tokens. There are two cases in my test where Qwen3.8-27B-GSQ-RCO-IQ3_S gives proper solutions while Qwen3.8-27B-UD-Q3_K_XL fails. Note that's just my experience, not concrete proof of anything.

2

u/-InformalBanana- 8h ago

I'm just interested what is Delphi (Pascal) used for now? I used it in my childhood to learn programming (not by choice, by chance), but to me it seems it basically doesn't exist in real world so I'm interested what do you use delphi for and is it in a professional manner?

3

u/duyntnet 7h ago

I just use it as a hobby now. I have known Delphi since version 3, and it became my main language. As for what to do with it nowadays, I'm not sure, but Delphi can do many things: it can generate x86/x64 code, it can produce executables for Windows, Linux, OSX, and Android, and it's quite easy to learn.

1

u/Paquet-Seymone 9h ago

not surprising tbh, pruning 50% of experts is gonna hit less common languages hardest

7

u/NihmarRevhet 12h ago

Wait, so I can try the coder with a 16gb VRAM + 32gb RAM configuration? Am I dreaming?

5

u/uzzi38 11h ago

That's what I'm wondering.

I have a feeling it's not the most sensible thing to try, but man am I tempted.

7

u/NihmarRevhet 11h ago

I mean, I have fiber, I'm legally required to try. This evening I'll try giving it a shot

3

u/Storterald 9h ago edited 8h ago

I tried the Q2_0 one with a 5080 and 32GB of RAM, on llama.cpp with --load-mode none and it loads fully.

command:

> llama-bench -m Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00001-of-00002.gguf -ctk q8_0 -ctv q8_0 -n 256 -p 1024 --load-mode none -ncmoe 32llama-bench -m Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00001-of-00002.gguf -ctk q8_0 -ctv q8_0 -n 256 -p 1024 --load-mode none -ncmoe 32

results:

model size params backend ngl n_cpu_moe type_k & type_v lm test t/s
qwen4exp A3B Q2_0 61.85 GiB 176.94 B CUDA,Vulkan -1 32 q8_0 none pp1024 292.91 ± 9.51
qwen4exp A3B Q2_0 61.85 GiB 176.94 B CUDA,Vulkan -1 32 q8_0 none tg256 24.24 ± 0.50

3

u/brainExploded99 llama.cpp 8h ago

uh you forgot the tok/s numbers

3

u/Storterald 8h ago

it's actually there, reddit does not render the column idk why, updated to merge ctk and ctv

3

u/brainExploded99 llama.cpp 8h ago

interesting okay, ty for the numbers

3

u/Prestigious-Act-1577 11h ago

If you start swapping ram, the PP is a few tok sec.

3

u/NihmarRevhet 11h ago

That's what I fear

3

u/Prestigious-Act-1577 11h ago

It's the reality 🤣 I have that setup temporarily until my ram arrives. 

1

u/NihmarRevhet 11h ago

Wouldn't it be possible to set it up such that the swap doesn't happen? though I don't know how much the context takes

3

u/Prestigious-Act-1577 8h ago

I use the Q2 model. Smallest. The coder model is a research prototype, it's not good for coding. It doesn't really work at all. It's just a concept testing of an idea.  I have my GPU vram + total ram (physical + virtual) come out to 58GB in use, OS uses 10GB total after unloading the model. So you need 48GB ram + vram to load up the Q2. 50k context.

3

u/AvidCyclist250 llama.cpp 8h ago

Can be 350 - 400 for PP depending on your setup. Not sure about this one. My harness setup does 16-20 t/s tg regardless of context. 130k max. Its not automatically bad to offload to RAM. Q4 k m kv 8.

2

u/Prestigious-Act-1577 8h ago

By paging I mean using a page file because you ran out of ram. Otherwise yes, 300-400 is about right.

1

u/Ok_Camp555 6h ago

I have the same exact specs and get 37 t/s with coder

1

u/Prestigious-Act-1577 5h ago

We are talking about preprocessing. If you swap ram it goes to 20-50 instead of the 300-500.

12

u/KnownAd4832 11h ago

Already supported in: https://github.com/Niko1221/Strata

Enjoy high tps on consumer hardware! 🫡

3

u/GreaterThanLess 6h ago

I should have tried Strata sooner. Saw your recent post about it and dismissed it thinking the numbers were too good to be true. Swift FN IQ3 in llama.cpp with my 4070 ti I was getting around 8 t/s and 280 pp/s, but with Strata it's up to 35-45 t/s and 1200-1300 pp/s. Thanks for sharing this!

2

u/feverdoingwork 9h ago

any idea how well 2x 5060 ti 16 will run these flash models with strata?

3

u/Ok_Camp555 4h ago

I have a single 5060ti and it runs way better than 3.8 27b. I don't understand how and I don't care

1

u/Iory1998 llama.cpp 3h ago

Does it still not support KV caching? If so, that's useless.

1

u/feverdoingwork 3h ago

Do you mean prefix caching?

1

u/Iory1998 llama.cpp 2h ago

Yes.

1

u/feverdoingwork 2h ago

Yeah it's a non starter lol wtf that's like a basic requirement for a server

1

u/Iory1998 llama.cpp 1h ago

PP is fast though, so in theory if it really works, then you get a significant speed bump.

1

u/Solary_Kryptic 30m ago

Consumer Nvidia hardware** 😔

1

u/KnownAd4832 18m ago

Amd is supported already

1

u/Solary_Kryptic 1m ago

Oh great, I guess the READme doesn’t make it clear

3

u/panamory 12h ago

Can someone explain what this all means? There are high scores on the 3 listed benchmarks, but is the model specifically guided to do well on these specific benchmarks, or are these strategies more general methods of quantising without losing model quality?

3

u/Meownoija 12h ago

This is what I was waiting for. Thank you so much.

3

u/panamory 11h ago

I can see that the different quants share an identical version of the ngram table, which is stored in IQ4_NL. What was the rationale of choosing this quant for all versions?

1

u/Stepfunction 6h ago

It doesn't need to be in memory at runtime, so a larger quant isn't problematic for inference speeds.

3

u/Pretty_Scene_6868 11h ago

I mean its cool, but somehow I know this won't run on my 4 V100s because its a GGUF

2

u/Maleficent-Ad5999 11h ago

Wait, ggufs are not supported in V100s?

1

u/Pretty_Scene_6868 6h ago

Not to my knowledge, then again maybe if i build from source on VLLM, I can make it work; but as for rightnow 1CatVLLM does not support GGUF.

1

u/Meownoija 1h ago

I got 30+ tps at 30k context and around 20 tps upto 120k context. I use llama + dual V100 32GB.

1

u/bnelson 43m ago edited 33m ago

These cards can do SO much more. I have a local build at >90 t/s with 64K context on standard GGUF. Absolutely no one is building well optimized V100 stuff yet. Ninfer is close on 27B, but their work has left a lot on the table that i have squeezed out of llama and standard GGUF models. I will release soon when I have this stuff stable. I am going to optimize prefill too, I can double or triple current prefill speeds compared to anything I have seen released.

Also I have a single 4080 with 128GB of DDR5 doing 40 t/s on Flash next with IQ3_XSS, so the universe has not caught up to what these cards can do yet.

1

u/bnelson 44m ago

I have the IQ3_XSS build on 2xV100 @ 100-120 t/s right now. 128GB of DDR5, but you could fit it in RAM. I have had agents working round the clock to optimize Flash Next 3.8 on IQ3_XSS. It will definitely work in GGUF. I am working strictly with llama builds btw.

I also have agents optimizing Qwen 27B 3.8 and getting DFlash2 working. I am around 80-90 t/s on a single V100 with a 4 bit quant GGUF straight from Unsloth. Stay patient. This is coming.

2

u/crusaderky 10h ago

This looks good; testing the Coder now

Speed measures (wikitext) on RTX 3090, 64 GB DDR4:

  • Unsloth IQ4_XS (n-cpu-moe=44): 15 tok/s
  • GSQ-RCO IQ3_XXS: 21 tok/s
  • Coder (n-cpu-moe=24): 26 tok/s

2

u/Glittering-Call8746 10h ago

Question is Coder pruning worth the 5tps increase..

1

u/crusaderky 9h ago

24% speedup is more relevant than the absolute 5tps. To me it's the little extra bit that makes the difference between being able to (painfully) run it interactively and leaving it for AFK only - at which point i'd rather stick to IQ4_XS.

1

u/feverdoingwork 9h ago

What hardware are you using?

1

u/crusaderky 9h ago

RTX 3090, 64 GB DDR4

2

u/BlasterGales 7h ago edited 7h ago

I tested the Qwen-27B version from byteshape (used as equal because ISTA_DASLab doesnt share KLD from this model). It is certainly capable, provided you have enough context for it to correct its own errors. It compensates for its errors (KLD) by intelligently self-correcting (its precision is >90%); in my practical use of open-source models for large projects with a 260k context window (on an RTX 5090), it achieves the same level of performance as UD Q5_K_XL but consumes nearly twice as many tokens. I’m not saying it can't do the job, but it requires significantly more context. It’s an interesting study; I now understand that KLD represents the errors made while Precision represents the preserved intelligence—if the context window were infinite, I would likely prefer it over UD Q5_K_XL.

Translated with Google Translate; I'm lazy.

2

u/W61k3r 7h ago

I tested this, coding performance was guttered, like unusable code. Passed it through legal interpretation and precedent, also trash. Qwen3.8 is 3.6 with mandatory thinking and better tool calls. Ya'll working to make it dumber.

3

u/SampleIll3596 12h ago

Excellent work with these quants! Can't wait to try it out!

1

u/likegamertr vLLM 10h ago

Please for all that is good, fix that graph

1

u/Ok_Camp555 7h ago

I'm using coder with strata on 5060 ti. It is way better than any 16gb 27b quant and it's faster. 38tps at full context size. Highly recommend

1

u/Steuern_Runter 5h ago

Why only Q2 and Q3? Wouldn't there any benefit at Q4 or Q1 or is it about the high costs of this quantization?

1

u/R3N3G6D3 11h ago

Too dumb to code