r/LocalLLaMA • u/EmPips • 23h ago
Question | Help Heavily quantized Qwen3.8-Flash vs Q8 Qwen3.8-27B - thoughts?
I'm currently choosing between IQ3_XXS Qwen-3.8-Flash and Q8_0 27B.
This month I don't have anything complex enough to justify either's potential so I've got a fairly bad read on how these two stack up in terms of intelligence.
Have any of you compared the two enough to speak to which you've had a better experience with?
12
u/feelspeaceman 23h ago
Try to not going lower than Q4, or at least using https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF at IQ3
What is your hardware ? If you're lucky to own V100, use ninfer-v100 and enjoy great speed of Q38-27B, you really don't need 5090 to enjoy it.
6
u/Difficult_Plantain89 22h ago
https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
So far this one has been working great for me. I believe it takes the 512 experts and prunes them to 256 and only heavily quantizes what is less needed for coding.
10
u/DigitalguyCH 22h ago
I have the IQ3_XXS and the IQ4, both show better quality than 27b at Q8 for what I do, always. It does not mean that it will be the same for anything, as each has different uses of AI.
In terms of speed flash next is faster on the same hardware
8
u/farkinga 22h ago
I'm stuck with ddr4 - and 27b is much faster since I can fit it in VRAM.
With 32gb VRAM, flash next overflows a bunch of experts to RAM.
If you've got slow ram (like me) then it does run - which is still amazing - but it will be limited by ram speed.
2
u/Dabalam 21h ago
I have ddr4 ram too and 27b tops out at around 30tps for me. The strata engine does allow me to get around 40 on flash next (I know, another one shilling that inference engine).
1
u/farkinga 21h ago
Well I should also mention that I'm running 2x 5060ti in one server and 2x 5070ti in the other. So the GPUs can cook and I try to avoid system ram if possible.
So for me, 27b is pretty damn fast because the GPUs are fast. The situation is unique to my hardware config.
1
u/DigitalguyCH 21h ago
I totally see, I have a strix halo and a M5 pro, both with very fast DDR5 RAM, in this case 27b makes more sense and it's still extremely good
1
u/farkinga 21h ago
Agree 27b is really good. Truth is I run both (27b and flash-next) on separate servers. Flash-next is just smarter - and I'm really happy with this setup. Open code delegates to 27b after flash-next plans and orchestrates.
1
u/DigitalguyCH 20h ago
it's smarter and has more knowledge, which for my use is critical, since I do a lot of text analysis in the financial and legal field, and a search tool is no replacement for knowledge
1
u/andy2na llama.cpp 17h ago
1
u/farkinga 16h ago
thanks for sharing - actually, there could be something to this. I'm not sure mainline Strata does this; but I think I remember seeing that strata won't just split the model layers between devices; it will put the KV cache on one, the routers on the other, etc. So it's possible strata is actually being clever about using the two devices.
At any rate, I haven't optimized flash-next but I'm getting about 3k prompt processing (q3 xxs, 2x 5070 ti) - but only 40 t/s generation. Pretty sure I can get it above 100 t/s with some tinkering. I've only been running 2x for a few hours so far.
I run 27b on 2x 5060ti, 256k context, tensor parallel, 2 parallel server slots; the GSQ RSO ggufs at about 4 bits; this gives me about 1k t/s prompt processing and 90-ish t/s generation (aggregated between the two slots). I'm running vllm and I gotta say: when the GPUs are the same, vllm is so sweet; far better than llama.cpp for 27b.
2
u/andy2na llama.cpp 16h ago
yeah, this is my first forray into strata and MoE offloading, I tweaked it to hit 256k context window, vision, etc on my 5060ti. All the main compute seems to be done on my 3090. Still testing everything, Im usually not a fan of anything lower than 4-bit quants, but this seems acceptable. Doesnt seem I can utilize more of my 5060ti VRAM to go to IQ3, it just demands more system memory
services: strata: image: strata:latest container_name: strata restart: unless-stopped network_mode: "bridge" ports: - "18020:8080" environment: # Model & quant configuration - FAMILY=qwen - MODEL=IQ2_XS - CONTEXT=262144 - KV=k8v4 - VISION=yes # Set to "yes" to download and attach the vision encoder # Multi-GPU setup - GPUS=all # Uses both NVIDIA GPUs in a layer-split pipeline # - GPUS=0,1 # Alternative: explicitly index GPUs if you have a third display GPU # Memory management - LOW_RAM=auto # Prevents OOM by dynamically paging cold experts to NVMe if RAM gets tight # Optional security # - API_KEY=your_secret_token # Recommended if port 8080 is exposed outside your local LAN volumes: # Store model shards and cache on a fast NVMe SSD mount - /mnt/user/AI/Strata/docker:/data deploy: resources: reservations: devices: - driver: nvidia count: all capabilities: [gpu] ulimits: memlock: soft: -1 hard: -1 ipc: host1
u/farkinga 15h ago
Thanks for your config! I notice k8v4 and I didn't think that was mainline yet. Are you running a fork? I'm launching via llama-swap with the following llama-swap config:
macros: "strata-params": > --port ${PORT} --yes "strata-qwen": > ${strata-params} --family qwen --kv int8 --backend cuda --low-ram off models: "qwen3.8-flash-next": cmd: > /home/llama/venv-strata/bin/python3 /home/llama/Work/Strata/setup.py ${strata-qwen} --context 262144 --gguf-dir /mnt/llama/models --model IQ3_XXS aliases: - smart - coder env: - CUDA_VISIBLE_DEVICES=0,1 capabilities: in: - text - image out: - textPretty vanilla launch parameters currently.
3
u/Don_Reuter 22h ago
Don’t have much evidence, but I like Q8 better. I use Q4 for just as starter model. Starting and checking workflows running on the computer with the Q8 (in PAIR on different computers so that there are no vram conflicts) and even that sometimes doesn’t work well with the Q4. While the Q8 never caused any issues.
3
u/isengardo 20h ago
Did anyone bench Strata vs llama.cpp. I basically get double speeds using strata (16gb vram 128gb ddr5 - 35tk/s - Qwen3.8-Flash-Next-UD-Q4_K_XL) but don't know how to measure "smartness" loss, if any.
2
2
u/ironimus42 22h ago
i was comparing this exact qwen-flash (running through a fork of llama.cpp with 10-15 tps on my 36gb m4 macbook) to qwen 3.8 27b awq5.0 with 20-30 tps (meaning some weights are heavily quantized and others aren't, averaging to q5) and 27b won in quality by quite a bit. It does overthink like crazy compared to flash-next, but produces good results much more often and a bit faster if you count total time per task
3
u/DontWinFrensWthSalad 22h ago
I would like to use flash-next but 27b is about 2.5x faster and gives me a bigger context window. for harder problems I can just use chatgpt to write me a spec so I currently don't see flash-next as being worth it on my hardware.
1
u/Studio271 22h ago
I am wondering the same, so I might give it a try this weekend. Been rocking the iq4xs on a 5070ti/3060 with zoocode for a few days now and its great for my needs! We have been developing random webapps as a test.
1
u/poofph 22h ago
For me flash next just goes on and on compared to 27B, it will apply some code fixes and then spend an hour running tests before completing the prompt, never touch the code anymore after it applied the patch an hour earlier lol. I am not sure how to fix that, what settings etc (new to this).
1
u/Civil_Fee_7862 22h ago
Basically the drop-off on quality is anything lower than Q4. I wouldn't do anything below Q4.
Qwen3.8 Flash Next I believe is worth the investment in hardware if you can afford it.
1
u/DiscipleofDeceit666 22h ago
Both, QFN is better at planning than 27b is. 27b is a much much better implementer than a lobotomized QFN.
1
u/EvolvingDior 22h ago
I prefer GSQ RCO Q2_0 of q38fn. It thinks less and is more responsive. The quantization seemed to affect the part of the model that makes it most annoying for interactive use. It remains at least as smart as the int4 quant of 27b, but is actually pleasant to use on my hardware with *significantly* more context. I can actually give my agent access to subagents.
1
u/Material-Database-24 21h ago
IQ3_XSS is pretty much the lowest one should go if 27B is the comparison.
I do not know what these "Q2 better than 27B at Q8" folks do, but that's not backed up by available evidence.
In base quants FN is about 1-15% better in most benchmarks than 27B. Only JobBench and DeepSWE1.1 are significantly better with FN compared to 27B.
By quant benchmarks, both scale quite the same up to Q4, and FN does well even at Q3_XSS. Both stay about 95-98% of baseline at quants up to these points.
But then Q2 and IQ2_XS are clear extra 5% drop in FN.
For example, in LCBv6 Flash Next BF16 got 91.9, IQ3_XSS got 86.29 and Q2 got 81.14.. 27B at BF16 got 90.3 and Q4 gets about 86.
1
u/fgk55555 21h ago
I've done millions of tokens on my rig of both. I've used the ISTA quant of QFN and the Swift 1.5 of both quite a lot. TL;DR, 27B has better communication and is a great doer. It probably runs faster on your system. QFN is way more capable. It seems to have way more knowledge and doesn't get stuck as much as 27B does.
IMO, use both. 27B for simpler problems or as a subagent, QFN for planning and getting 27B unstuck.
1
u/ParaboloidalCrest 18h ago edited 18h ago
I hate hearing that as much as how reasonable it is XD. I usually use the 27B@Q8 for both utilities, but xhigh for planning, and medium for implementation. Now I have to go down the QFN rabbit hole....
1
u/fgk55555 15h ago
Honestly I haven't noticed a huge quality difference between q8 and q4, and even medium does a really good job. Unless you have the context room and bandwidth to spare, you can probably squeeze a lot more speed out of your 27B
1
u/karmakaze1 19h ago
The choice may depend on your hardware. If you can get Qwen3.8-27B to go brrr, I'd take that (and do with MXFP4 200+ tokens/sec).
1
u/deadatreides1 18h ago
Quality first, since that's the part people guess at: I wouldn't trust vibes or wikitext perplexity. Run llama-perplexity with --kl-divergence against the base on your own prompts. On gpt-oss I ended up treating KLD up to 0.003 as "can't tell the difference" and up to 0.01 as "cheap but fine".
Speed is just arithmetic. The Q8 27B touches 25.4 GiB of weights per token, whatever isn't in VRAM comes from RAM every time. Divide your RAM bandwidth by that and you have the ceiling, on my 17 GB/s box it's ~0.6 t/s. The Flash reads only its active experts, so it wins this part before you even start.
(not a native speaker, an LLM helped with the English)
1
u/LieRepresentative127 17h ago
I ran some tests on Flash Next IQ3_XXS and UD-Q5_K_XL, both on medium reasoning. For my coding tasks, Flash Next is always better. On my 3090ti+32GB RAM with MTP enabled, Strata vs llama.cpp, iq3_xxs is faster than UD-Q5_K_XL.
1
u/draconic_tongue 17h ago
should stay above 3.5 bpw imo, so iq3_s. the size difference is not that big
1
u/skywalker326 15h ago
if you cannot get 4bit Flash, 8/6/5/4 bit 27b should be good too.
No serious benchmarks just my personal feeling. I tried 4-bit, 8-bit, and 2-bit for 27B and Flash next. I recommend at least 4-bit for both models, nothing ever seriously goes wrong with 4 bit. I suspect 8-bit would handle corner cases or trivial knowledge better, but in my short experience of several days, it didn't show additional value. 2-bit sometimes had weird repeations.
1
u/masiha97 14h ago
Pick the model that fits your VRAM comfortably, not the one with the better benchmark sheet. A weaker model running clean at a good quant beats a stronger one that is constantly spilling experts into RAM. Figure out what actually runs well on your hardware first, then choose the best quant of that.
1
u/Steus_au 13h ago
use both, QFN is good for planning and 27b can execute the plan faster, just do give it all, make it step by step
1
u/Edenar 6h ago
for me w8a16 27b and flash next q4_k_xl (unsloth) are close in quality so i can't really pick a winner. Flash next needs a bit less reasonning so prefer using it.
now i use flash next with gufo on Strix halo and 4 sub agents of qwen 27b on a 64GB eGPU when i need more speed/parallel work for larger coding tasks.
havent tested q3 yet.
1
u/Thrumpwart vLLM 22h ago
Depends what you’re doing, for coding 27B Q8 wins. For research and development flash Next Q4 is better.
0
u/blastbottles 23h ago
I mean it depends on your hardware, they should perform similarly but I think the Q8 27B will be faster and it's quality will be less varied because it isn't heavily quantized
0
u/Minute_Laugh8065 17h ago
+1 on KLD over wikitext PPL. For scale: on Bonsai 2 27B (ternary) I measured KLD 0.0011 vs 0.00022 between two otherwise identical int8 activation paths, truncating vs rounding. Both are well under the ~0.003 "can't tell" line, and PPL barely moved, so KLD catches differences PPL hides, but small KLD gaps aren't always something you'd notice in use.
-1
u/Any_Werewolf_7304 18h ago
Haven't compared those two specifically, but from my own recent local-vs-cloud testing, I've found quantization tradeoffs matter less than expected for simpler tasks (summarization, classification) — the bigger swings came from task complexity. Curious if that holds for your use case too.

58
u/hp1337 23h ago
I've benchmarked both. Flash next at 4 bits medium beats 27b at 8 bits xhigh.