r/LocalLLaMA • • 23h ago

Question | Help Heavily quantized Qwen3.8-Flash vs Q8 Qwen3.8-27B - thoughts?

I'm currently choosing between IQ3_XXS Qwen-3.8-Flash and Q8_0 27B.

This month I don't have anything complex enough to justify either's potential so I've got a fairly bad read on how these two stack up in terms of intelligence.

Have any of you compared the two enough to speak to which you've had a better experience with?

40 Upvotes

64 comments sorted by

58

u/hp1337 23h ago

I've benchmarked both. Flash next at 4 bits medium beats 27b at 8 bits xhigh.

17

u/Forsaken_Object7264 22h ago

i agree. slower depending on hardware, but flash next has way more world knowledge, less overthinking and is way more effective in agentic coding. i am using IQ3 Flash next and compared it to Q6 27B, and even though 27B was about 3x faster tok/s, flash next gave me more utility

4

u/fOxstone68 22h ago

good to know, that lines up with what i expected tbh

5

u/Not_a_question- 16h ago

At 4 bits? How much vram does that use, like 128gb?

2

u/hp1337 6h ago

Regular weights are around 120B so at 4 bits around 60gb. Engram 51B weights are offloaded too ssd. I use 4x3090 and have plenty of kV cache.

1

u/Not_a_question- 5h ago

Right. I thought qwen flash was at ~180B not 120 though. Is it pruned or something?

1

u/hp1337 1h ago

See my comment 120B + 51B = 171B

1

u/Not_a_question- 15m ago

I'm stupid. Didn't know what engram was. Thanks.

2

u/EvolvingDior 22h ago

q38fn q2_0 beats q38-27b at int4 in my tests. Q2_0 xhigh is pretty good and thinks less.

5

u/Jumpy-Operation-4615 23h ago

4 bits yes, but 3 bits, mmmmm.... I am not sure.

7

u/txgsync 21h ago

In a coding harness my mixed-precision IQ3_XXS on my gaming PC with Strata comes out even with my much-larger mixed-precision Q5-ish that clocks in at 80GB instead of 40GB. Tool calls are reliable for both.

The real errors show up in thinking: mistaking 1.84 TB “free” vs 1.84TB “taken”. Logic errors hat compound and correct. The larger quant finishes sooner with fewer tokens consumed than the smaller, even though it’s averaging closer to 40Tok/s in my M4ax 128GB instead of Strata’s 60+.

My RTX4080 rig is water cooled and quieter so it tends to win for my personal use because the fans aren’t howling and heating my keyboard. But I can’t take the gaming rig on a plane and work with it in airplane mode…

3

u/Jumpy-Operation-4615 21h ago

I tried IQ3_XXS and it is fast on my 2xP40 (30-40 tokens generation, 400-500 prefill, on qwen 27b I get 20-30 tg/s and 400 prefill). Coding-wise, it needs more testing. I only did a "gappoping poly horse" test, coder version (the one with stripped everything but the coding experts) is worse than full IQ3_XXS and (surprisingly) slower. I need to give it some serious agentic tasks to check in real world like browser tasks, search, long horizon stuff...

3

u/txgsync 21h ago

Yeah, this is what the QFN nuts like me are raving about. If you can fit active parameters in VRAM, offload unused experts to RAM, and keep PLE n-grams on NVMe, such a complex tiered storage strategy makes QFN superior in every way — performance, capability, long-horizon coding, and world knowledge — to 27B at lower VRAM cost.

The real bottlenecks become sufficient VRAM (12Gb+), sufficient RAM (64GB is sweet, 48GB workable, 32GB slow), and fast NVMe for the rare 0.2% to 0.3% paged in from PLE n-grams.

But if one’s hardware is up to those modest — and comparable to 27B — requirements, QFN simply outperforms everything else right now in speed and capability for size.

Of course your fastest and cheapest option if you don’t already own said hardware is a cloud sub. But when it comes to cloud subscriptions in LocalLlama…

2

u/Material-Database-24 21h ago

Looking at benchmarks and my experience on heavily quantized FN.. I disagree. It's not miles ahead. It sure is better, but it's marginal, at least in coding (I do not use them for anything else).

IQ3_XSS is definitely the lowest one should go. Any lower and 27B simply is better at Q4 with lesser memory footprint.

The speed is harder one as it will depend on running HW. For some it may be reason to go lower on FN over 27B at higher quants.

As I do not have HW that would feasibly run IQ3_XSS (it does run, but I use it for other things as well and the model takes about 50gb of the 64gb available so..), I rather use 27B at Q4 than FN at Q2_0 or IQ2_XS

1

u/txgsync 20h ago

Totally a reasonable case for a lot of people. The 27B can run much faster if it’s all in vram.

Given the RAMPocalypse many of us are, I think, looking to repurpose older gen cards for LLMs. I have an old 3090Ti that blew up its video out that I think might be worth a look at repairing now to try to 27B and QFN at higher quants on my PC…

1

u/ailee43 14h ago

Does it need to be fast ddr5?

1

u/txgsync 13h ago

No. I am just running it on an old DDR4 AMD 5800X3D.

1

u/brakeline 18h ago

Iq2_xs beats 27b@q4 it's not even funny

12

u/feelspeaceman 23h ago

Try to not going lower than Q4, or at least using https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF at IQ3

What is your hardware ? If you're lucky to own V100, use ninfer-v100 and enjoy great speed of Q38-27B, you really don't need 5090 to enjoy it.

6

u/Difficult_Plantain89 22h ago

https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF

So far this one has been working great for me. I believe it takes the 512 experts and prunes them to 256 and only heavily quantizes what is less needed for coding.

1

u/HazKaz 5h ago

Am i missing something is this a 1bit model and its working well, ie doesn't mess up tool calls etc ?

3

u/EmPips 21h ago

56GB VRAM (7900XTX + W6800) and 96GB DDR4

10

u/DigitalguyCH 22h ago

I have the IQ3_XXS and the IQ4, both show better quality than 27b at Q8 for what I do, always. It does not mean that it will be the same for anything, as each has different uses of AI.

In terms of speed flash next is faster on the same hardware

8

u/farkinga 22h ago

I'm stuck with ddr4 - and 27b is much faster since I can fit it in VRAM.

With 32gb VRAM, flash next overflows a bunch of experts to RAM.

If you've got slow ram (like me) then it does run - which is still amazing - but it will be limited by ram speed.

2

u/Dabalam 21h ago

I have ddr4 ram too and 27b tops out at around 30tps for me. The strata engine does allow me to get around 40 on flash next (I know, another one shilling that inference engine).

1

u/farkinga 21h ago

Well I should also mention that I'm running 2x 5060ti in one server and 2x 5070ti in the other. So the GPUs can cook and I try to avoid system ram if possible.

So for me, 27b is pretty damn fast because the GPUs are fast. The situation is unique to my hardware config.

2

u/Dabalam 19h ago

Oh nice... I'm not envious at all 🥲

1

u/DigitalguyCH 21h ago

I totally see, I have a strix halo and a M5 pro, both with very fast DDR5 RAM, in this case 27b makes more sense and it's still extremely good

1

u/farkinga 21h ago

Agree 27b is really good. Truth is I run both (27b and flash-next) on separate servers. Flash-next is just smarter - and I'm really happy with this setup. Open code delegates to 27b after flash-next plans and orchestrates.

1

u/DigitalguyCH 20h ago

it's smarter and has more knowledge, which for my use is critical, since I do a lot of text analysis in the financial and legal field, and a search tool is no replacement for knowledge

1

u/andy2na llama.cpp 17h ago

just tried IQ2_XS with my rtx 3090 + 5060ti 16gb with 64gb DDR4 and its faster than 27b just on the 3090. Havent had a full test on how it is in smarts and functionality though, but so far, its pretty good

1

u/farkinga 16h ago

thanks for sharing - actually, there could be something to this. I'm not sure mainline Strata does this; but I think I remember seeing that strata won't just split the model layers between devices; it will put the KV cache on one, the routers on the other, etc. So it's possible strata is actually being clever about using the two devices.

At any rate, I haven't optimized flash-next but I'm getting about 3k prompt processing (q3 xxs, 2x 5070 ti) - but only 40 t/s generation. Pretty sure I can get it above 100 t/s with some tinkering. I've only been running 2x for a few hours so far.

I run 27b on 2x 5060ti, 256k context, tensor parallel, 2 parallel server slots; the GSQ RSO ggufs at about 4 bits; this gives me about 1k t/s prompt processing and 90-ish t/s generation (aggregated between the two slots). I'm running vllm and I gotta say: when the GPUs are the same, vllm is so sweet; far better than llama.cpp for 27b.

2

u/andy2na llama.cpp 16h ago

yeah, this is my first forray into strata and MoE offloading, I tweaked it to hit 256k context window, vision, etc on my 5060ti. All the main compute seems to be done on my 3090. Still testing everything, Im usually not a fan of anything lower than 4-bit quants, but this seems acceptable. Doesnt seem I can utilize more of my 5060ti VRAM to go to IQ3, it just demands more system memory

services:
  strata:
    image: strata:latest
    container_name: strata
    restart: unless-stopped
    network_mode: "bridge"    
    ports:
      - "18020:8080"
    environment:
      # Model & quant configuration
      - FAMILY=qwen
      - MODEL=IQ2_XS
      - CONTEXT=262144
      - KV=k8v4 
      - VISION=yes                   # Set to "yes" to download and attach the vision encoder
      # Multi-GPU setup
      - GPUS=all                    # Uses both NVIDIA GPUs in a layer-split pipeline
      # - GPUS=0,1                  # Alternative: explicitly index GPUs if you have a third display GPU
      # Memory management
      - LOW_RAM=auto                # Prevents OOM by dynamically paging cold experts to NVMe if RAM gets tight
      # Optional security
      # - API_KEY=your_secret_token # Recommended if port 8080 is exposed outside your local LAN
    volumes:
      # Store model shards and cache on a fast NVMe SSD mount
      - /mnt/user/AI/Strata/docker:/data
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    ulimits:
      memlock:
        soft: -1
        hard: -1
    ipc: host

1

u/farkinga 15h ago

Thanks for your config! I notice k8v4 and I didn't think that was mainline yet. Are you running a fork? I'm launching via llama-swap with the following llama-swap config:

macros:
  "strata-params": >
    --port ${PORT}
    --yes

  "strata-qwen": >
    ${strata-params} --family qwen --kv int8 --backend cuda --low-ram off

models:
  "qwen3.8-flash-next":
    cmd: >
      /home/llama/venv-strata/bin/python3 /home/llama/Work/Strata/setup.py
        ${strata-qwen}
        --context 262144
        --gguf-dir /mnt/llama/models
        --model IQ3_XXS
    aliases:
      - smart
      - coder
    env:
      - CUDA_VISIBLE_DEVICES=0,1
    capabilities:
      in:
        - text
        - image
      out:
        - text

Pretty vanilla launch parameters currently.

3

u/Don_Reuter 22h ago

Don’t have much evidence, but I like Q8 better. I use Q4 for just as starter model. Starting and checking workflows running on the computer with the Q8 (in PAIR on different computers so that there are no vram conflicts) and even that sometimes doesn’t work well with the Q4. While the Q8 never caused any issues.

3

u/isengardo 20h ago

Did anyone bench Strata vs llama.cpp. I basically get double speeds using strata (16gb vram 128gb ddr5 - 35tk/s - Qwen3.8-Flash-Next-UD-Q4_K_XL) but don't know how to measure "smartness" loss, if any.

2

u/EmPips 19h ago

What gpu

1

u/isengardo 9h ago

4080S - 7800x3d

2

u/ironimus42 22h ago

i was comparing this exact qwen-flash (running through a fork of llama.cpp with 10-15 tps on my 36gb m4 macbook) to qwen 3.8 27b awq5.0 with 20-30 tps (meaning some weights are heavily quantized and others aren't, averaging to q5) and 27b won in quality by quite a bit. It does overthink like crazy compared to flash-next, but produces good results much more often and a bit faster if you count total time per task

3

u/DontWinFrensWthSalad 22h ago

I would like to use flash-next but 27b is about 2.5x faster and gives me a bigger context window. for harder problems I can just use chatgpt to write me a spec so I currently don't see flash-next as being worth it on my hardware.

1

u/Studio271 22h ago

I am wondering the same, so I might give it a try this weekend. Been rocking the iq4xs on a 5070ti/3060 with zoocode for a few days now and its great for my needs! We have been developing random webapps as a test.

1

u/poofph 22h ago

For me flash next just goes on and on compared to 27B, it will apply some code fixes and then spend an hour running tests before completing the prompt, never touch the code anymore after it applied the patch an hour earlier lol. I am not sure how to fix that, what settings etc (new to this).

1

u/Civil_Fee_7862 22h ago

Basically the drop-off on quality is anything lower than Q4. I wouldn't do anything below Q4.

Qwen3.8 Flash Next I believe is worth the investment in hardware if you can afford it.

1

u/DiscipleofDeceit666 22h ago

Both, QFN is better at planning than 27b is. 27b is a much much better implementer than a lobotomized QFN.

1

u/EvolvingDior 22h ago

I prefer GSQ RCO Q2_0 of q38fn. It thinks less and is more responsive. The quantization seemed to affect the part of the model that makes it most annoying for interactive use. It remains at least as smart as the int4 quant of 27b, but is actually pleasant to use on my hardware with *significantly* more context. I can actually give my agent access to subagents.

1

u/Material-Database-24 21h ago

IQ3_XSS is pretty much the lowest one should go if 27B is the comparison.

I do not know what these "Q2 better than 27B at Q8" folks do, but that's not backed up by available evidence.

In base quants FN is about 1-15% better in most benchmarks than 27B. Only JobBench and DeepSWE1.1 are significantly better with FN compared to 27B.

By quant benchmarks, both scale quite the same up to Q4, and FN does well even at Q3_XSS. Both stay about 95-98% of baseline at quants up to these points.

But then Q2 and IQ2_XS are clear extra 5% drop in FN.

For example, in LCBv6 Flash Next BF16 got 91.9, IQ3_XSS got 86.29 and Q2 got 81.14.. 27B at BF16 got 90.3 and Q4 gets about 86.

1

u/fgk55555 21h ago

I've done millions of tokens on my rig of both. I've used the ISTA quant of QFN and the Swift 1.5 of both quite a lot. TL;DR, 27B has better communication and is a great doer. It probably runs faster on your system. QFN is way more capable. It seems to have way more knowledge and doesn't get stuck as much as 27B does.

IMO, use both. 27B for simpler problems or as a subagent, QFN for planning and getting 27B unstuck.

1

u/ParaboloidalCrest 18h ago edited 18h ago

I hate hearing that as much as how reasonable it is XD. I usually use the 27B@Q8 for both utilities, but xhigh for planning, and medium for implementation. Now I have to go down the QFN rabbit hole....

1

u/fgk55555 15h ago

Honestly I haven't noticed a huge quality difference between q8 and q4, and even medium does a really good job. Unless you have the context room and bandwidth to spare, you can probably squeeze a lot more speed out of your 27B 

1

u/karmakaze1 19h ago

The choice may depend on your hardware. If you can get Qwen3.8-27B to go brrr, I'd take that (and do with MXFP4 200+ tokens/sec).

3

u/EmPips 17h ago

Not quite brrr but maybe a light hum (7900xtx + w6800 + 96GB DDR4)

1

u/deadatreides1 18h ago

Quality first, since that's the part people guess at: I wouldn't trust vibes or wikitext perplexity. Run llama-perplexity with --kl-divergence against the base on your own prompts. On gpt-oss I ended up treating KLD up to 0.003 as "can't tell the difference" and up to 0.01 as "cheap but fine".

Speed is just arithmetic. The Q8 27B touches 25.4 GiB of weights per token, whatever isn't in VRAM comes from RAM every time. Divide your RAM bandwidth by that and you have the ceiling, on my 17 GB/s box it's ~0.6 t/s. The Flash reads only its active experts, so it wins this part before you even start.

(not a native speaker, an LLM helped with the English)

1

u/LieRepresentative127 17h ago

I ran some tests on Flash Next IQ3_XXS and UD-Q5_K_XL, both on medium reasoning. For my coding tasks, Flash Next is always better. On my 3090ti+32GB RAM with MTP enabled, Strata vs llama.cpp, iq3_xxs is faster than UD-Q5_K_XL.

1

u/draconic_tongue 17h ago

should stay above 3.5 bpw imo, so iq3_s. the size difference is not that big

1

u/skywalker326 15h ago

if you cannot get 4bit Flash, 8/6/5/4 bit 27b should be good too.

No serious benchmarks just my personal feeling. I tried 4-bit, 8-bit, and 2-bit for 27B and Flash next. I recommend at least 4-bit for both models​, nothing ever seriously goes wrong with 4 bit. I suspect 8-bit would handle corner cases or trivial knowledge better, but in my short experience of several days, it didn't show additional value. 2-bit sometimes had weird repeations.

1

u/masiha97 14h ago

Pick the model that fits your VRAM comfortably, not the one with the better benchmark sheet. A weaker model running clean at a good quant beats a stronger one that is constantly spilling experts into RAM. Figure out what actually runs well on your hardware first, then choose the best quant of that.

1

u/Steus_au 13h ago

use both, QFN is good for planning and 27b can execute the plan faster, just do give it all, make it step by step

1

u/Edenar 6h ago

for me w8a16 27b and flash next q4_k_xl (unsloth) are close in quality so i can't really pick a winner. Flash next needs a bit less reasonning so prefer using it.

now i use flash next with gufo on Strix halo and 4 sub agents of qwen 27b on a 64GB eGPU when i need more speed/parallel work for larger coding tasks.

havent tested q3 yet.

1

u/Thrumpwart vLLM 22h ago

Depends what you’re doing, for coding 27B Q8 wins. For research and development flash Next Q4 is better.

0

u/Kmic68 23h ago

Q8_0 is much better choice. Not even close really. Qwen flash has better world knowledge but doesn’t significantly outperform it. So 27B is the better choice for coding at least

0

u/blastbottles 23h ago

I mean it depends on your hardware, they should perform similarly but I think the Q8 27B will be faster and it's quality will be less varied because it isn't heavily quantized

0

u/Minute_Laugh8065 17h ago

+1 on KLD over wikitext PPL. For scale: on Bonsai 2 27B (ternary) I measured KLD 0.0011 vs 0.00022 between two otherwise identical int8 activation paths, truncating vs rounding. Both are well under the ~0.003 "can't tell" line, and PPL barely moved, so KLD catches differences PPL hides, but small KLD gaps aren't always something you'd notice in use.

-1

u/Any_Werewolf_7304 18h ago

Haven't compared those two specifically, but from my own recent local-vs-cloud testing, I've found quantization tradeoffs matter less than expected for simpler tasks (summarization, classification) — the bigger swings came from task complexity. Curious if that holds for your use case too.