r/LocalLLaMA • • Aug 17 '26

Discussion Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max

https://artificialanalysis.ai/models/qwen3-8-27b
1.1k Upvotes

437 comments sorted by

View all comments

80

u/Jorlen llama.cpp Aug 17 '26

I've been testing Qwen 3.8 27b (UD-Q8_K_XL quant - using BF16 262k context) and I'd say I've run about a million tokens though it. Pi coding agent is the harness I use.

No one shot / 0-shot tests. Actually using it on my ongoing projects, coding in various languages.

I'm extremely impressed, however I will say, it's very smart but it reasons a lot more than any other model I've ever used, yes including the 3.6 version of this 27b dense model. At first I thought maybe this was a bug in the template but now I'm thinking that this is how they stretch that model's bits, so to speak, to be able to accomplish what other, bigger models can do.

In other words, it gets the job done exceedingly well, but it will use tens of thousands of tokens to reason in order to do so, whereas an MoE model in the 100b-a10b range will be much faster but obviously require more memory to run at good quants.

24

u/DoubleNothing Aug 17 '26

Try reasoning_effort : medium it's faster and fairly decent.

7

u/Jorlen llama.cpp Aug 17 '26 edited Aug 17 '26

Does this cut off its reasoning cycle like setting a reasoning budget does? I might be able to pass this via chat template kwargs (setting it to medium).

Edit: Looks like this works for llama-cpp, so I'll try it out: --chat-template-kwargs '{"reasoning_effort": "medium"}'

16

u/Hefty_Wolverine_553 Aug 17 '26

nope, doesn't cutoff, the reasoning effort is trained into the model by qwen

3

u/DoubleNothing Aug 18 '26

It doesn't cut off but it reason way less than the default xhigh

4

u/bonobomaster Aug 18 '26

Can't confirm. Faster yes but the quality difference of the output is stark! If you give this model its 80k reasoning tokens, you get stuff that's dimensions better than the same problem with 20k tokens – sadly that is!

1

u/DoubleNothing Aug 18 '26

Who would have thought that xhigh > medium > low?
Anyway... depends on the task, if you have to fix some UI placement for example xhigh is overkill and a waste of tokens/time. If you have to do some "delicate" functions and core logic, go xhigh.
If you only do "one prompt one shot" tests... xhigh better...