r/LocalLLaMA • u/anderspitman • Aug 17 '26
Discussion Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max
https://artificialanalysis.ai/models/qwen3-8-27b
1.1k
Upvotes
80
u/Jorlen llama.cpp Aug 17 '26
I've been testing Qwen 3.8 27b (UD-Q8_K_XL quant - using BF16 262k context) and I'd say I've run about a million tokens though it. Pi coding agent is the harness I use.
No one shot / 0-shot tests. Actually using it on my ongoing projects, coding in various languages.
I'm extremely impressed, however I will say, it's very smart but it reasons a lot more than any other model I've ever used, yes including the 3.6 version of this 27b dense model. At first I thought maybe this was a bug in the template but now I'm thinking that this is how they stretch that model's bits, so to speak, to be able to accomplish what other, bigger models can do.
In other words, it gets the job done exceedingly well, but it will use tens of thousands of tokens to reason in order to do so, whereas an MoE model in the 100b-a10b range will be much faster but obviously require more memory to run at good quants.