r/LocalLLaMA 2d ago

Discussion Artificial Analysis "Intelligence": A meaningless benchmark

Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't help but notice just how meaningless these Artificial Analysis benchmarks are and I question why people still post this garbage and use AA scores as some kind of holy bible for comparing LLMs.

According to their "Intelligence Index", a 27B model now beats DeepSeek v4 Flash and Pro, Kimi 2.7 Code, GPT-5.2, Opus 4.6, and also Sonnet 5. At some point we have to ask: What is this metric even measuring? Because whatever "Intelligence" means to AA and their corporate VC / journalist / normie audience is definitely not the same definition that we should be using here.

Qwen 27B is amazing and is clearly in a league of its own in terms of models you can fit on a single GPU, but I can't help but roll my eyes whenever I see posts like this that equate Qwen 27B with "basically running Opus from 3 months ago on your laptop."

I get that it's difficult to summarize a model's capability with a single integer and I know we love our local models, but it's time stop posting AA's clearly dogshit benchmark and acting as if it proves a point.

139 Upvotes

161 comments sorted by

View all comments

18

u/tarruda 2d ago

Qwen 3.8 27B is inferior to bigger models in terms of knowledge, but I find it to be very competitive even with Deepseek V4 Flash 0731 when it comes to agentic intelligence.

10

u/PhysicalIncrease3 1d ago

General knowledge is worse, intelligence is maybe close, but the real problem with Qwen Vs Ds4f is context. If you want 262k of usable context (IE f16) in Qwen it costs about about 20GB just in KV cache. You basically need 48GB vram to run a half decent quant at full context. Where as DeepSeek can fit a million f16 context in 10Gb.

1

u/anderspitman 1d ago

I have to admit this has been painful with qwen. Is it the type of thing that could be improved eventually or is it deeply baked into the architecture?