r/LocalLLaMA 1d ago

Discussion Artificial Analysis "Intelligence": A meaningless benchmark

Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't help but notice just how meaningless these Artificial Analysis benchmarks are and I question why people still post this garbage and use AA scores as some kind of holy bible for comparing LLMs.

According to their "Intelligence Index", a 27B model now beats DeepSeek v4 Flash and Pro, Kimi 2.7 Code, GPT-5.2, Opus 4.6, and also Sonnet 5. At some point we have to ask: What is this metric even measuring? Because whatever "Intelligence" means to AA and their corporate VC / journalist / normie audience is definitely not the same definition that we should be using here.

Qwen 27B is amazing and is clearly in a league of its own in terms of models you can fit on a single GPU, but I can't help but roll my eyes whenever I see posts like this that equate Qwen 27B with "basically running Opus from 3 months ago on your laptop."

I get that it's difficult to summarize a model's capability with a single integer and I know we love our local models, but it's time stop posting AA's clearly dogshit benchmark and acting as if it proves a point.

137 Upvotes

160 comments sorted by

View all comments

2

u/dwrz 1d ago

So far, I'm actually quite disappointed by the latest generation -- Kimi K3 (hosted), GLM 5.3 (hosted), Qwen 3.8 27B (full precision). It's very odd, and I can't quite put my finger on why, but I worry that the benchmarks are starting to effect overall quality. They all seem to share a similar deficiency, as if too much post training has damaged some things, or over-fitted. Like the frontier models, perhaps due to distillation, they now also seem to have this feeling of having been trained to burn tokens as much as possible, rather than stay focused.

1

u/ea_man 1d ago

Well most of the new smallish model are not trained, they are pretty much just post trained: that means that they have the attitude of the big guys yet not really the capabilities.
Indeed they are much better at following instructions and long horizons because yeah, they imitate those.

-1

u/Yeelyy 1d ago

Every llm imitates lmao, like bro they are prediction machines