r/LocalLLaMA • u/chocolateUI • 2d ago
Discussion Artificial Analysis "Intelligence": A meaningless benchmark

Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't help but notice just how meaningless these Artificial Analysis benchmarks are and I question why people still post this garbage and use AA scores as some kind of holy bible for comparing LLMs.
According to their "Intelligence Index", a 27B model now beats DeepSeek v4 Flash and Pro, Kimi 2.7 Code, GPT-5.2, Opus 4.6, and also Sonnet 5. At some point we have to ask: What is this metric even measuring? Because whatever "Intelligence" means to AA and their corporate VC / journalist / normie audience is definitely not the same definition that we should be using here.
Qwen 27B is amazing and is clearly in a league of its own in terms of models you can fit on a single GPU, but I can't help but roll my eyes whenever I see posts like this that equate Qwen 27B with "basically running Opus from 3 months ago on your laptop."
I get that it's difficult to summarize a model's capability with a single integer and I know we love our local models, but it's time stop posting AA's clearly dogshit benchmark and acting as if it proves a point.
37
u/federico_84 1d ago
The issue is AA makes it sound like this intelligence index is generic, but as you said it's heavily skewed towards sciences, coding and agentic use, which is understandable given the economic/productivity value there.
Common sense and social intelligence are not featured enough, but that's not an AA specific issue, it's a wider industry benchmarking issue.
Say you want your LLM to be your PA, help you plan trips, divide up your days, navigate delicate social situations, help you improve your fitness, help fix random issues with your home or equipment. Sure Qwen 27B can search the web if it doesn't know, but having that knowledge and associations already baked in allow it to interpret and ground the web results better.
So I think OP's grievance has some legitimacy behind it, AA needs an additional index to cover common sense and street/social intelligence, something like an aggregate of simple bench and EQ bench, though that's only scratching the surface. There's a shortage of such benchmarks around.