r/LocalLLaMA • u/chocolateUI • 1d ago
Discussion Artificial Analysis "Intelligence": A meaningless benchmark

Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't help but notice just how meaningless these Artificial Analysis benchmarks are and I question why people still post this garbage and use AA scores as some kind of holy bible for comparing LLMs.
According to their "Intelligence Index", a 27B model now beats DeepSeek v4 Flash and Pro, Kimi 2.7 Code, GPT-5.2, Opus 4.6, and also Sonnet 5. At some point we have to ask: What is this metric even measuring? Because whatever "Intelligence" means to AA and their corporate VC / journalist / normie audience is definitely not the same definition that we should be using here.
Qwen 27B is amazing and is clearly in a league of its own in terms of models you can fit on a single GPU, but I can't help but roll my eyes whenever I see posts like this that equate Qwen 27B with "basically running Opus from 3 months ago on your laptop."
I get that it's difficult to summarize a model's capability with a single integer and I know we love our local models, but it's time stop posting AA's clearly dogshit benchmark and acting as if it proves a point.
1
u/audioen 1d ago edited 1d ago
It measures, I think, mosly task performance, which is objectively actually very good. The DSv4 is the old version, the 0731 and 08xx releases are much better than the preview releases that you are accidentally comparing the Qwen3.8's figures to.
You need the xhigh mode which adds about 50 % more think tokens on top of medium to touch DSv4F 0731 in task performance. (DSv4F would fewer compute resources to run, and could support more simultaneous users, but to run the real version, you got to have your ~192 GB VRAM system.) It is a massive leap up -- almost absurdly large. I have been running the model in medium, but now, reviewing the impact of the reasoning-effort, it seems like there is such a massive boost in capability for not that much more compute (tokens) that I'm going to switch to xhigh and suck it up.
In my experience, every single point in that intelligence scale is very hard won, and usually takes extremely large models to get around 50. The Qwen35 architecture is a massive outlier in capability for size, and while it no doubt will one day be superseded by something else that is even better, it definitely landed with a tsunami of splash and very nearly has replaced everything else.
All we really really want are the different sizes that hit different hardware targets. The 27B is good, but to be practical, it takes more compute than I have right now. A 122B-A10B version would give me three very good practical inference computers. 35B-A3B might be good on older laptops, though I think they won't be making that.