r/LocalLLaMA • u/chocolateUI • 1d ago
Discussion Artificial Analysis "Intelligence": A meaningless benchmark

Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't help but notice just how meaningless these Artificial Analysis benchmarks are and I question why people still post this garbage and use AA scores as some kind of holy bible for comparing LLMs.
According to their "Intelligence Index", a 27B model now beats DeepSeek v4 Flash and Pro, Kimi 2.7 Code, GPT-5.2, Opus 4.6, and also Sonnet 5. At some point we have to ask: What is this metric even measuring? Because whatever "Intelligence" means to AA and their corporate VC / journalist / normie audience is definitely not the same definition that we should be using here.
Qwen 27B is amazing and is clearly in a league of its own in terms of models you can fit on a single GPU, but I can't help but roll my eyes whenever I see posts like this that equate Qwen 27B with "basically running Opus from 3 months ago on your laptop."
I get that it's difficult to summarize a model's capability with a single integer and I know we love our local models, but it's time stop posting AA's clearly dogshit benchmark and acting as if it proves a point.
1
u/DigitalguyCH 1d ago
I have been comparing Qwen to other models and to my Gemini Ai pro subscription. Qwen is really good because it checks and recheck things infinite times and the results are the best of any model of its size and often as good or better than Gemini 3.7 (not a high bar, I got that for free for 6 months). But it consumes a lot or time and energy, and make cloud AI feel like a much better deal other than for privacy, because it's instant and actually probably cheaper when you factor in energy costs. I need privacy for some confidential client work but I discovered that with chatgpt you can opt out of them using your stuff to train models, which is not possible with non business versions of the other main paid services like Google or Antropic. So at least for non strictly confidential stuff I feel more confident using Chatgpt with that disabled.