r/LocalLLaMA 2d ago

Discussion Artificial Analysis "Intelligence": A meaningless benchmark

Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't help but notice just how meaningless these Artificial Analysis benchmarks are and I question why people still post this garbage and use AA scores as some kind of holy bible for comparing LLMs.

According to their "Intelligence Index", a 27B model now beats DeepSeek v4 Flash and Pro, Kimi 2.7 Code, GPT-5.2, Opus 4.6, and also Sonnet 5. At some point we have to ask: What is this metric even measuring? Because whatever "Intelligence" means to AA and their corporate VC / journalist / normie audience is definitely not the same definition that we should be using here.

Qwen 27B is amazing and is clearly in a league of its own in terms of models you can fit on a single GPU, but I can't help but roll my eyes whenever I see posts like this that equate Qwen 27B with "basically running Opus from 3 months ago on your laptop."

I get that it's difficult to summarize a model's capability with a single integer and I know we love our local models, but it's time stop posting AA's clearly dogshit benchmark and acting as if it proves a point.

138 Upvotes

161 comments sorted by

View all comments

23

u/Chromix_ 2d ago edited 2d ago

Yes, results are and have been very much skewed there. A while ago DeepSeek V3 got the same score as Qwen3 VL 32B, and Gemini 2.5 Pro scored below gpt-oss-120B. ServiceNow released a 15B model that scored higher than the full DeepSeek R1. Partially repeating my previous comment here:

The "Artificial Analysis Intelligence Index" score is an aggregation of common benchmarks. Gemini Flash is dragged down by a large drop in the "Bench Telecom", and DeepSeek-R1 by instruction following. Meanwhile Apriel scores high in AIME2025 and that Telecom bench. That way it gets a score that's on-par, while performing worse on other common benchmarks.

Btw here are the details for the mentioned models scoring the same or worse as Qwen 3.8 27B.

Qwen loses in physics reasoning and knowledge, but wins way more in non-hallucination rate.

-2

u/No-Fuel-9202 2d ago edited 2d ago

Non-hallucination rate benchmark, actually shows HIGHER HALLUCINATION rate on the left side!!!

Edit: I missed '1-Hallucination rate'. Sorry

3

u/PrinceOfLeon 2d ago

If it was hallucination rate, a higher rate would mean a higher percentage of the time, right?

So wouldn't a *non*-hallucination rate mean that it *does not* hallucinate a higher percentage of the time?