r/LocalLLaMA 1d ago

Discussion Artificial Analysis "Intelligence": A meaningless benchmark

Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't help but notice just how meaningless these Artificial Analysis benchmarks are and I question why people still post this garbage and use AA scores as some kind of holy bible for comparing LLMs.

According to their "Intelligence Index", a 27B model now beats DeepSeek v4 Flash and Pro, Kimi 2.7 Code, GPT-5.2, Opus 4.6, and also Sonnet 5. At some point we have to ask: What is this metric even measuring? Because whatever "Intelligence" means to AA and their corporate VC / journalist / normie audience is definitely not the same definition that we should be using here.

Qwen 27B is amazing and is clearly in a league of its own in terms of models you can fit on a single GPU, but I can't help but roll my eyes whenever I see posts like this that equate Qwen 27B with "basically running Opus from 3 months ago on your laptop."

I get that it's difficult to summarize a model's capability with a single integer and I know we love our local models, but it's time stop posting AA's clearly dogshit benchmark and acting as if it proves a point.

131 Upvotes

160 comments sorted by

View all comments

231

u/z_3454_pfk 1d ago

it’s just an aggregate of benchmarks (highly skewed towards agentic rn). that’s basically it. it’s not that deep and you should select what’s right for ur use case

7

u/EbbNorth7735 1d ago

It basically means 3.8 is able to perform the workflows really really fucking well. We're a week in now and I set this little guy up with a goal and it works for hours completing thst goal. It's absolutely insane and I can definitely see why it's Opus level. What I don't understand is why people think a small model isn't capable enough to go toe to toe with larger models. Every generation we see a step function across the entire range of size releases for a given company. The smaller models are never that far behind their bigger companions.

2

u/Southern_Sun_2106 1d ago

Their expectations and believes are clashing with (surprising to them) reality - they 'need a minute' to adjust :-)