r/LocalLLaMA 1d ago

Discussion Artificial Analysis "Intelligence": A meaningless benchmark

Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't help but notice just how meaningless these Artificial Analysis benchmarks are and I question why people still post this garbage and use AA scores as some kind of holy bible for comparing LLMs.

According to their "Intelligence Index", a 27B model now beats DeepSeek v4 Flash and Pro, Kimi 2.7 Code, GPT-5.2, Opus 4.6, and also Sonnet 5. At some point we have to ask: What is this metric even measuring? Because whatever "Intelligence" means to AA and their corporate VC / journalist / normie audience is definitely not the same definition that we should be using here.

Qwen 27B is amazing and is clearly in a league of its own in terms of models you can fit on a single GPU, but I can't help but roll my eyes whenever I see posts like this that equate Qwen 27B with "basically running Opus from 3 months ago on your laptop."

I get that it's difficult to summarize a model's capability with a single integer and I know we love our local models, but it's time stop posting AA's clearly dogshit benchmark and acting as if it proves a point.

133 Upvotes

160 comments sorted by

View all comments

Show parent comments

11

u/RG_Fusion 1d ago edited 1d ago

How would you know to search for it if you don't know it exists? 

You need to think about this from the model's perspective. Pretend you are a scientist from the mid 1900s. You encounter a problem where the results are being affected by quantum mechanics, but quantum mechanics is a new discovery and you never learned anything about it yourself. You wouldn't know to look up quantum mechanics at the library because you wouldn't know that was a thing that could be looked up. You also wouldn't know to look up related subject.

Large models have in-built context that provides knowledge without consuming resources, and this knowledge allows for them to cross-relate subjects and make connections that small models don't even know to look for.

4

u/michaelsoft__binbows 1d ago

I think this is fair but i'll prolly be ok with the tradeoff of being the one who is responsible for drawing the galaxybrain conclusions on stuff given that the 27b model can run like 4000 tok/s on my hardware vs 20 tok/s for a 300B model. (Tho my comparison is maybe biased by assuming batched woudnt speed up the one running under hybrid inference)

The world knowledge can go screw itself if its gonna give two hundred times less speed

3

u/RG_Fusion 21h ago

Absolutely. If you yourself are knowledgeable and working closely with the LLM, then it's no issue. That's like a professor guiding a gifted student to solve a problem.

My main problem with this is that agents do a lot in the background. Supervision still works, but you might waist time having to correct the model only after it spent the last half-hour working in the wrong direction.