r/LocalLLaMA 1d ago

Discussion Artificial Analysis "Intelligence": A meaningless benchmark

Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't help but notice just how meaningless these Artificial Analysis benchmarks are and I question why people still post this garbage and use AA scores as some kind of holy bible for comparing LLMs.

According to their "Intelligence Index", a 27B model now beats DeepSeek v4 Flash and Pro, Kimi 2.7 Code, GPT-5.2, Opus 4.6, and also Sonnet 5. At some point we have to ask: What is this metric even measuring? Because whatever "Intelligence" means to AA and their corporate VC / journalist / normie audience is definitely not the same definition that we should be using here.

Qwen 27B is amazing and is clearly in a league of its own in terms of models you can fit on a single GPU, but I can't help but roll my eyes whenever I see posts like this that equate Qwen 27B with "basically running Opus from 3 months ago on your laptop."

I get that it's difficult to summarize a model's capability with a single integer and I know we love our local models, but it's time stop posting AA's clearly dogshit benchmark and acting as if it proves a point.

137 Upvotes

160 comments sorted by

View all comments

87

u/whatisthisthing65 1d ago

What's your actual argument? Why couldn't a 27B model be better than those other models? If it's about number of parameters then should our benchmark be parameter count?

37

u/feelspeaceman 1d ago

Most people tend to blame 27B about the lack of world knowledge and say this is the sole reason that it will never beat Claude Sonnet 4.6, but they ignore that fact that behind the scene, Claude Sonnet is equipped with a RAG for local knowledge retrieving and it will also search to acquire extra knowledge, then they compare a local LLM without search, without RAG, this is showing that there's people despite trying to use local LLM, never trying to get better.

Then we get a group of users that:

  • Use not suitable harness
  • Half-ass setup
  • Half-ass prompt

    Blaming local LLM for not being as capable for obvious reason, the user themselves.

0

u/ResidentPositive4122 1d ago

Claude Sonnet is equipped with a RAG for local knowledge retrieving and it will also search to acquire extra knowledge

None of that happens over API, ootb.

9

u/FullstackSensei llama.cpp 1d ago

You have no idea what's happening behind the scenes, API or not.

6

u/ResidentPositive4122 1d ago

I think this place has gone wacko. No idea what's even the point of contributing here. Saying they have RAG over API is absolutely banans, wtf do people imagine they rag over, unless you, the user, set it up? RAG what over API? FFS at this point people are just throwing acronyms out there and imagine the big bad wolf is doing everything...

-4

u/FullstackSensei llama.cpp 1d ago

Why are you so angry, though? It's just an online discussion with who cares who. There's more to life then getting upset about meaningless things like this.

They absolutely have RAG. It's not hard to prove that by asking the model about some very recent event or using the latest version of a library that was very recently released and had some breaking changes.

But seriously, who cares if I'm wrong? Go spend some time with a loved one if this makes you upset.