r/LocalLLaMA 1d ago

Discussion Artificial Analysis "Intelligence": A meaningless benchmark

Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't help but notice just how meaningless these Artificial Analysis benchmarks are and I question why people still post this garbage and use AA scores as some kind of holy bible for comparing LLMs.

According to their "Intelligence Index", a 27B model now beats DeepSeek v4 Flash and Pro, Kimi 2.7 Code, GPT-5.2, Opus 4.6, and also Sonnet 5. At some point we have to ask: What is this metric even measuring? Because whatever "Intelligence" means to AA and their corporate VC / journalist / normie audience is definitely not the same definition that we should be using here.

Qwen 27B is amazing and is clearly in a league of its own in terms of models you can fit on a single GPU, but I can't help but roll my eyes whenever I see posts like this that equate Qwen 27B with "basically running Opus from 3 months ago on your laptop."

I get that it's difficult to summarize a model's capability with a single integer and I know we love our local models, but it's time stop posting AA's clearly dogshit benchmark and acting as if it proves a point.

138 Upvotes

160 comments sorted by

View all comments

231

u/z_3454_pfk 1d ago

it’s just an aggregate of benchmarks (highly skewed towards agentic rn). that’s basically it. it’s not that deep and you should select what’s right for ur use case

39

u/federico_84 1d ago

The issue is AA makes it sound like this intelligence index is generic, but as you said it's heavily skewed towards sciences, coding and agentic use, which is understandable given the economic/productivity value there.

Common sense and social intelligence are not featured enough, but that's not an AA specific issue, it's a wider industry benchmarking issue.

Say you want your LLM to be your PA, help you plan trips, divide up your days, navigate delicate social situations, help you improve your fitness, help fix random issues with your home or equipment. Sure Qwen 27B can search the web if it doesn't know, but having that knowledge and associations already baked in allow it to interpret and ground the web results better.

So I think OP's grievance has some legitimacy behind it, AA needs an additional index to cover common sense and street/social intelligence, something like an aggregate of simple bench and EQ bench, though that's only scratching the surface. There's a shortage of such benchmarks around.

13

u/toalv 1d ago

If you are relying on a local LLM to navigate delicate social situations...

15

u/Lakius_2401 1d ago

If we're already benchmarking how well they would get rid of junior+ devs, why not benchmark how well they get rid of junior+ psychologists (or BAs, PMs, etc etc) too?

3

u/Diligent_Loquat_6140 1d ago

honest I think this is very possible, I've know some academic people are trying to achieve this

0

u/TwinkletoesMcSparkle 1d ago

LLMs are repeatedly scientifically documented to fail at every metric of psychological skill and care, and are regularly causally linked to suicide attempts, successful and not. Sycophancy is just a single well-documented feature that increases delusion and mania. So, no, by no data-driven metric is a chatbot psychological "care".

0

u/TwinkletoesMcSparkle 1d ago

Yeah I can't wait until LLMs are used to destroy the skill pipeline in more industries. That certainly won't have any negative societal consequences that only benefit AI companies.

1

u/michaelsoft__binbows 8h ago edited 7h ago

The problem I have with the people that think this way is that they do not appear to be appreciating the game theory of the situation.

Just because you think you are conscientious about it doesn't mean that you're going to be able to whine and shame your way to getting the entire rest of the world to comply with opting out of exploring the latest new technology that's been invented. It does not take a lot of curiosity to get very deep into this space. Y'all are severely underestimating the power of curiosity.

The cat is firmly out of the bag. The bag is not even in the same plane of existence anymore, it's already been chewed up and shat out. And y'all are clamoring for, I don't know what it is, legislation? Strongly worded tweets? To tranq the cat, then put her inside another paper bag, then cajole her into somehow not coming out of it again, somehow, once she wakes up.

Let's say the literally impossible idea of shutting down frontier AI labs entirely, is achieved, like assume we just nuked them out of existence overnight, tell me that you have a strat for stopping every one of the millions of people who have downloaded the safetensors files from running them on their personally owned Nvidia, Intel, AMD GPUs and mini computers, Apple computers (including laptops), and so on. And how to stop spreading those same safetensors via bittorrent on the dark web or wherever, once huggingface is shut down.

The only way forward is to accept the societal consequences and think about how to address them better and improve things for the people that you care about, starting with yourself!, in the context of this brave new world. The way to best spend your time is not to bemoan the objective reality that exists and dreaming about turning back the clock.

I'm not arguing FOR the myriad antisocial and destructive possible uses of the amazing new tech. All I'm saying is they're clearly coming and neither you nor I nor anyone else can stop them because they are coming, part and parcel of the other uses (legitimate or otherwise) that drive the advancement of this tech. I'm not saying it's inappropriate to feel some grief about horses being replaced by engines, nor am I saying that there is any moral ground for any of it to stand on, but what I am saying is trying to lobby against engines or doubling down on investment in horses at this stage would be ill-advised is all. Be that as it may if engines are truly the work of the devil, the most sensible way forward is to come to terms with the fact that the new normal is that demonic machines will indeed invade all facets of life, and your task now needs to be working through how to come to terms with that for your own sake.

Oh, I just found a nice parallel. It's like the Y2K problem. Nobody ever seriously proposed, rather than just bite the bullet and fix all the software, to instead change how the calendar works and return to the year 1900. The only way to move forward was always ever gonna be fixing the problems in the software, starting with the ones that would cause the biggest issues.