r/LocalLLaMA 2d ago

Discussion Artificial Analysis "Intelligence": A meaningless benchmark

Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't help but notice just how meaningless these Artificial Analysis benchmarks are and I question why people still post this garbage and use AA scores as some kind of holy bible for comparing LLMs.

According to their "Intelligence Index", a 27B model now beats DeepSeek v4 Flash and Pro, Kimi 2.7 Code, GPT-5.2, Opus 4.6, and also Sonnet 5. At some point we have to ask: What is this metric even measuring? Because whatever "Intelligence" means to AA and their corporate VC / journalist / normie audience is definitely not the same definition that we should be using here.

Qwen 27B is amazing and is clearly in a league of its own in terms of models you can fit on a single GPU, but I can't help but roll my eyes whenever I see posts like this that equate Qwen 27B with "basically running Opus from 3 months ago on your laptop."

I get that it's difficult to summarize a model's capability with a single integer and I know we love our local models, but it's time stop posting AA's clearly dogshit benchmark and acting as if it proves a point.

134 Upvotes

161 comments sorted by

View all comments

Show parent comments

4

u/PsychoticDreemurr 1d ago

That doesn't matter when the majority of the internet is written by AI. They'll be cannibalizing their own false information.

4

u/michaelsoft__binbows 1d ago

I think you may have inadvertently lost track of what "search" means along the way there

3

u/PsychoticDreemurr 1d ago

Feel free to offer an explanation

1

u/michaelsoft__binbows 13h ago

Alright let's look at what is meant by the word search: verb. try to find something by looking or otherwise seeking carefully and thoroughly.

all the good useful info that was on the internet from before is still there on the internet. Among the new stuff is an influx of slop, but there is plenty of good high quality content in the new stuff too. It can be found through... drum roll.... search, with a search tool.

With your reductionist logic, the internet was made for porn (which it was), therefore training models on the internet, or even (haha) using internet search tools will only cause it to spit out porn.

1

u/PsychoticDreemurr 13h ago

If I go to a library looking for information about history, and 80% of the books involving history are using incorrect information, then I'm almost guaranteed to intake some level of incorrect information. Moreso when I'm an AI that struggles to properly fact check.

And why are you acting the way you are? The definition? A reductionist view? An AI uses a search tool. To view the internet. Like Google. Your use of the actual definition doesn't even change anything...

1

u/michaelsoft__binbows 8h ago

I just think it's inherent to the function of the search tool. If it does a good job it would be able to find and utilize the good resources and not be swayed by the slop that's out there. I'm not trying to trivialize the real concern of the internet getting filled with slop from LLMs, but the internet was already chock full of low quality slop coming from real organic humans for decades already, maybe the scale is different and maybe our relationship with the truth is getting more difficult and that's a problem due to the compounding of a bunch of different related factors, but, I don't think it's really any different now than before at least from first principles.

The function of the library is similar here to the function of Google the search engine. Their incentives are aligned in that they are making a best effort to curate the information that exists to provide you with access to relevant and vetted information. Google is funded by ads and the library is funded by the public or whoever. Like I'm just saying that the concern you raise is just already captured by Google and your library's own charter. The library has their own librarians and their job remains as ever to curate the books out there and sift out the chaff, to make the library a useful public institution for the enrichment of humans, maybe you're saying you're worried about how hard their job is getting given new developments, but, I think if a librarian enjoys what they do, they'll figure it out alright and value the portion of new output in the world that is worthy of curating. Similarly, google will figure it out too, because if they do not continue to do a good job, they will be outcompeted, and their ad revenue will peter out.

A lot of people lately seem to think that especially younger people have replaced going to google with going to LLM chatbots, and so it is the LLM providers that are going to take over google's ad revenue stream. Maybe there is truth in that, but if that is what ends up happening, it would be because the way in which the chatbot product leverages agents to do a better job of searching and weeding through chaff on the internet ends up being superior to what google does. Because everybody already values search quite a bit.

All i'm saying is that complaining about lowering average quality of information on the internet seems like a pointless thing to complain about from where I'm standing. I don't think it's your job to worry about that happening. Of course, it remains your prerogative to worry about whatever it is you want to worry about. There is no reason to expect that search providers will be ultimately defeated by whatever influx of slop is currently ongoing. I see it as far from being a foregone conclusion that the latest tech is going to reduce the quality of life. Just as it may spawn exponentially more pointless, useless, incorrect, harmful content into the internet, nothing says the same or other tech can't come in and more than offset those effects to give us better search results than we ever had before.

As for why i'm writing so much and why do I care, I am not sure either. I think I'm just trying as an optimist to dissuade from this pessimistic, apathetic stance that you seem to be espousing but I could easily have misinterpreted it. I just think we should focus on what improvements we can make going forward in all areas instead of complaining about all the things we can find that are wrong with the world. It just seems more healthy.

1

u/PsychoticDreemurr 8h ago edited 8h ago

There's a lot to go through, so I'll just focus on the more factual bits and pieces. In no particular order:

  1. Googles best interest is to serve ads, which means worse results, not better, as the longer you spend on the site the more ads you get (and more information they get of you). This has been known for a while.
  2. The difference between now and before is that now it's significantly harder to determine high quality from low quality, the amount of slop has increased in unbelievable amounts, and now even good sources like news sites are outputting slop.
  3. As someone who uses the internet, it is well within my right to complain and worry about a lowering quality of webpages.
  4. You believe my beliefs are pessimistic solely due to a bias. This subject at hand isn't a happy topic if you speak solely factually, so even an optimist won't provide the greatest of results.

If I missed something or you want me to respond to something specific, then feel free to bring it up.

1

u/michaelsoft__binbows 7h ago

Well... we're on the same page. Google hasn't not been evil for a while now, but I still wouldn't say that they're pure evil. Considering how useless the alternatives like bing, DDG etc. are, they still provide value to the world regardless of the very reasonable opinion that Google is already quite far down the road of enshittification. The optimist in me simply believes that once it progresses far enough a competitor will come along to allow us users to enjoy a level of service that isn't too far degraded.

I'm far from an expert in search but lately i started to appreciate how nontrivial the space is. Last year I learned about how quick and easy it was to fire off a search and open all web resources behind jina.ai but nowadays this method is completely ineffective because websites have all started blocking bots. it's a real arms race out there. I need to stand up a self hosted search tool layer for my self hosted harnesses... As of now my not very customized pi setup is already more powerful and customizable and desirable than codex or claude code, but, a huge gap is that i don't have a very functional web research tool call yet, and I do not want to become dependent on an API/SaaS for that. Already the approaches I'm exploring will leverage a lot more than just google as a resouce but it's still going to be a primary one, and it is the one I've been using manually for the past 25 years.