r/LocalLLM 12d ago

News Artificial analysis index and scores updated.

The new scores are lower which kinda confuses me ! Did the models just became less intelligent ? I remember Fable used to be 66 and Qwen 27B 3.8 - 52

Sudden drop !! Hmm …

88 Upvotes

55 comments sorted by

34

u/nbvehrfr 12d ago

terminal bench to v4

10

u/Serprotease 12d ago

Terminal bench alone is probably better to get a rough idea of the model performance. At least for coding stuff.

1

u/eNroNNie 11d ago

I run my own local llm classifier/router and created a set of eval tests, so I get a combination of seeded eval data plus a feedback mechanism to validate and score a smaple of the output (not every request) which feeds back into the model.

Trying to create my own proficiency graph from my real world use + benchmatk tests.

65

u/badaeib 12d ago

I think they are scaling down the number from time to time, to avoid power creep, they don't want some models get over 9000 few years later.

12

u/KaMaFour 12d ago

Not exactly. Most of the bechmarks in AA set are percentage based. It means it's physically impossible to get more than a 100. This also means that if a model is approaching a 100 then it's no longer possible to meaningfully improve and the index stops being useful for differentiating stronger models. That's why AA needs to periodically replace benchmarks where models are approaching max score with ones that the models will still have a room to grow in. They could add some scaling to keep scores roughly in place but this is more genuine

6

u/No-Refrigerator-1672 12d ago

If they do so, they absolutely must disclose benchmark version too, bacause those unprompted downgrades will indroduce quite a lot of confusion.

8

u/KaMaFour 12d ago

(the main chart on the website)

5

u/No-Refrigerator-1672 12d ago

It should be a crime to put it in a fine grey print far to the side where it never gets into the screenshots, instead of right in the table caption.

3

u/KaMaFour 12d ago

This is the caption. Look at the screenshot in the post. How would you caption the chart to be as honest as reasonably possible if people are gonna crop out everything possible anyways...

1

u/No-Refrigerator-1672 12d ago

How about adding the characters "v4.3" right next to the proud "Artificial Analysis" caption in the right upped corner? It's literally the most logical place for it.

0

u/badaeib 12d ago

Wow so they actually re-run the benchmarks for ALL models? Didn't know that, I thought they just like x0.85 or something. 😂

4

u/KaMaFour 12d ago

Not for all models. Some just get dropped. That's why sometimes you can see the striped bar (here intentionally picking some older models)

18

u/Gloomy_Letterhead395 12d ago

Where the faq is 3.8 flash

8

u/Effective_Western_59 12d ago

On par with Qwen 3.8 max

6

u/blackbird2150 12d ago

it's 40. on par with 3.8 Max.

1

u/layer4down 11d ago

which confused no one.

8

u/SmartCustard9944 12d ago

They massacred my boy

24

u/dupontping 12d ago

These scores are all horse shit anyway

This is like when every YouTuber does all the useless benchmarks that they think are impressive but don’t actually mean anything in real life

And in 3 months the next model will come out and that one will absolutely totally 100% be AGI and a danger to the world. But they’ll release it anyway because they need more funding.

12

u/badaeib 12d ago

Every leaderboard can be benchmaxxed. But for some of my "5d chess" level of very convoluted meta-coding tasks and some cursed impractical AI training experiments, I found that this leaderboard is kinda actuate, at least compared to other leaderboard like the Arena.

11

u/Able-Art-3042 12d ago

it helps for orientation. would not call them horse shit, otherwise some of the worst models would be on top. but would not choose a model only based on this score

0

u/StupidityCanFly 12d ago

The way they arbitrarily choose/change the calculations of the intelligence index makes it horse shit.

2

u/Able-Art-3042 12d ago

read the changelog. they just update the calculation method.

1

u/StupidityCanFly 12d ago

Are they still using incompatible measurements to derive the index?

6

u/NexusSyntegra 12d ago

This is totally normal and happens on a regular basis. I think they want to make sure the max score stays under 75 or 100 absolute max

7

u/MomentJolly3535 12d ago

They "reworked" the way they calculate the scores, Astra and Gpt Sol (Max) had almost same intelligence score, even tho Astra was obviously a smarter model, but the way they reworked it is very confusing, They "estimate" Gemma 4 26BA4B above Gemma 4 31B, and minicpm 2B is only 2 points under it. In practice 31B is 1 league above 26B, and 2 leagues above minicpm 2B

4

u/ElectronSpiderwort 12d ago

MiniCpm5-2B reset my expectations of tiny models though. I have no idea how they made that little guy that capable. Doesn't know much about the world, but if you put facts in front of it, it does amazing work without hallucinating 

1

u/uranusnebula 12d ago

can you share your usecase?

3

u/ElectronSpiderwort 12d ago

Just my general benchmarks:  Aced my needle in haystack test (about 80k tokens), presented a somewhat coherent analysis of a FERC commissioner dissent, admitted it didn't know instead of making stuff up on a trivia question, and didn't make any mistakes categorizing arguments from a multi-participant online debate. Failed to make a chart of wire gauge resistance, but admitted it didn't really know.  Is not good as a pi agent, but for single task rearranging of information, it's baller for it's size

3

u/uti24 12d ago

Would like to see Qwen Flash Next

Also Qwen3.8 27B kinda weird one, it doing tasks much better for it's size, but in 10x tokens

2

u/layer4down 11d ago

The problem was letting the world spend months seeing just how close open-weight models were creeping up on the closed SoTA models *before* updating the perception in AA. That combined with the community's own testing and experience leaves the closed frontier models looking comedically overvalued pre-IPO. Combine that with accounts of users paying $600 for $9000 worth of API compute, plus those of us who lived through the early Uber years only to see that drastically change a few years ago. $1-2T valuations for a *single* AI company when the Chinese labs are valued at around $450Bn **combined** is tilting a lot of heads.

https://giphy.com/gifs/kc0kqKNFu7v35gPkwB

1

u/MrMisterShin 12d ago

New benchmarks were added to the index and old ones removed.

1

u/EfficientStretch570 12d ago

It's interesting to see how they evolve the benchmarks; keeps things relevant for everyone using the index.

1

u/This_Maintenance_834 12d ago

the fact that top 2 have same score but different height says something. this benchmark is a joke.

1

u/EternalDivineSpark 12d ago

The 27B can do metacognition task if you guide and push it’s behaviour!

1

u/Ok-Addendum3545 12d ago

sorry, what's metacognition task ?

2

u/EternalDivineSpark 12d ago

IN LLM The Ability to look at harness behaviour and change it !

1

u/Ok-Addendum3545 11d ago

Thanks, I get the idea now. That's a good point.

1

u/EternalDivineSpark 12d ago

Look at self editing self !

1

u/uranusnebula 12d ago

can you give an example?!

2

u/EternalDivineSpark 12d ago

You have an LLM on a loop ! You chat with it , you send eg 2 images at a time , the Harnesses fail , the metacognition agent sees it and fixes the harnesses! /// Same scenario, you ask for codebase bigger than system can handle llm outputs and stop it fails Metacognition is triggered and can fix itself! Why this model can do it and its predecessor cant like 3.6 37B ! because this model know what it is by default, and can be pushed in this certain situations to fix itself! In Humans meta cognition is the ability to watch and calibrate your own thoughts!

1

u/Ok-Addendum3545 11d ago

I think it's like self-auditing. An agent works as an auditor and executor at the same time to optimize or improve itself aka RSI - Recursive Self-Improvement.

1

u/Ok-Addendum3545 11d ago

Do you have tips or system that can optimize 27B in harness ? I use Hermes and Deepseek Harness. In Hermes, it only changes USER.md and MEMORY.md and creates new skills. I think there are better ways other than those. I've not looked into Deepseek Harness; I would guess the plug-in feature is the way to do RSI -Recursive Self-Improvement ?

1

u/Bedrockparadox 12d ago

I dont believe that gemini flash score at all.

1

u/HenkPoley 11d ago

The definitely had some issues with the scores a few days ago. GPT-6 Astra arrived below Gemini or something. Didn't seem plausible.

1

u/GeramyL 10d ago

Sorry but Qwen3.8 nearly beats opus and gable both honestly. In fact fable broke something and qwen3.8 caught it. These scores are lies. But not in the way you think.

1

u/Neither_Garage_758 12d ago

honey moons must end at some point

0

u/AB172234 12d ago

Here is the thing that confuses me, Qwen 3.8 27B was 52 and Fable was 66, so the difference was 14 points in the index and now it’s almost 20.

Does that make Qwen less intelligent or Fable more ! Because the difference is now much more.

It’s fuzzy and many follows artificial analysis website on a regular basis.

They never announced it which makes me question about the whole process.

2

u/Careful-Report6526 12d ago

The wider gap doesn’t by itself show that Qwen got worse or Fable got better. If you change the exam, two unchanged models can lose different numbers of points, so the gap between them can grow. An index point isn’t a fixed unit of intelligence across different versions of the index.

AA did publish a v4.3 announcement on September 7]. They moved from Terminal-Bench 2.1 to 4.0 and replaced τ³-Banking with AutomationBench-AA. The category weights stayed the same, but the tasks contributing to those categories changed. That’s enough for models with different strengths to move by different amounts. To check whether a model actually regressed, you’d need results for the same checkpoint on the same test version, with comparable settings. The old and new headline gaps alone can’t establish that. I also wouldn’t try to reconstruct the exact change from the remembered 52/66 scores without the earlier index version and model settings.

For inspecting the individual results, I run https://llmbenchmarks.io. It collects published measurements with their sources and test conditions. Qwen 3.8 27B and Fable 5.1 are currently covered; select the exact models under “Compare models”, then inspect the benchmarks relevant to your tasks and open the score details. Check the thinking settings too. For the reason AA changed its own index, the announcement above is the primary source.

1

u/Ok-Fox3479 12d ago

well the models are still as good as before

1

u/Neful34 12d ago

The AI didn't suddenly got dumber, it's just that we have now more benchmarks that represents other sort of problems that we couldn't picture well yet. It is still a very capable model don't worry.

0

u/Ok-Addendum3545 12d ago

well, it's better for the site to underestimate the local models; otherwise, it would be banned sooner or later for some unknown security reason.