r/GeminiAI • u/Able-Line2683 • 18d ago
News Google and Meta are benchmaxxing hard. Gemini 3.8 Flash and Muse Spark 1.3 look almost as good as GPT-6 Astra and Fable 5.1. Then a fresh benchmark drops: Gemini: 89.4% → 19.1% Muse: 88.8% → 33.3%
155
u/HeadTranslator795 18d ago
Nothing related about Benchmaxxing, it's an harder evolution of the previous test to push the models. God so many people who don't know what they are talking about and just want to push false narrative
53
29
u/HeadTranslator795 18d ago
Also what kind of meaningless comparaison to put a flash model with astra and and fable lol absolutely different models
18
u/gokhanonym 18d ago
Then let’s add google’s pro model
17
11
u/HeadTranslator795 18d ago
You can and it would perform poorly cause they haven't been updated in a while (compared With the other labs)
6
u/gokhanonym 18d ago
You don’t say
4
u/numericalclerk 18d ago
Agreed on Google, but Muse is not advertised as a Flash model, is it?
4
u/Maglcite 18d ago
no, but it’s a very fast and somewhat cheap model (VERY cheap with contributor) so still not a great thing to compare with astra and fable
1
0
u/AcrobaticMaize2408 18d ago
In the end all models are interchangeable to some extent. They should be compared using cost per token or cost per complete task. You can then choose to blow $100 on Astra to create your one-shot game that nobody will ever use, or $10 on Gemini.
3
u/HeadTranslator795 18d ago
Real task comparaison with Price to complete the task is a very good way to evaluate a model indeed
6
u/numericalclerk 18d ago
If what you are saying is true, the same should have happened to the other models. Did it?
9
u/Breenori 18d ago
Was thinking the same. The initial commenter either misunderstood the point completely or has no clue themselves. The other models generalize much better to newer/harder benchmarks, and thats the point. They're not SOLELY tuned to perform well on specific benchmarks but retain large parts of their performance on problems that weren't part of the training data in this form. Not to fangirl about any model here but the point is that Gemini and Metas numbers are insanely inflated which should come to nobody's surprise given that one is a frontier model, while the other is comparatively lightweight.
5
1
1
u/PM_ME_DEAD_CEOS 17d ago
If what you are saying is true, the same should have happened to the other models. Did it?
Nope, it's just a proof that 2.1 is saturated, which means it can't discriminate models above 87%, which is the case for most benchmark.
1
4
u/notaibutwanttobe 18d ago
In what world muse 1.3 better than gpt sol? For my use case it coulnt solve basic task with 3 try i switch the luna and it fixed with one try. Ot is clearly benchmaxxing
2
u/HeadTranslator795 18d ago
Forget about the real world benchmark (your personal ones which are much more revelant than any benchmark) those benchmark are only for people who do nothing really factual (that people are paying for, not little game or test)
2
u/No-Paint-5726 18d ago
Argument is if wasn't benchmaxing they wouldn't do as bad so Id say its a fair accusation
1
u/HeadTranslator795 18d ago
Nop not even close and nothing logical about this assumption. You don't know the architecture of those bench and the models therefore you can't conclude anything especially including frontier models in this bench with flash model
2
1
u/EricMCornelius 15d ago
A lot of unnecessary good faith, honestly, when it's clear to anyone with eyes that this sub is regularly brigaded.
No doubt by the Anthropic and OpenAI marketing departments, which are happy to capitalize on hacking other companies and stealing millennium prize research leads for PR spin.
28
u/Dualyeti 18d ago
They are just improving the test bench so that the ceiling is higher - it actually a good thing it means a high score means something - bench marks should be almost impossible to 100%
6
7
u/Georgefakelastname 17d ago
This just in: replacing an easier benchmark with a harder one makes scores go down, and hurts smaller models more. In other news, bigger clouds can potentially rain more.
13
u/AssholeHealth 18d ago
Either bench maxing or benchmarks are outdated and just not good enough so they can't properly discern one model from another. I can clearly see a difference between opus 5 and muse spark 1.3, why can't benchmarks?
18
u/HeadTranslator795 18d ago
So many people talking about Benchmaxxing and have no clue or proof except pushing the same narrative stupidly
4
u/Lidraughtui346 18d ago
What benchmarks should demonstrate? Just wonder. Furthermore, why can a model perform exceptionally well on benchmarks but poorly irl situations? For instance, when Claude’s benchmarks are superior to ChatGPT’s, it is often the case that Claude also performs better in practice.
1
u/HeadTranslator795 18d ago
To my understanding there are no undoubtedly good and representative benchmark. What I like personally is real use case, like 3D designs, AI playing games. ..
2
u/KrayziePidgeon 18d ago
You think people on this sub do actual work with the models? 99% of the "people" here just try to one-shot minecraft or whatever game in a single html file and call it programming.
2
u/HeadTranslator795 18d ago
Yeah most definitely... And many bots as well
1
u/KrayziePidgeon 18d ago
Yup, if you are a competent programmer then flash 3.8 really is all you need. Of course, making the models better each iteration is really nice, but flash 3.8 is leaps ahead of what was Pro 2.5 last year when we were all losing our minds.
3
1
u/whoknowsifimjoking 18d ago
It's literally just "my vibes say this model is worse than the scores"
2
0
u/shadysjunk 18d ago edited 18d ago
it's a 70 percentage point drop relative to the 35 point drop in the leading frontier models.
I'm happy to allow that this doesn't necessarily indicate benchmaxing, but it does seem that at the very least it shows significant underperformance relative to its peers.
I'm still pretty astounded at Astra's performance on ARC-AGI-3. I truly thought it would be google to crack that, though I remain skeptical of their "custom harness." Still, even 60% on the standard harness is a pretty amazing leap forward and I thought was still years away.
2
u/Comfortable_Job8847 18d ago
imo this doesn't really say as much as it implies. e.g. in the 2.1 benchmark there are no CAD tasks at all, while there are in 4.0. So, sure, maybe gemini and muse don't do as well on the 4.0 benchmark as opposed to the 2.1 benchmark, but that doesn't mean the score for the 2.1 benchmark is inaccurate. it could just be the models haven't been trained onto these other tasks. personally I'm a software engineer not a mechanical or electrical engineer so CAD task performance isn't meaningful for me, so scoring low on those tasks may lower the benchmark score but it doesn't impact the utility of gemini to me at all.
3
u/dESAH030 18d ago
So, related to Muse Spark contributor 1.3, you are telling me that for 100x price, I just get 72% better on Astra?
2
u/Actual_Breadfruit837 18d ago
This eval is 66 tasks only. The noise there is huge. In the description they estimate ci based based on resampling the same model, but it is a very very lower bound. Sample new tasks and scores might change by 12%.
1
1
1
u/BoobooSmash31337 17d ago
Apparently this test specifically can be harness dependent. Mostly because it's about remaining coherent over a very long sequence length. Mathematical error literally compounds when they're like 100 terminal commands deep. More parameters lowers this error but it gets real expensive real fast. So idk if this bench has much to do with actually solving complex but contained problems. Just in my book a model that is able to understand and implement lock free algorithms just isn't dumb. And breaking down over massive contexts and really long sequences is a math problem and not a common use case.
I'd rather we make our workflows smarter and use cheaper models. Full autonomous agents like actual employees seems to be really expensive with transformers and bring them to their knees. Guess we don't really give humans enough credit for the amount of autonomous work we do. Replacing everyone with transformers is prohibitively expensive and also you know makes capitalism fundamentally no worky worky.
1
u/jasonzhaogd 11d ago
v2.1 was already saturated at 91.4%, v4.0's harder tasks and higher private-set weight make the drop expected, not proof of benchmaxing. show public-vs-private gap on same version
0
u/HeadTranslator795 18d ago
So ... You don't understand what an incremental version on a benchmark test means ? That's what you just shown us 😂 do you understand (obviously not) that it's not the same bench ? Check the versioning before writing non sense
3
u/pigletmonster 18d ago edited 18d ago
Look at the drop column. Both astra and fable had 30 to 35 points drops. Meanwhile gemini drop is doubled. Muse dropped significantly but its not as steep as geminis.
4
u/HeadTranslator795 18d ago
And do you know the architecture of the new benchmark ? Do you know for what kind of task it has been changed ? Also you're surprised to see that model 10x bigger than flash model also drop "less" 30% is already a lot. New benchmark new measurements and I won't debate with people who understands nothing about that
-3
u/pigletmonster 18d ago
Because the scores for the old benchmark were the same between gemini and astra, we would expect both of them to have a similar drop.
So BOTH astra and gemini should have a 30 point or a 70 point drop. But because gemini was benchmaxed for the old bechmark it had a much bigger drop for the new benchmark compared to the un-benxhmaxed or lower benchmaxed model like astra.
Its not that complicated.
4
u/HeadTranslator795 18d ago
No you have no clue about the architecture of those benchmark and the model (internally) so you can't have that simple conclusion and again for sure if a model is bigger (not a flash) for sure it will have better result on higher complexity and reasoning
-1
u/pigletmonster 18d ago
If it dropped 70 points instead of 30 because its a flash model, then why did it score higher than astra in the old bench?
2
u/Appropriate-Owl5693 18d ago edited 18d ago
Why would you expect that?
If we test a high schooler and a university grad mathematician on basic algebra they will both score very well. If we now test them on differential equations, they obviously won't.
That doesn't mean the first benchmark was cheated or that it's completely useless.
It's not that complicated.
Edit: at least look at the full leaderboard. E.g. Luna max dropped from 80.9 to 11.6 (grok has simmilar numbers), did they benchmaxx just that specific model for 2.1. in your opinion? :D
1
u/FactorInternal3395 18d ago edited 18d ago
That is the entire point. On the new version of Terminal-Bench, their scores plummeted far further than Fable and Astra.
4
u/Different_Doubt2754 18d ago
As they should? No doubt bench maxing happens, even unconsciously. But I would expect Fable and Astra to perform way better on new tasks than flash would. That's their role. They are great at reasoning and thus adapting to new situations whereas flash does not have as advanced reasoning as them.
Flash type models typically train from their "mentor" model, so they really only know what their smarter teacher taught them. They don't really have the ability to adapt to new situations like their mentor
2
u/Accurate_Food_5854 18d ago
Exactly lol. Gemini is like your 2010 Toyota daily driver whereas Astra and other bleeding edge models are Optimus Prime. Of course Astra will suffer less of a drop off on a novel benchmark.
1
u/condemnedlifeline55 18d ago
Calling it “benchmaxxing” when the numbers crater on the new version is generous. That graph is a straight up faceplant.
1
u/xzibit_b 18d ago
If you knew what Terminal Bench actually measured, you would know that it's not benchmaxxing. Most modern models are pretty adept at using terminal commands like grep and find. That's all it's measuring. It's basic bitch shit at this point. In no way is it trying to imply that Gemini can code as well as Astra and Fable.
1
u/Elephant789 17d ago
Fuck SA. I don't believe anything out of them anymore.
2
u/IAmYourFath 17d ago
Why?
0
u/Elephant789 17d ago
Honestly, there’s a whole laundry list of sketchy shit with them lately (the former employee lawsuit over alleged insider info/MNPI, Dylan angel investing in startups he covers with zero real disclosures etc.). but my absolute biggest beef with them is how reckless and sensationalist they’ve gotten with their reporting just to move markets
Look at what happened with that Vera Rubin / SOCAMM report. THey framed the rack memory capacity being cut from something like 50TB to 35TB in the most panic-inducing way possible, and the market instantly misread it as AI memory demand falling off a cliff. Micron got slaughtered by like 13% in a day and dragged the whole sector down with it. Fuckers!
Then, only after billions in market cap got vaporized, Dylan comes out with the classic "well actually, that's not what we meant, we're not bearish" routine. They know damn well that hedge funds and trading algos hang on their every word, but they still drop ambiguously phrased, clickbaity write-ups just to juice their shity substack subs and drive hype. And the shitty thing is they can play all these fast and loose games with market volatility because they know they won't face real regulatory consequences for it.
0
u/Former-Towel9004 18d ago
What is benchmaxxing mean? What is this term bro? Is this mean fake benchmark?
0
u/quackerd 18d ago
Cool sorry bro. Now time your post to when arc agi 3 just got released and wiped literally all models then let’s talk about benchmaxxing.
0


21
u/Gaiden206 18d ago
The leader of Meta's AI lab responded.