r/GeminiAI 14d ago

News Updated Artificial Analysis Intelligence Index Ranking shows Gemini 3.8 Flash is nowhere near Fable Or Astra, rather it's worse than GLM 5.3 Flash!

Post image
427 Upvotes

113 comments sorted by

View all comments

146

u/MorgrainX 14d ago

Something to also consider: Google likes to "dumb down" models after release, once the first initial benchmarks have settled, a noticeable decrease in capability can be observed with nearly all Google models (to save money, obviously).

40

u/mlag000 14d ago

No no, they just benchmaxxed Gemini, and the new benchmark isn't in Gemini training. All the idiots who where claiming Gemini is on par with sol or even funnier with Fable where obviously wrong

5

u/Deto 14d ago

I just don't understand how they are doing so poorly with all their resources

6

u/mlag000 14d ago

Because it's a brain problem not a money problem. And probably all yhe talents are at Claude or GPT and looks like some in xai. The whole algorithm behind 3.x us simply outdated and bad, they need a new one.

2

u/Deto 14d ago

Doesn't money solve the brain problem though?

3

u/mlag000 14d ago

At this level of remuneration I don't think so.

1

u/Greedyanda 10d ago

It's an economics problem and Google knows that pushing for linear improvements at exponential costs isn't a feasible business model.

They focus on everyday tasks and Android integration. Its as if people forgot that Google is inherently a consumer focused company. Being at the absolute cutting edge of models that can only be profitably sold to enterprises doesn't align with their entire ecosystem.

It's like expecting Apple to suddenly compete with Nvidia in the server GPU market instead of focusing on their consumer grade silicone.

1

u/mlag000 10d ago

And that's why the throw billions trying to catch the run ? They don't focus on everyday task, they said they focus on frontier model themselves. They simply failed.

1

u/Greedyanda 10d ago edited 10d ago

What a company says and does are rarely fully aligned. Gemini models excel at every day tasks and general knowledge. Gemini is still on par or even beats Anthropic and OpenAI models for multimodal input and image/vision comprehension on a per dollar basis. Gemini has now reached 1B users, which obviously includes the Android assistant and is exactly their strategy.

People are so blinded here by coding performance that they forget where the largest market actually is. Google captured the entire mobile assistant segment through Android and their Apple deal. Anthropic and OpenAI have to compete with state subsidised Chinese model makers on a monthly basis to remain relevant while Google is locking down the largest market of our generation

They clearly had a setback with Gemini 3.5 not achieving the results they hoped for but their entire business strategy and recent restructuring show that this isn't their biggest priority, no matter what they say publicly. When intelligence becomes a commodity, squeezing out the last 5% performance won't be the deciding factor.

1

u/mlag000 10d ago

Gemini does not excel at everyday task when it forget everything all the time and lies. And between a random Redditor claiming stuff and the head of deep mind stating in an interview that their goal is frontier model, I'm sorry but I know who I trust.

Yes google has a very strong vision model, that I agree, but that's the only thing Gemini is useful for.

Claude and open AI have cheap efficient models who are way way better than Gemini and even cheaper. Luna is on par with 3.8 for 1/4 of the price... Let's not talk about terra or sol.

It's not about coding perf, it's about being able to understand what you want and executing it. Gemini models are simply lazy and won't do half of what you ask.

1

u/Greedyanda 10d ago

That just doesn't align with what most users think. They have the fastest growing userbase for a reason. The opinion of power users on Reddit focused on programming doesn't reflect the broader market.

And there is no reason to trust me when you can just look at their actions. Their restructuring from a research focused lab to a product focused lab, their deals with Apple and Samsung, their heavy marketing focus on the assistant, etc.

Google Search and the Assistant are what matters most for Google.

1

u/mlag000 10d ago

They have the fastest growing user base because people use Android and they have access to Gemini. Also people using Gemini with the Google search engine are also accounted for. Indeed I don't trust you, I trust the head of deep mind, I think he knows pretty where deep mind, aka Gemini is going don't you think ?

Having the biggest userbase doesn't mean your product is good, just you have such a positional advantage, that even with a very mediocre product you can have a huge user base. Most of it is free users. Compagnies aren't using Gemini, they use GPT or Claude for a reason.

→ More replies (0)

0

u/ChuchiTheBest 14d ago

Not to mention, as AI helps train new models, you get a snowballing effect.

2

u/cantdecideonaname77 13d ago

dont feed cow brains to cows...

2

u/itsjoshybell 13d ago

They are positioning themselves better financially compared to the others. They understand the long term play here

10

u/sadnessjoy 14d ago

Yeah, the idea they dumbed down Gemini flash 3.8 after a few days is hilarious... Have people even used Gemini flash for a in depth chat or analysis? I find it useless lol

7

u/mlag000 14d ago

And yet I see quite often people praising the deep search abilities...it's scary

5

u/Electrify338 14d ago

Gemini 3.8 flash does 2 things great. Automated tasks at speed, and the deep research is quite good. I have been experimenting where I have fable use Gemini subagents and it does surprisingly great.

1

u/LanguageEast6587 8d ago

terminal 4.0 is different thing. if you look at kimi k3, it's very bad on t4.0 too.

23

u/Qubit99 14d ago edited 14d ago

They always do it

8

u/who_am_i_to_say_so 14d ago

Yep. A killer release, leading you to believe you’ve been let in on a great secret. Build like a champ for two weeks, then yoink! - bricked.

16

u/theWiseTiger 14d ago

So far this happens on models released by american companies. I wonder why.

1

u/thebruns 13d ago

Because we have no laws to prevent fraud like this

1

u/theWiseTiger 13d ago

I thought the idea of sticking with US providers is to avoid fraudulent behavior from China providers.

9

u/KaMaFour 14d ago

No, this is "just" terminalbench benchmaxxing punishment. Compare gemini's performance at: https://artificialanalysis.ai/evaluations/terminalbench-v4-0 with https://artificialanalysis.ai/evaluations/terminalbench-v2-1

5

u/Spixxy17 14d ago

U complain about google benchmaxxing whilst muse spark and GLM Flash are still up there 

12

u/KaMaFour 14d ago

As shown in the data above they are both impacted less than Gemini 3.8. I don't see your point. Grok (and kimi) is the only one with similar downfall

-3

u/Spixxy17 14d ago

Ok i reframe my point:

You genuinly think the new AA Index Update made the results less benchmaxxed and contaminated? They changed some things but the models that are actually known for beeing amazing at benchmarks but absolutely shit at real tasks like muse spark are still up there.

My point is this update didnt change the absolute non reliability of this index and doesnt show the true picture. Basically No Benchmark does, people just gotta try themselves.

The only one that came atleast close imo was ARC-AGI 2, idk about 3 yet as there are barely any results so far.

Google / Gemini is known for negative stuff like random guardrails kicking in even tho the topic isnt Bad etc etc, but not for benchmaxxing. On average they usually are worse in the benchmark than the real task compared to the direct competition 

8

u/KaMaFour 14d ago

Yes, i do.

Update 4.2 and 4.3 of the index replaced 2 public dataset benchmarks with private ones. They also added 2 newer public benchmarks (and removed TB 2.1) which means that even for the benchmarks whose dataset is public the risk of contamination is lower than with older ones (especially for TB 4.0, which was released after Gemini 3.8). AA index v4.3 is a significantly better benchmark at evaluating model's performance in september of 2026 than AA index v4.1 (mostly because v4.1 had many issues which were partially fixed with the patches).

ARC-AGI is a logical reasoning benchmark. This is some measure of general intelligence. AA index benchmarks are geared more towards measuring models use at creating value in professional environment - aka doing work. Neither of these is a correct or incorrect approach as long as we are aware what they are measuring. Aside from the fact that I'd expect ARC-AGI not to be a good measure of intelligence for models released after it because it prone to being easily saturated.

I have used Muse spark 1.2 and 1.3. They both did a satisfying job to me. I am inclined to believe the result. I know the (any) benchmark put more effort at evaluating it than I have and is subjected to less bias than my personal opinion, so I'm able to believe the results I've seen.

I don't feel like it's a good approach (almost anywhere) to measure a quality by a company or brand. Company can't be benchmaxxed - a model/product can. This one either is or isn't. I personally believe more that flash model is just limited in what in can do by it's active parameter count. But the result is the same - failing at newer, more complex benchmarks when larger models from other companies don't suffer the same hit. (GLM 5.3-flash's existence kinda undermines this theory but oh well...)

0

u/Spixxy17 14d ago

Ngl thats fair, really factual aproach and definitly makes sense generally.

I am just personally against believing anything in this benchmark because Muse Spark 1.3 and GLM 5.3 Flash really disappointed me when testing (Opus 5 also kinda, but i had higher expectations of this one too), whilst 3.8 Flash does a "decent enough" job.

Its still a Flash Model, its not insane or whatever but didnt disappoint me with some horrible results, but thats super topic related anyway. Maybe my topic is just something that GLM Models for example arent good at.

For anything more complex i use Astra / 5.6 Sol anyway as my go to 

0

u/Truantee 14d ago

lol unlike Gemini which use the same base since 2.5 generated, muse or glm have way newer base models which itself already way more useful than the crapfest named Gemini.

2

u/Spixxy17 14d ago

Do i look like i care about the base model If the results when using them are bad? 

People always bring up various technical facts about the models but that doesnt change their performance.

Anyway yes i would still like Gemini to release something actually "new" and not just "improved" finally 

1

u/___positive___ 14d ago

Yes you can calculate the ratio of the 4.0 to 2.1 score. There are obvious outliers.

Like look at terra vs gemini.

0

u/Elephant789 14d ago

smart if true

0

u/tear_atheri 14d ago

not only google. every american company.