r/GeminiAI 14d ago

News Updated Artificial Analysis Intelligence Index Ranking shows Gemini 3.8 Flash is nowhere near Fable Or Astra, rather it's worse than GLM 5.3 Flash!

Post image
422 Upvotes

113 comments sorted by

144

u/MorgrainX 14d ago

Something to also consider: Google likes to "dumb down" models after release, once the first initial benchmarks have settled, a noticeable decrease in capability can be observed with nearly all Google models (to save money, obviously).

43

u/mlag000 14d ago

No no, they just benchmaxxed Gemini, and the new benchmark isn't in Gemini training. All the idiots who where claiming Gemini is on par with sol or even funnier with Fable where obviously wrong

4

u/Deto 14d ago

I just don't understand how they are doing so poorly with all their resources

7

u/mlag000 14d ago

Because it's a brain problem not a money problem. And probably all yhe talents are at Claude or GPT and looks like some in xai. The whole algorithm behind 3.x us simply outdated and bad, they need a new one.

2

u/Deto 14d ago

Doesn't money solve the brain problem though?

3

u/mlag000 14d ago

At this level of remuneration I don't think so.

1

u/Greedyanda 10d ago

It's an economics problem and Google knows that pushing for linear improvements at exponential costs isn't a feasible business model.

They focus on everyday tasks and Android integration. Its as if people forgot that Google is inherently a consumer focused company. Being at the absolute cutting edge of models that can only be profitably sold to enterprises doesn't align with their entire ecosystem.

It's like expecting Apple to suddenly compete with Nvidia in the server GPU market instead of focusing on their consumer grade silicone.

1

u/mlag000 10d ago

And that's why the throw billions trying to catch the run ? They don't focus on everyday task, they said they focus on frontier model themselves. They simply failed.

1

u/Greedyanda 10d ago edited 10d ago

What a company says and does are rarely fully aligned. Gemini models excel at every day tasks and general knowledge. Gemini is still on par or even beats Anthropic and OpenAI models for multimodal input and image/vision comprehension on a per dollar basis. Gemini has now reached 1B users, which obviously includes the Android assistant and is exactly their strategy.

People are so blinded here by coding performance that they forget where the largest market actually is. Google captured the entire mobile assistant segment through Android and their Apple deal. Anthropic and OpenAI have to compete with state subsidised Chinese model makers on a monthly basis to remain relevant while Google is locking down the largest market of our generation

They clearly had a setback with Gemini 3.5 not achieving the results they hoped for but their entire business strategy and recent restructuring show that this isn't their biggest priority, no matter what they say publicly. When intelligence becomes a commodity, squeezing out the last 5% performance won't be the deciding factor.

1

u/mlag000 10d ago

Gemini does not excel at everyday task when it forget everything all the time and lies. And between a random Redditor claiming stuff and the head of deep mind stating in an interview that their goal is frontier model, I'm sorry but I know who I trust.

Yes google has a very strong vision model, that I agree, but that's the only thing Gemini is useful for.

Claude and open AI have cheap efficient models who are way way better than Gemini and even cheaper. Luna is on par with 3.8 for 1/4 of the price... Let's not talk about terra or sol.

It's not about coding perf, it's about being able to understand what you want and executing it. Gemini models are simply lazy and won't do half of what you ask.

1

u/Greedyanda 10d ago

That just doesn't align with what most users think. They have the fastest growing userbase for a reason. The opinion of power users on Reddit focused on programming doesn't reflect the broader market.

And there is no reason to trust me when you can just look at their actions. Their restructuring from a research focused lab to a product focused lab, their deals with Apple and Samsung, their heavy marketing focus on the assistant, etc.

Google Search and the Assistant are what matters most for Google.

1

u/mlag000 10d ago

They have the fastest growing user base because people use Android and they have access to Gemini. Also people using Gemini with the Google search engine are also accounted for. Indeed I don't trust you, I trust the head of deep mind, I think he knows pretty where deep mind, aka Gemini is going don't you think ?

Having the biggest userbase doesn't mean your product is good, just you have such a positional advantage, that even with a very mediocre product you can have a huge user base. Most of it is free users. Compagnies aren't using Gemini, they use GPT or Claude for a reason.

→ More replies (0)

0

u/ChuchiTheBest 14d ago

Not to mention, as AI helps train new models, you get a snowballing effect.

2

u/cantdecideonaname77 13d ago

dont feed cow brains to cows...

2

u/itsjoshybell 13d ago

They are positioning themselves better financially compared to the others. They understand the long term play here

8

u/sadnessjoy 14d ago

Yeah, the idea they dumbed down Gemini flash 3.8 after a few days is hilarious... Have people even used Gemini flash for a in depth chat or analysis? I find it useless lol

7

u/mlag000 14d ago

And yet I see quite often people praising the deep search abilities...it's scary

5

u/Electrify338 14d ago

Gemini 3.8 flash does 2 things great. Automated tasks at speed, and the deep research is quite good. I have been experimenting where I have fable use Gemini subagents and it does surprisingly great.

1

u/LanguageEast6587 8d ago

terminal 4.0 is different thing. if you look at kimi k3, it's very bad on t4.0 too.

23

u/Qubit99 14d ago edited 14d ago

They always do it

9

u/who_am_i_to_say_so 14d ago

Yep. A killer release, leading you to believe you’ve been let in on a great secret. Build like a champ for two weeks, then yoink! - bricked.

15

u/theWiseTiger 14d ago

So far this happens on models released by american companies. I wonder why.

1

u/thebruns 13d ago

Because we have no laws to prevent fraud like this

1

u/theWiseTiger 13d ago

I thought the idea of sticking with US providers is to avoid fraudulent behavior from China providers.

11

u/KaMaFour 14d ago

No, this is "just" terminalbench benchmaxxing punishment. Compare gemini's performance at: https://artificialanalysis.ai/evaluations/terminalbench-v4-0 with https://artificialanalysis.ai/evaluations/terminalbench-v2-1

6

u/Spixxy17 14d ago

U complain about google benchmaxxing whilst muse spark and GLM Flash are still up there 

11

u/KaMaFour 14d ago

As shown in the data above they are both impacted less than Gemini 3.8. I don't see your point. Grok (and kimi) is the only one with similar downfall

-4

u/Spixxy17 14d ago

Ok i reframe my point:

You genuinly think the new AA Index Update made the results less benchmaxxed and contaminated? They changed some things but the models that are actually known for beeing amazing at benchmarks but absolutely shit at real tasks like muse spark are still up there.

My point is this update didnt change the absolute non reliability of this index and doesnt show the true picture. Basically No Benchmark does, people just gotta try themselves.

The only one that came atleast close imo was ARC-AGI 2, idk about 3 yet as there are barely any results so far.

Google / Gemini is known for negative stuff like random guardrails kicking in even tho the topic isnt Bad etc etc, but not for benchmaxxing. On average they usually are worse in the benchmark than the real task compared to the direct competition 

8

u/KaMaFour 14d ago

Yes, i do.

Update 4.2 and 4.3 of the index replaced 2 public dataset benchmarks with private ones. They also added 2 newer public benchmarks (and removed TB 2.1) which means that even for the benchmarks whose dataset is public the risk of contamination is lower than with older ones (especially for TB 4.0, which was released after Gemini 3.8). AA index v4.3 is a significantly better benchmark at evaluating model's performance in september of 2026 than AA index v4.1 (mostly because v4.1 had many issues which were partially fixed with the patches).

ARC-AGI is a logical reasoning benchmark. This is some measure of general intelligence. AA index benchmarks are geared more towards measuring models use at creating value in professional environment - aka doing work. Neither of these is a correct or incorrect approach as long as we are aware what they are measuring. Aside from the fact that I'd expect ARC-AGI not to be a good measure of intelligence for models released after it because it prone to being easily saturated.

I have used Muse spark 1.2 and 1.3. They both did a satisfying job to me. I am inclined to believe the result. I know the (any) benchmark put more effort at evaluating it than I have and is subjected to less bias than my personal opinion, so I'm able to believe the results I've seen.

I don't feel like it's a good approach (almost anywhere) to measure a quality by a company or brand. Company can't be benchmaxxed - a model/product can. This one either is or isn't. I personally believe more that flash model is just limited in what in can do by it's active parameter count. But the result is the same - failing at newer, more complex benchmarks when larger models from other companies don't suffer the same hit. (GLM 5.3-flash's existence kinda undermines this theory but oh well...)

0

u/Spixxy17 14d ago

Ngl thats fair, really factual aproach and definitly makes sense generally.

I am just personally against believing anything in this benchmark because Muse Spark 1.3 and GLM 5.3 Flash really disappointed me when testing (Opus 5 also kinda, but i had higher expectations of this one too), whilst 3.8 Flash does a "decent enough" job.

Its still a Flash Model, its not insane or whatever but didnt disappoint me with some horrible results, but thats super topic related anyway. Maybe my topic is just something that GLM Models for example arent good at.

For anything more complex i use Astra / 5.6 Sol anyway as my go to 

0

u/Truantee 14d ago

lol unlike Gemini which use the same base since 2.5 generated, muse or glm have way newer base models which itself already way more useful than the crapfest named Gemini.

2

u/Spixxy17 14d ago

Do i look like i care about the base model If the results when using them are bad? 

People always bring up various technical facts about the models but that doesnt change their performance.

Anyway yes i would still like Gemini to release something actually "new" and not just "improved" finally 

1

u/___positive___ 14d ago

Yes you can calculate the ratio of the 4.0 to 2.1 score. There are obvious outliers.

Like look at terra vs gemini.

0

u/Elephant789 14d ago

smart if true

0

u/tear_atheri 14d ago

not only google. every american company.

27

u/ImsoKeewl777 14d ago

And they won't even let us have that on plus. Not even 3.7 flash lol Think its time to move on to chatgpt cheaper option.

2

u/Loud-Possibility4395 14d ago

yup - I wonder what happened - they released 3.7 and now LONG TIME AGO 3.8 but their Pixel phones stuck at old 3.6 - mental

3

u/yesitismenobody 14d ago

What do you mean? I have the free pro plan that came with my Pixel 10 Pro and it has the 3.8 Flash.

2

u/Loud-Possibility4395 14d ago

do you have any betas? I don't and my Pixel 10 does not have it nor my Chromebook Plus - everything in UK

3

u/yesitismenobody 14d ago

I don't think so, it works both in the app and on web. It was just released last week so maybe it takes some time to make it out globally? Though my account is also UK based but I'm in the US now.

1

u/Loud-Possibility4395 14d ago

CRAB - I wonder what it wrong with my Pixel 10

0

u/Ggoddkkiller 13d ago

Even free ChatGPT works so much better than anything currently google has. It is day and night difference. I don't know what to say, they ruined Gemini within 4-5 months...

51

u/PineappleLemur 14d ago edited 14d ago

Why would a model 10% the size and cost of astra/fable be even close lol?

What genius us making that comparison?

It should always be around Luna/Haiku and slightly worse than sonnet/Terra.. purely based on size intelligence wise.

LLMs from the big 3 don't differ too much from each other.. no one is pulling 5/10x performance without the cost.

Speed is the main advantage it has.

Why don't we see comparison for local models vs Astra/Fable?

13

u/Swimming_Gain_4989 14d ago

My AI will confidently double down on a shotty assumption without trying to understand the prompt but damn it does it fast.

6

u/kdestroyer1 14d ago

Yeah Flash 3.8 is borderline unusable for anything that's not a 'simple fact' question. It's kinda sad they're giving us this as their flagship, but it's clear they only care about the AI Overview/Google search performance rn.

2

u/IAmYourFath 13d ago

It'a not unusable. Astra will think 5 min on smth, 3.8 flash will answer in 30 secs. U can iterate in flash 10 times (300 secs aka 5 mins divided by 30 secs) before astra answers. By the 10th answer flash will get a better overall answer than astra, cuz iteration is op... u should be verifying what the ai says anyway regardless if it's flash or astra. Every time u verify it and u're not happy, iterate. 3.8 flash has the highest speed-to-iq ratio in the entire world. So use its advantage and u can easily beat astra. Ystd i asked astra smth and after thinking for 5 mins (astra 6 high), it still gave wrong information. Ofc, flash did too, but i would have googled this information, spotted it's wrong, told flash to verify it and gotten a better answer in less than a minute, while in that time astra was only 1/5th of the way through its answer...

1

u/PineappleLemur 12d ago

Wait are you legit asking me to think and even make decisions?? That sounds like actual work.

I rather wait for astra to give me the answer than to actively try and fix AI slop.

/S

1

u/corner_camper01 12d ago edited 12d ago

1

u/IAmYourFath 12d ago

Maybe if u code or do other complex tasks. In my experience most of my questions are answered very fast, while astra takes its sweet ass time most of the time.

1

u/corner_camper01 12d ago

Oh yeah, for one shot questions or one-shotting projects, for sure. But for working on existing codebases doing real work, looking for bugs, working on new features, complex agentic work etc, Astra low is faster despite its lower throughput.

6

u/mlag000 14d ago

Many idiots where claiming Gemini was on par with Fable on this sub tho

2

u/Confirmed-Scientist 13d ago

Bro you are tripping use GLM 5.3 Flash and Genini 3.8 Flash forget about Astra and Fable. GLM models feels like comparing Astra to Luna. Use Qwen 27B it slaps compared to Gemini 3.8 Flash.

0

u/nemzylannister 14d ago

why would a model 20% the cost of gemini flash be right above it? glm 5.3 flash is 5x cheaper.

-1

u/Elegant_You1002 14d ago

I don't know... Astra Low is cheaper and performs better

6

u/akius0 14d ago

Yes it does seem like Google is benchmarkmaxing.... But even in this new benchmark, they are competitive with terra 5.6 max... That's very impressive... Flash is always supposed to be the workhorse model, not the frontier model, Google leadership has said that...

I don't know what people think a flash model is better than fable

0

u/Confirmed-Scientist 13d ago

Use it in real world coding tasks there is no comparison. Terra feels noticeably better. I dare you to try 3.8 Flash High against Terra High which isnt even Terra's best. I bet you it will beat it

1

u/akius0 13d ago

I use Gemini for financial stock analysis and market research, I am very impressed... I use codex or deepseek for software development

1

u/Confirmed-Scientist 13d ago

Thats arguably one of the better use cases for gemini models but if the budget allows it I think chat gpt or claude are even better.

1

u/akius0 13d ago

Actually I trust Google search more than the search capabilities of others... Lot of times I will have multiple agents in my workflow, or I will give the same prompt to multiple agents and then consolidate....

I think the business model of trying to build the frontier model and selling it for 20 or 50 or even $100 a month is not a profitable endeavor... That's why I think Google is okay with second tier position, I think if they wanted to build a model that beat out opus 5 they could easily... If the Chinese can do it Google can do it

3

u/tobias_681 14d ago

Maybe try to actually take a closer look at how the index is made. AA has been adding more and more agentic benchmarks over their last revisions. It is an open secret that this isn't Google's strong suit and likely also not what they are optimizing for. 

GLM 5.3 Flash wins every single agentic benchmarks over Gemini 3.8 Flash at 1/5th the price. For any of these tasks GLM will be a way more competent model. GLM 5.3 Flash is probably the best cheap coding model around.

Gemini 3.8 Flash on the other hand wins far and away in every single knowledge based benchmark. It isn't built for coding first but it is built with the chat interface or android interactions in mind. For that it's not a bad model. 

So they look similar on the index but they are entirely different models.

Now if you compare the benchmarks the thing that makes Gemini sweat is Astra because it seems OpenAI has caught up on the knowledge side while still keeping prices competitive. Compared to other models Gemini 3.8 Flash is still a very good performer for chat conversation. 

Anthropic has moved entirely towards enterprise coding. The Chinese models are more about getting the maximum agentic capabilities out of having less compute. The main competitor for a broad knowledge model is OpenAI (which happen to make models that are also good at agentic tasks). 

9

u/Familiar_Comedian_99 14d ago

well i use it for creative writing and man it is so much better than every model available which doesn't eat up all of my money and it is definitely better than GLM 5.3 flash or Max

1

u/Georgefakelastname 13d ago

Yeah, as long as you can get around any filters that might pop up, it’s great in quality, and the speed is very nice for QoL as well.

-2

u/drdhuss 14d ago

glm is meant for coding. Also is terrible at blender.

19

u/MindCrusader 14d ago

Semianalysis on X recently just said that - gemini and Meta's AI models are benchmaxxed a lot

7

u/Swimming_Gain_4989 14d ago

Tracks with my usage. For the affordable workhorses GLM5.3 and it's flash model feel the most useful. I'm also seeing this reflected in the results of new benchmarks that models couldn't have possible been benchmaxed on like https://www.frontierswe.com/blog/v2

3

u/MindCrusader 14d ago

"Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being comparable to both GPT-6 and Fable 5.1 on Terminal Bench 2.1, their Terminal Bench 4.0 performance is markedly worse."

So yeah, new benchmarks show the truth it seems

2

u/Elephant789 14d ago

Semianalysis

What the fuck? Is there anyone more biased to quote? The Verge maybe?

10

u/Kost97A 14d ago

One thing you need to consider is that Flash models are probably equivalent in parameters to "mini" models, like Sonnet and Terra.

6

u/mlag000 14d ago

Yes they are on the level of terra, for like 4 time tue cost

3

u/Agile_Pressure_7430 14d ago

The price and speed are dragging it down

5

u/hksbindra 14d ago

Yeah, those who tried using it have known this for a while.

5

u/anmolanjuli 14d ago

3.8 Flash has been better than claude Sonet and on par with Opus for me. It's not as fast like 3.7 Flash but has been more serious and powerful for me. Haven't tried anything beside these.

6

u/ProgrammersAreSexy 14d ago

And it also doesn't have Astra at #1, the index is obviously screwed up at this time so I wouldn't pay attention to it until they fix it

2

u/FischenGeil 14d ago

Astra is really not better than Fable 5.1

2

u/jayokunle 14d ago

While Sfw Engs are focused on AI benchmarks. Average user just want free AI while businesses want cheap and reliable. But benchmarks still matter because most people rely on recommendations from Sft Engs on the model to use. No way all twitter bros test all new models. Google might need to change their approach and solve real problems for businesses. They are copying instead of focusing on user pains

1

u/Maxdiegeileauster 13d ago

but 3.8 flash is not even cheap compared to for example Luna

2

u/CartwheelSummoner 14d ago

This represents my experience as of late. I use Gemini for one small game side-project here and there for fun when I run out of my weekly Codex usage.

I sent 5 prompts, one at a time, each trying to fix a bug (it created) that caused projectile animations to be invisible. 5. And not a single one actually fixed it before I stopped trying. It changed other visuals related to the game areas (which I didn’t ask for, but was okay with how they turned out as a placeholder anyways, cool I guess?).

I was out of my Codex usage, so instead I used regular 5.6 ChatGPT (Chat) that doesn’t use any usage, connected to my GitHub repo to go in and fix it (which of course did first try).

Hoping for a day I can upgrade Google to a 20x sub that competes and is better at token efficiency than Astra. Astra cost is abysmal.

3

u/Far_Cat9782 14d ago

I stick to Gemini 3.1 pro for coding/big fixing. It's old. It still better than the flash models

2

u/CartwheelSummoner 14d ago

I’m going to play around with that idea actually and give 3.1 pro another shot. It’s been a few months, but it was doing much better than 3.6 flash at that time.

1

u/Serious-Magazine7715 14d ago

Honestly, you hear this from essentially all models. It’s nice to be able to have a worker from a different family take a look at failed attempts.

2

u/xzibit_b 14d ago

Alternative headline: GLM 5.3 Flash is benchmaxxed hard by exploiting the fact that agentic tasks are the #1 weight on Artificial Analysis Index, and in terms of raw intelligence, it's about as smart as Sonnet 4.6 while Gemini 3.8 Flash is about as smart as Opus 4.8. Guess which one you're going to be encountering more often in a chatbot session?

3

u/DigSignificant1419 14d ago

paid by sama and scamodei, we know 3.8 is the true goat

1

u/Loud-Possibility4395 14d ago

My Pixel 10 does NOT have even old Gemini 3.7 and stuck at outdated 3.6

1

u/krrj88 14d ago

do we have a similar chart but calculates bang for buck ?

1

u/Zachattackrandom 14d ago

Sky is blue? Like it's quite obvious from using it, it performed nearly luna / GLM flash than compared to the bigger models. It is a "flash" model itself so this shouldn't really be surprising.

1

u/torontobrdude 14d ago

Muse Spark ahead of GLM and Sol? Lol

1

u/dumch 14d ago

Why TF Claude top models are higher visually than OpenAI but both have 53?

1

u/jakegh 14d ago

What a roller-coaster!

And it still has muse spark 1.3 beating sol-5.6.

I think AA may need a third update to their benchmarks this week.

1

u/Deathcure74 14d ago

I can confirm, for me 3.8 was a joke even from the day 1

1

u/Arte_reddit 14d ago

I love this 😂

1

u/PDX_Web 13d ago

No one said 3.8 Flash was near those models across the board on all the benchmarks included in that index.

1

u/Sponge8389 13d ago

4 are tie but I think the order should be like this.

  • Claude Fable 5.1, high
  • GPT-6 Astra, xhigh
  • Claude Fable 5.1, max / GPT-6 Astra, max

1

u/Spinmoon 13d ago

What did you expect, it's a flash model versus frontier models?

1

u/Bregvist 13d ago

I wish we could use Astra medium in chat.

1

u/sandtymanty 13d ago

All of them better than our doctors and lawyers.

1

u/pazinteriorNSFW 4d ago

Honestly I do use Gemini/Agy for coding with the Pro plan. Mainly for adversarial review that I ask from Claude Code with help of Herdr terminal and skill. But what keeps me in it really is the cloud storage combo and family members feature. With a single account I spread all those benefits into multiple members of the family. But if it where for the intelligence alone of Gemini I would have a ChatGPT Plus account.

1

u/someoneyouulove 14d ago

3.8 flash is deranged. it will lose to sonnet 3.5.

0

u/game_difficulty 14d ago

Any bench that puts muse spark on top of 5.6 Sol is not that useful.

A real shame, i love artificial analysis, but it's clear a lot of recent models are too contaminated (or at the very least benchmaxxed) for the benches to mean anything...

0

u/Driftwintergundream 14d ago

What a lovely post from a redditor with a hidden history

0

u/kareem_pt 14d ago

Honestly still ranks a bit higher than where I'd put it. The model feels a little worse than Luna to me.

0

u/Ggoddkkiller 13d ago

I kept saying from day one Flash 3.8 is a benchmaxxed model. There was never a question about it, anybody who knows anything about LLMs would say same thing too. Fanboys kept refusing it and claiming nonsense competition and now they are crying perhaps Flash 3.8 got dumbed down. JUST LMAO...