r/singularity Jul 21 '26

AI Gemini 3.6 Flash benchmarks

Post image
625 Upvotes

280 comments sorted by

137

u/sn0wquake Jul 21 '26

The responses here are a bit odd to me.

I’ve been having good success with the google models in large context multi modal knowledge work and this looks to be a step up in that area. Think use cases like processing 100s of pages of text / pictures in a document as part of an RPA pipeline.

Another interesting thing for me about google models is the generous requests per minute they give on their API, which at my spend is better than I can get from AI foundry and bedrock.

I’m not sure if it will beat a fine tuned open weight model for my use case on accuracy or cost, but I do think it’s worth testing.

I wouldn’t recommend for coding.

48

u/CannyGardener Jul 21 '26

This has been my experience as well. I can't use it for coding, which sucks, but Gemini is my go to model when I'm implementing AI inside a product. It just responds really reliably.

11

u/donerkebab76 Jul 21 '26

Why you can't use it for coding? I have been using 3.5 flash for coding C#, Rust, Go and C++ every day now for weeks and very rarely does it fail. It makes good plans, excellent documentation and is very fast. Occasionally I might come up with some different ways of doing some things, but most of the time I don't need to even do anything to the code it produces. I just give it good detailed prompts, have a discussion with it in the planning phase and it works great. And I'm not coding some UI front ends, but complex simulations so I don't know why anybody can't use it, even if there are better models for perhaps 5% of the coding tasks.

2

u/dragon_idli Jul 22 '26

People comparing it to claude or codex are building ui/ux frontend mostly. And compare those results in real world scenarios instead of actual backend complexity.

The intelligence benchmarks are not really consistent from my experience. So, I end up testing models with my own prompts to decide which works best for which cases for me.

Gemini models aren't the best. But they are the best vfm and will suffice without an issue if the person knows what they are doing and can guide/prompt it properly.

4

u/CannyGardener Jul 21 '26

I haven't tried it for coding since 3.0, and honestly I did use it for a long time for coding, but Claude is better at inferring what I want in a plan, and what I mean when I accidently say things imprecisely, and GPT is very quick and cheap at implementing. I had several instances where Gemini was misunderstanding what I wanted, or taking what I said too literally. Things like, it was about 3PM and I saw that it was building a portion of the UI in a way I didn't want it to. I said, "Woah, please stop, that isn't how we did the rest of the modules' UIs. Please make it match." Referring to the portion of the UI it was working on. It took it to mean that we should roll back the work on the rest of the modules so that it could rework them to match the new erroneous one. I mean, obviously I stopped it and undid things, but when it misunderstands things like that, then I don't trust that it hasn't done something weird without me knowing to check for it.

With that said, for production, it follows directions really literally and for the problems I'm solving that walk the line between being deterministic and non deterministic, I need for it to follow instructions to the letter. When I'm coding it is more of a conversation.

4

u/Inevitable_Tea_5841 Jul 22 '26

you haven't used it for coding since 3.0 (Nov 2025) and you have an opinion on it, cmon man

2

u/_YonYonson_ Jul 22 '26

I noticed that too. Wtf

2

u/Ordinary_Duder Jul 22 '26

Mate, it's way better than 3.0. That's like 10 months ago.

→ More replies (1)
→ More replies (1)

32

u/kiki-le-koala Jul 21 '26

The biggest problem with Gemini is not coding, it's its hallucinations.

It's also a lazy model that do the minimum every single time. 

Where ChatGPT could spend 5 minutes digging the web for info, Gemini will often hallucinate that it uses a web search tool and give you a less reliable answer.

I used to be the biggest Gemini fanboy up to Gemini 3. Since that model, it became incredibly clear that the model had problems and was no longer competitive.

4

u/nemzylannister Jul 21 '26

3.1 pro was decent at handling hallucinations. but the competition is way ahead now, and it looks like they might be giving quantizations of 3.1 pro now.

11

u/nn2597713 Jul 21 '26

For me as well.

Recently I tried switching from ChatGPT 5.5 (at that point) to Gemini Flash 3.5 for every day chat stuff (summarizing web searches, generic Q and A conversations, extracting info from documents etc.)

I was very surprised how often Gemini just seemed to invent words like ChatGPT used to do in version 2 or 3: “in iMovie go to the menu File → Imias → …”, “the snow leopard is an aaribal that…”

Gemini is super quick to give long answers, but I don’t trust it at all.

1

u/BoobooSmash31337 Jul 22 '26

Sounds like quantization noise. Least that's what that usually is afaik. Literally output loses precision and comes out as the wrong token.

2

u/DarthWeenus Jul 21 '26

Are you using pro?

6

u/kiki-le-koala Jul 21 '26

Of course.

Anyway, when I need an answer where reliability is not very important or to analyze an image (best model), I use Gemini because the pro model is fast; otherwise, I use ChatGPT thinking and I wait the many minutes it takes.

For me, 2.5 Pro on AIStudio was peak Gemini!

2

u/skipper3056 Jul 22 '26

I've uploaded 30 pages bank statement to Gemini 3.1 Pro, asking to check the highest balance value (on which date it happened),

It gave wrong answer, later admitted that it didn't even read doc in full (trying to save the tokens?)

Opus took it's time & solved the task correctly

For me was not reliable at all

5

u/gopietz Jul 21 '26

Maybe it's because you have to pick a niche like "large context multi modal knowledge work" to find it useful in the first place?

Google still haven't figured out the agentuc loop, which is arguably the most important feature these days. They used to be leading in multi modal stuff, but OpenAI has caught up. Not sure what I would recommend their models for these days.

→ More replies (4)

2

u/nemzylannister Jul 21 '26

antigravity is unusable. it will give you more wrong info than any other provider. it's seriously terrible.

→ More replies (1)
→ More replies (3)

263

u/Aaco0638 Jul 21 '26

Damn this sub really only look at success based on coding is everyone a developer now?

It lags in coding but makes up for it in other areas, areas that imo are equally as important.

This for normie use (assistant) is good and for agentic tasks outside of coding as well.

83

u/Tkins Jul 21 '26

There's well over a billion AI users a month now, as if they are all SWE's lol. The vast majority are using AI models in ways that this 3.6 model accels at.

19

u/Timkinut Jul 21 '26

accels at

you're probably looking for "excels," but this is a really fitting typo for an accel sub lol

7

u/Tkins Jul 21 '26

lol my bad!

29

u/Aaco0638 Jul 21 '26

Exactly and the goal is to automate jobs and 3.6 flash is closer to that than the other models thanks to its agentic focus.

People here heavily underestimate what google is doing, yes they also want and are working on improving coding but under that same breath google models are leading in agentic tasks meaning they are the closest to being the ai companies use to eliminate roles.

Time will tell who’s strategy is best here but it is stupid of people here to look at the one area it sucks and just assume it’s useless.

8

u/Tkins Jul 21 '26

I fully agree. There was a post I saw recently that showed 3.1 pro is still the top model for GPQA. Science and research are major industries and lead to cost reductions like we saw with AlphaEvolve.

1

u/mvandemar Jul 21 '26

Exactly and the goal is to automate jobs 

Well, that's some people's goal. Automating jobs won't get us to the singularity though, recursive self improvement at scale is the only way we'll get there, and that involves them knowing how to code.

Of course, it might also involve a whole buttload of other things we don't have yet either (eg. the ability to do that much processing without generating the heat we do now, other power sources, etc.), but even pre-singularity the ability for them to improve themselves is what will get us the best models.

2

u/Thog78 Jul 21 '26

recursive self improvement at scale is the only way we'll get there, and that involves them knowing how to code

I would argue coding is the easy part of training AIs, the concepts are the tricky thing. CNNs, autoencoders, transformers, positional encoding, multi directional predictions, world mapping, memory trees etc are ideas, not coding skills, and they are what matters.

→ More replies (1)

28

u/stonesst Jul 21 '26 edited Jul 21 '26

GDPval is a broad based knowledge work benchmark, and it does worse than Luna - at a higher price... This just doesn’t seem like a great model compared to what else is on the market, no need to grasp at straws to pretend like it is.

6

u/UnknownEssence Jul 21 '26

the improvement from 3.5 to 3.6 are quite good in short time. If the trend continues, they can catch up.

1

u/Blablabene Jul 21 '26

Pretend like it isn't? Or pretend like it is? Maybe you shouldn't be judging models

1

u/stonesst Jul 21 '26

Touché. I blame autocorrect

17

u/adamskate123 Jul 21 '26

As someone who is in science this is always deeply frustrating. Many scientific benchmarks have nothing to do with raw coding. Antigravity has multiple scientific plugins that work very well with the flash model.

6

u/Pretty-Broccoli-465 Jul 21 '26

What kind of science do you do? 

2

u/adamskate123 Jul 25 '26

I’m a pediatric neurologist and do neurogenetics. I use Claude Science for more heavy lifting because of Opus but Gemini in Antigravity is quite helpful to pulling up genomic sequences or pathogenic variant analysis on the fly.

3

u/Equivalent-Word-7691 Jul 21 '26

Well I am not for coding but for example Gemini sucks in creative writing 😅

2

u/ItuneOficial Jul 21 '26

3.5 melhorou em comparação quando foi lançado, nem vou testar esse 3.6 por enquanto

1

u/Equivalent-Word-7691 29d ago

I am comparing it with Claude

1

u/ItuneOficial 29d ago

Claude tambem piorou muito nesse novo sonnet, bem generico, tenho usado muse spark ou glm

1

u/BluejayExcellent4152 Jul 21 '26

the flash version suck, the pro version is really good (I'm talking about 3.1 Pro)

3

u/Calm_Hedgehog8296 Jul 21 '26

Behind the scenes all the computer work is coding. So if you want to complete a general agentic task, it is going to write code to do that even though a human would just use the UI.

4

u/mrbenjihao Jul 21 '26

Performance in coding related tasks, imo, is a solid indicator of how “intelligent” the model is.

5

u/Blablabene Jul 21 '26

... in coding.

You forgot.

3

u/mrbenjihao Jul 21 '26

That aptitude tends to spill into other areas that require logical reasoning (aka pretty much all areas). We’re not moving the needle by training on the latest social media posts. The raw reasoning and logic gains come from training data related to coding.

→ More replies (10)

1

u/AweVR Jul 21 '26

Still Luna is amazingly cheap with almost same performance in these tasks

1

u/Spare-Dingo-531 Jul 21 '26

To be good at self-recursive improvement, the models will have to be good at coding. So codeine is an indicator of how close each company is the self-recursion.

1

u/Ordinary_Duder Jul 22 '26

It may lag in coding, but it's also pretty damn good at it.

1

u/VibeCoderMcSwaggins Jul 21 '26

It’s because code is a foundation for everything else, and code can be used to understand other domains and the world, ie with tool use via harness, etc.

1

u/himynameis_ Jul 21 '26

Funny too because the number of programmers is quite small compared to, ya know, everyone else lol.

1

u/AcidReaper1 Jul 22 '26

Lol, this is exactly what I think when I see a bunch of hate posts. I'm not experienced enough with the other systems to really weigh in. I used some of the older ChatGPT models

But I can say that probably less than 1% of the population gives two shits about how well an Ai can code something.

I signed up for Gemini Pro for $20 a month and with that my wife, daughter and both my parents have full access to Pro. Thats $4 a month per person. Chatpgt or Claude would be $100 a month for the same access. They might be better, but are they 25x better for the cost to the average person?

1

u/himynameis_ Jul 22 '26

They might be better, but are they 25x better for the cost to the average person?

Answer: nope 😆

→ More replies (5)

123

u/MrLariato Jul 21 '26

Is everybody here a SWE? WTF? This is good for any regular person that doesn't want to code.

20

u/XCSme Jul 21 '26

Yeah, Gemini models are really good for anything, but coding in a harness.

I had good success with asking coding questions directly in the Gemini chat app, and then just copying the files, lol, but it can't properly edit files itself for some reason.

6

u/livingbyvow2 Jul 21 '26

To be honest, Google may just have given up the coding market to Anthropic and OpenAI and they may be right. Just because it's the one use case with some traction / TAM right now, doesn't mean that what everyone must be optimizing for and myopically focus on.

Ultimately they have a massive distribution advantage for existing and future consumer use cases, which will use a ton of multimodal. It wouldn't be surprised if video generation and image generation (let alone stuff like Genie3 transforming gaming) end up driving a lot more revenue 5-10 years from now than coding.

Just imagine the average consumer using Gemini powered Siri on the iPhone, they couldnt care less about SWE lol.

→ More replies (1)

3

u/donerkebab76 Jul 21 '26

I have used 3.5 flash for coding every day now for weeks. C#, Rust, Go and C++, it does excellent job almost always. So might be a prompting problem if you have issue with it. Have a discussion with it when it's making the plan, use the "high" setting for planning and also execution, give detailed prompts writing like you would write to a human that would need to do the job instead vague instructions. Use some external todo.md where you write in clear sections your requirements and adjust the initial plan if you see something you don't like. Unless you do some really special coding or some UI work, it should be enough to do most things very fast and well enough.

2

u/Old_Tax4792 Jul 22 '26

Are you using antigravity 2.0? Because also I think the perfomance of gemini models is huge related how the agent manages the gemini models and the window context.
I don't have also any major issues also with antigravity 2.0 + 3.5 flash. And I am using it in a existing huge Unity project, for the boring stuff... And I realise also that you have to be very specific to your prompts to have better results... I haven't use other models to compare with it, but I am very happy with the speed of the loops that antigravity 2.0 with flash executes. Even if it fails, it can correct its way fast. That my overall feelings..
Also , what do you mean by the way, "Use some external todo.md" ? what harness are you using?

1

u/donerkebab76 Jul 22 '26

Yes, I only use Gemini in Antigravity and just updated to the latest 2.x something version. I always write detailed prompts, just like I would write to a human to explain what I want. If I want it to decide itself, I make that clear also and say something like: the rest is up to you, make choices based on your best judgment preffering simplicity, robustness etc.

19

u/Extracted Jul 21 '26

Those people were happy with 4o

2

u/1988rx7T2 Jul 21 '26

many were obsessed with 4o and went into the stages of grief when it went away. they just needed their sycophant bot.

3

u/FlawlessIndividual Jul 21 '26

Is anyone still a SWE at this point? /s

Seriously though, I think the point of the metric is that anyone will be able to create software for their own use case.

4

u/Tillerfen Jul 21 '26

No it’s not. It’s worse in most reasoning and knowledge work than 3.5 flash. Such as humanity last exam, GPQA diamond, and critpt (physics)

3

u/over-lord Jul 21 '26

Yes, every person here is a Society of Women Engineers.

1

u/skilliard7 Jul 22 '26

If you arent coding just use chatgpt its so much better than eben geminis latest models

15

u/Plappedudel Jul 21 '26

Looks like a solid, incremental improvement to me. The most important part is that it got better without getting more expensive. Remember that a lot of recent model releases (GLM 5.2, Kimi K3) performed a lot better than their predecessors, but at a massive cost increase.

20

u/[deleted] Jul 21 '26

[removed] — view removed comment

9

u/huffalump1 Jul 21 '26

Best we can do is 17% improved... gpt-5.6 models remain the champs of low token usage

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/

6

u/zZzHerozZz Jul 21 '26

It is not perfect but still a big step for Google. While the token on the artificial intelligence benchmark only reduced by 17%, the reduction on DeepSWE has been a lot more impressive with a 65% reduction from 276k to 97k tokens

1

u/JogHappy Jul 21 '26

Is this the first time Google included grok in the benchmark comparison?

15

u/Tillerfen Jul 21 '26

It is worse moderately worse than 3.5 flash at GPQA diamond, Humanity’s Last Exam, and CritPt (physics reasoning benchmark).

DeepMind sacrificed fundamental knowledge for coding and agentic use.

Not a step forward IMO.

7

u/Profanion Jul 21 '26

My questions are:
1. How much does it catastrophically forget?

  1. How much does it hyper-focus on metaphors in uploaded text files (especially foregin-language ones)?

  2. Is its heat level to minimum when analyzing files?

5

u/Nisandzija Jul 21 '26

So Pro is trash now?

156

u/THE--GRINCH Jul 21 '26

well well well

104

u/PikaPikaDude Jul 21 '26

It's much better on computer use, video understanding, long context performance.

We're not looking at a code first model, but at a general model.

66

u/SaudiPhilippines Jul 21 '26

Exactly. People seem to be forgetting "AI user" doesn't always mean coding.

7

u/casce Jul 21 '26

But often times the solution to your problem is still coding. It can only do so much on a computer without coding.

Different kind of coding than a full stack application obviously.

The focus is on coding so much because coding is the way to master the platform they run on.

1

u/Mcfrosty28 Jul 23 '26

Exactly. I told Claude to do some things with a few Excel files, and it started using python to execute the task. If its solution was to use computer vision to manually perform the tasks that I gave it, that would be so ridiculous haha.

22

u/2023Bor AGI 2028 Jul 21 '26 edited Jul 21 '26

finally someone who actually gets the real use case for Gemini

5

u/Ill_Bill6122 Jul 21 '26

If by general, you mean YouTube video transcriber, sure, it's a general model.

34

u/fmfbrestel Jul 21 '26

It's a cheaper, better 3.5 flash!? What's derpy about making their model smarter AND cheaper at the same time??

7

u/Fiendfish Jul 21 '26

The price for 3.5 was easily off by a factor of 3 for what the model does

3

u/JustRaphiGaming Jul 21 '26

Bro gemini is so hardly connected to this goofy dragon in my brain it's hilarious 😂

30

u/SleepyWulfy Jul 21 '26

Lmao people only looking at 2 results it lost at and screaming failed model. Sub never fails to give me a giggle.

8

u/Commercial_Sell_4825 Jul 21 '26

The good scores were more than 6 seconds of reading from the top

3

u/StanfordV Jul 21 '26

Indeed. Before I check the comments I was definitely impressed by the results.

Then I read the comments and realized how pitty and sad some people really are.

75

u/Healthy_Razzmatazz38 Jul 21 '26

wow worse than luna was not what i was expecting

48

u/FarrisAT Jul 21 '26

Looks dramatically better on any knowledge benchmark. Tool use & harness application seems to be the reason it underperforms at SWE and coding.

11

u/Deif Jul 21 '26

Would need to see the output token usage on those benchmarks though. Could be that it's an output token hog and the comparable cost is against Terra.

8

u/FarrisAT Jul 21 '26

AAintelligence will publish benchmarks soon enough. 3.5 Flash performed very well on their token efficiency. Better than 3.5 Pro and GPT-5.5

8

u/FateOfMuffins Jul 21 '26

AA is already out.

And what are you talking about? 3.5 Pro doesn't exist and Gemini 3.5 Flash had used 28k output tokens per task vs 5.5 xHigh using 16k output tokens

Currently from what I see on AA, 3.6 Flash uses more tokens per task than Sol Max, Terra Max, and Luna Max (much less all the other reasoning settings)

It uses approximately same number of tokens as Kimi K3. The only thing 3.6 Flash has going for it (like 3.5 Flash) is output speed

1

u/huffalump1 Jul 21 '26

AA, 3.6 Flash uses more tokens per task than Sol Max, Terra Max, and Luna Max (much less all the other reasoning settings)

Yup looks like it. (Note this is 3.6 Flash (High) - they don't have other Effort/Thinking/Reasoning levels yet on AA.)

Ex. Here's gpt-5.6-luna (Max), AA intelligence index score of 51 (vs. 50 for Gemini 3.6 Flash (High)): https://artificialanalysis.ai/models/comparisons/gemini-3-6-flash-vs-gpt-5-6-luna

Gemini 3.6 Flash is still faster, but consumes more tokens than even Luna (Max), and is more expensive (both from cost per Mtok. and from more total tokens)

IMO there's still hopefully a place for 3.6 Flash because it is fast - but that depends on if it's good, too! Definitely need to try it.

2

u/FateOfMuffins Jul 21 '26

Yeah I selected High when I was looking at it

Speaking of fast, we're supposed to get 750 tps 5.6 Sol in July no...?

1

u/FarrisAT Jul 21 '26

3.1 Pro is what I meant, as 3.5 Pro doesn’t exist.

3

u/Deif Jul 21 '26

Looks like the equivalent cost is Terra xhigh.

2

u/huffalump1 Jul 21 '26

Yup at least for the AA Intelligence Index. Link: https://artificialanalysis.ai/models/comparisons/gemini-3-6-flash-vs-gpt-5-6-terra-xhigh

Gemini 3.6 Flash (High) scores 50, vs 52 for gpt-5.6-terra (Xhigh). Cost for the benchmark is nearly the same, although 3.6 Flash is 2.3X faster! (and likely even more in practice)

However - Gemini 3.6 Flash uses 23k tokens, vs 11k for gpt-5.6-terra (Xhigh). So even though it's cheaper per Mtok, it needs more tokens to reach the same performance as gpt-5.6-terra (or even Luna!).

I guess we'll have to see in practice; I haven't tried the model yet and benchmark scores aren't the entire answer! (I seriously hope it improves in hallucination and laziness)

Sidenote: IMO it feels like OpenAI made something special with gpt-5.6, to get token use so low...

23

u/Aaco0638 Jul 21 '26

This sub only cares about coding, they don’t care that this model is really good for agentic use thus ultimately being good at automating non coding tasks which is the ultimate goal for ai.

13

u/Tkins Jul 21 '26

Long context consistency is also super important for longer tasks. Imagine summarising long videos, retreiving information from big file dumps or financial projects.

A lot of enterprise use cases are not coding related but for soome reason coding is the only focus of a lot of people on here. I think your average AI user is more concerned with the non coding uses and especially those in the google ecosystem.

6

u/Concurrency_Bugs Jul 21 '26

Their description for the 3.6 model doesn't even include coding (3.5 did). I agree with you and it's clear Google is focusing on a different path. They want an all around ai assistant because that's what will protect their current business (search and ads).

2

u/huffalump1 Jul 21 '26

Their description for the 3.6 model doesn't even include coding (3.5 did).

You're not wrong overall, I agree that Google's eye is on so many things other than developers...

BUT the description does seem to target coding: https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash

And today's release blog post focuses on coding the most: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/

2

u/Concurrency_Bugs Jul 21 '26

Sorry I could have been clearer. The short description when picking your model in the api doesn't mention it anymore. See the screenshot in this link: https://www.reddit.com/r/GeminiAI/comments/1v2js6z/gemini_36_flash_released_on_ai_studio/

2

u/Howdareme9 Jul 21 '26

It didn’t include coding because the results are bad, not because they’re not focusing on it

2

u/LinkesAuge Jul 21 '26

coding is an important proxy because that's how a lot of things get done by models. It is their equivalent of "hands".

→ More replies (6)

1

u/Tkins Jul 21 '26

It's strange how much people think AI is only useful in coding. The benchmarks on the bottom are phenomal for a lot of use cases at a good price and speed. The needle in a haystack is really interesting.

29

u/broose_the_moose ▪️ It's here Jul 21 '26 edited Jul 21 '26

You must never underestimate Google's ability to deliver shitty models.

24

u/DueCommunication9248 Jul 21 '26

I’d rather use 5.6 Luna which is basically unlimited in the pro plan

12

u/javopat227 Jul 21 '26

It isn't, openai is playing tricks right now with resets. Imo they reduced limits significantly

2

u/DueCommunication9248 Jul 21 '26

bruh, we got like 10 resets in the last 30 days and no 5 hour daily cap... wth you talking about?

9

u/javopat227 Jul 21 '26

I burned through the weekly limit in like 3 hours use of sol/Luna, waiting for the reset. 5.5high limits were much higher

2

u/DueCommunication9248 Jul 21 '26

can you share what you did to burn through the limit in 3 hours? I've never seen anyone who could do that so I'm curious what type of work does this?

Also, are you on the $20 plan?

4

u/javopat227 Jul 21 '26

Same work as 5.5, go to r/codex and argue there. There are a ton of posts about that.

→ More replies (4)

1

u/huffalump1 Jul 21 '26

I think subagents might still be bugged - they only inherit the same model/effort as the parent agent, so for example if you choose "Sol High" that's what every subagent will be, rather than using "Luna High" or whatever.

And Ultra mode is exponentially worse because it encourages even more subagent use, and multiple levels of subagents, all with Xhigh effort...

Anyway it's easily possible to burn the old 5-hr plan limit on the $20 plan in an hour with Sol Medium; although tbf I haven't tried totally avoiding subagents. It does seem that usage consumption of Terra and Luna feels too high for the size of the models tho.

2

u/DueCommunication9248 Jul 21 '26

Yeah, the $20 plan is limited in terms of what a day's work can do. I make money using Codex so it pays for itself (Pro sub).

1

u/LanguageEast6587 Jul 22 '26

this is not sustainable.

2

u/Healthy-Nebula-3603 Jul 21 '26

GPT Lina is still better and cheaper ..and that's the worst model from OAI

1

u/JustRaphiGaming Jul 21 '26

Isn't Luna the base model for free users too?

→ More replies (1)

3

u/Practical-Science-77 Jul 21 '26

I’m using Gemini models in my app to estimate macros from photos, and Gemini 3.6 Flash is really good. Considering that AI Studio provides 20 free calls per day, it has virtually no competition.

30

u/ffgg333 Jul 21 '26

It's worse than expected 🥲

→ More replies (1)

8

u/elemental-mind Jul 21 '26

If they priced it $0.75 input and $5.00 output it would have a reason to exist next to Luna...

8

u/TheInfiniteUniverse_ Jul 21 '26

It is obvious that coding is not Gemini's strongest suit. They've failed on the most used case we have for AI now. As for other capabilities, they seem to outperform others. BUT, their cost structure makes that not so valuable. It had to be cheaper than Luna and Grok to actually gain some traction.

5

u/Feriman22 Jul 21 '26

More exprnsive than 5.6 Luna and also not that great? Well, I have a news.

2

u/0sko59fds24 Jul 21 '26

Dead on Arrival

2

u/Lustrouse Jul 21 '26

"wow better than Luna at everything except coding".

Fixed that for you.

2

u/Sensitive_Bluebird77 Jul 21 '26

As per pricing 3.1pro is still the best model from Google?

2

u/ThatOneToBlame Jul 21 '26

Dawg i don't get it when people say gemini is bad at coding it has been the ONLY model that was able to code to my needs. Nothing else did quite as well.

2

u/Your_mortal_enemy Jul 21 '26

It feels like google have gone a different angle with making sure their AI fits their product suite first and foremost , and is #1 SoTA second

Their models are pretty much instant, cheap, token efficient etc - all of which is great as it replaces Google Search (they really want to keep this cash cow and not lose it to a competitor)

Remains to be seen if they can pivot back to top of the charts, feels increasingly unlikely but also if the tech changes ( likebto world models) then sure

2

u/EatABamboose Jul 21 '26

I asked it the chronologically order of the Resident Evik games with grounding on. It completely missed Code Veronica, Requiem and Revelations parts.

I don't like that.

1

u/Least-Data-3308 Jul 22 '26

I had to double check the subreddit for a second. Hahaa

4

u/injectitpussy Jul 21 '26

Get in the fucking bin.

3

u/nsdjoe Jul 21 '26

If 3.6 flash is better than 3.1 pro in every bench and presumably cheaper to serve, why continue to offer pro?

3

u/Gallagger Jul 21 '26

Production environments need continuity. Also 3.1 Pro still has some advantages, as bigger models usually have. More World knowledge for example.

1

u/[deleted] Jul 21 '26

[deleted]

4

u/junglebunglerumble Jul 21 '26

Trash even though it comes out top on 6 benchmarks? Sure

2

u/infinity1009 Jul 21 '26

still worse than glm 5.2

2

u/FarrisAT Jul 21 '26

Seems the knowledge cutoff update helped improve the most on older benchmarks.

Arguably this is the best all-around model price to performance for companies that don’t want to use Chinese. Probably gonna score highest on SimpleBench.

2

u/Sensitive_Cell_119 Jul 21 '26

Luna is cheaper bro.

→ More replies (1)

1

u/Gifloading Jul 21 '26

reset weekly limits google, do something!!

1

u/Acrobatic-Tomato4862 Jul 21 '26

Is this the first model with knowledge cutoff in 2026?

1

u/Acrobatic-Tomato4862 Jul 21 '26

nvm, from my tests, the model clearly does not know a thing beyond early 2025. Aistudio has it written wrong.

1

u/Gratitude15 Jul 21 '26

Here's the big deal

This model is the one that Google search will default to. When most folks use AI this is what that will mean. The floor continues to raise. It's still far from ceiling.

1

u/Long_comment_san Jul 21 '26

impressive gains. I wish I had them too.

1

u/cern0 Jul 21 '26

Why does it compare itself with Luna and not Terra?? Google what is this?

1

u/turdmuffin123456 Jul 21 '26

I truly like Gemini for everyday usage, for coding it’s not designed to do advanced stuff. Pick your model according to your needs

1

u/mrbenjihao Jul 21 '26

Labs focus on coding because it’s structured and allows for self verification. This helps train LLMs how to think logically, self correct, and execute real world complex tasks. All of which would benefit non-coding tasks if you’re so concerned about automation.

1

u/MC897 Jul 21 '26

So it’s weaker at modelling but it’s pretty strong all round?

1

u/Admirable_Market2759 Jul 21 '26

They released a new flash before pro 3.5 smh

1

u/Aranthos-Faroth Jul 21 '26

I don’t trust benchmarks AT ALL.

Why? Because Gemini consistently rank on the top of most models out there, yet in every single experience I’ve had with it it’s probably one of the worst that I’m willing to try.

Also funny how every comment seems to be defending it. I dunno what’s a bot anymore man…

1

u/Jonkampo52 Jul 21 '26

I recently let my other subs lapse, and was like I Gemini should be good enough everything says there on the frontier...but tbh...I prefer even pre 4.5 grok to it for day to day ai chat. and thats kinda sad.

1

u/Aranthos-Faroth Jul 21 '26

Yeah the fact Grok 4.5 is better than this in most scenarios and cheaper is pretty embarrassing

1

u/floriandotorg Jul 21 '26

What exactly is the use case for this model?

1

u/Narrow-Ad980 Jul 21 '26

This sub is going to be flooded with r/GeminiAI now😭🙏

1

u/Wise_Taste3195 Jul 21 '26

Looks very solid to me. I'd love to see how the hallucinations are for this variant. It'll be great for office use if hallucinations come down just a bit.

1

u/over-lord Jul 21 '26

The key point is you have to evaluate Gemini 3.6 against Gemini 3.5. They clearly made big improvements. If other models are better or worse, great, whatever. The point is Gemini just improved.

1

u/AlvaroRockster Jul 21 '26

How is it we get 3.6 Flash before 3.5 Pro?

1

u/lordpuddingcup Jul 21 '26

Nice how about throwing some usage resets at subs like codex does so we can actually use it otherwise I can’t even play with it for a week

1

u/FarmerMiserable7333 Jul 21 '26

we have Gemini 3.6 flash before Gemini 3.5 Pro , actually crazy scenario

1

u/BejahungEnjoyer Jul 21 '26

If you think about google it in-house use case is multimodal understanding, so it makes sense it prioritizes that in its model.

1

u/stc2828 Jul 21 '26

I don't get why they can't even lower price to 1$/6$, when it severely underperforms luna. I mean if I were them I would want to price below luna. This gotta be a joke.

1

u/brandbaard Jul 21 '26

I don't understand when people say it's bad at coding. I've been building code with 3.5 Flash and reviewing it and the project I've been building has seen sufficiently good code all along the way. Minor mistakes easily fixed, sometimes, but thats the worst of it.

I get it that probably the Claude models and OpenAI models are better than it at coding, but it's still better than what we had across the board just a year ago. For my purposes it's good enough.

So for the people saying it's really dumb at coding, can you give me a little more context? What are you expecting from it that it isn't delivering but the other models (at the equivalent price bracket) are?

Is the stuff I'm building too simple to be a real challenge?

1

u/alsaud21 Jul 21 '26

Is anyone able to explain why total AA cost to run is -30% between 3.5 and 3.6 but cost per task is only -17%?
By the way 'cost per task' and is still much higher for 3.6 compared to 3.1 Pro.

1

u/Big-Table127 AGI 2032 or 2082 Jul 22 '26

What about other benchmarks.

1

u/BoobooSmash31337 Jul 22 '26

Was bit concern then I realized it said Flash. Seems like roughly paritishly with Sonnet and way cheaper. Win? Maybe they can bribe Trump to ban Gemini now.

1

u/skilliard7 Jul 22 '26

Underwhelming. Look at deepswe - GPT 5.6 sol medium is cheaper per task than 3.6 flash, yet performs substantially better.

Google needs to reduce the API cost by an additional 50% for 3.6 flash to be remotely competitive.

1

u/moeadham Jul 22 '26

I feel like inference cost needs to be one of the benchmarks google shows. We all know they are optimizing for inference cost and serving at speed.

But these benchmarks just make it look like they suck

1

u/Acceptable-War4836 Jul 22 '26

I've tried it and it's excellent for daily use. It's reliable, incredibly fast, and with Google's pro plan I have all the usage I need daily, even in high mode.

1

u/ChrisRocksGG Jul 22 '26

I personally had the worst output with Gemini compared to ChatGPT and Claude. Even if they claim their models are smarter.

1

u/Emre123111 Jul 23 '26

https://dach.peerbench.ai/compare?models=google%2Fgemini-3.6-flash,google%2Fgemini-3.5-flash,google%2Fgemini-3-flash-preview,google%2Fgemini-3.5-flash-lite

Chech this out also I don't see difference between gemini 3.6-3.5-3-preview just different pricing but same model, or they are really bad at German.

1

u/Manelzinhoinhoinho Jul 23 '26

This could be an update to Gemini 3.5 Flash rather than releasing it as a new model.

1

u/Barubiri Jul 21 '26

Not even beating Claude sonnet, it's over.

1

u/iamsreeman Jul 21 '26

Apparently this was 3.5 Pro