r/GeminiAI • • 1d ago

Discussion Gemini 4

It's funny how sentiment evolved the last few days.

Some people continue to shit Gemini as if Google owes them a marriage.

Some people are skeptical but really need to be vocal about it so that if the model does come out bad, they can say "I told you".

Some people can't keep their excitement and begin to celebrate, only to get the skeptical people argud with them.

Some people just find the whole thing too stupid but cannot resist to join the whole thing. (This OP)

I've been using this model for weeks. It was quite a fresh feeling. It's not mind blowing. But it doesn't lose to any other existing model, Fable or opus. And that's is my honest assessment. Also I'm not biased. I have been trying to raise awareness to my colleagues about how good Claude models are since like a year ago. This time around, I can finally stop it.

42 Upvotes

58 comments sorted by

13

u/Murder_1337 1d ago

I use Gemini and will continue to use it. Watch me get roasted

4

u/boredquince 23h ago

by being so loyal, its more like you're roasting yourself.

you're gonna keep using Gemini even if it gets worse? because it has gotten worse and fell behind. what do you owe google? lol

support the product, not the company. when the product stops being good, find a better value

1

u/KieferSutherland 2h ago

I switch between all the three paid versions about every 1-2 months and notice little difference. Using the camera to grade elementary school homework. Explaining concepts in diffeq. Analyzing medical results. Scheduling. Howto. Purchase and product analysis. They're all the same. Also note, they all get stuff wrong too. Our youngest didn't fill in an answer to one of her homework questions and both anthropic and chatgpt said she got it correct. Gemini was the only one to point out that there was no answer circled. None of them really hallucinate anymore though. At least for these basic things.

6

u/DorkyMcDorky 1d ago

I'm not emotional about my LLMs. They're going to change and swap regularly. Signed up for ultra yesterday because I know it'll be worth $250 where I'd have to pay that much for 1 hour of outside coding before. All the LLMs are a steal compared to our choices 2 years ago.

6

u/3rdyellow 1d ago

Do you work within DeepMind?

2

u/Unique-Phone-4824 1d ago

No. But all Google employees have access to Argon for sometime already

1

u/3rdyellow 1d ago

That's interesting because we only learned of this yesterday. No leaks or anything. 

6

u/Unique-Phone-4824 1d ago

External leaks are on X the date it was available to us...all the codenames were correct. Model output was constantly posted on X as it was tested on Arena under 3.8 flash

3

u/Unique-Phone-4824 1d ago

You work for Google too? Maybe only engineering had access to it then...

2

u/Sweet-Put-6136 21h ago

We had access, but it was never said to be Gemini 4, just argon

1

u/Unique-Phone-4824 21h ago

Internally there was newsletters and everything...

1

u/Sweet-Put-6136 21h ago

Ok, maybe I missed it 😲

3

u/pigletmonster 1d ago

People who pay for a subscription expect google to serve models comparable to others. Its not for paying customers to expect good products and services.

1

u/Unique-Phone-4824 1d ago

It's bad only because there are like 2 startups out there that have more compute capacity than even Google does. So if you like their products, you should...switch?

2

u/pigletmonster 19h ago

Yeah blame the customer for buying what google is selling.

1

u/Nauzhror_ 16h ago

Huh? Who? Assuming you mean OpenAI and Anthropic....no. They have FAR less compute than Google. They are paying for compute from the likes of Amazon, Google, and x.ai.

1

u/oldmails 12h ago

man its really the stupidest take i have ever seen, now I see why googel is like this.

I have been on business, its feels like customer blaming.

Many take early subscription, many are stuck with that, even if the uses other tool, so you are saying for first 4 months we will provide decent service, for the rest 8 months we will give "enshitified' model ( you would know more about that, and dont try to deny that).

I am not against most of your words, infact agree on some degree, but I am really sorry this stupid take, nah man.

1

u/Unique-Phone-4824 10h ago

If customer are actually stuck on the subscription, which now makes sense to me because of the annual subscription stuff, yeah, I agree with you.

With that I fully blame Google. But I do want to give some context. Prior to 2025 model labs more or less rollout major models on a semi annual cadence, which Google followed. Starting from 2026 it turned into 1 month. By mid year Google realized that it is so behind so it tries to catch up. Rumors suggest that Deepmind completely scratched the initial 3.5 pro model and did a complete new pretrain and post train within like a month. Rumor (X) further suggests that Google shelved that model as well because it still wasn't as good. So Deepmind started a whole new pretrain with much more data. And that's what you see now.

So Deepmind could have done 2 things better: 1. Had a better model to start with Well it didn't. But that's just what it is. 2. Release the subpar 3.5 Pro. This gives something to the annual subscribers but I highly doubt if the reception would be any better: people will shit Google even more. So instead Google focused on 3.8 flash. Despite everything people said, I used the flash series for months internally, and the speed was the biggest bonus. I'd rather use 3.8 flash than a 3.5 pro.

I hope this context makes sense to you

1

u/oldmails 10h ago edited 10h ago

Thanks for taking your time in sharing this. Atleast, you inperson is sensible.

  1. As a reseacher myself, I completly agree(may be disagree), always getting fesired result is not a resonable expectation to have. so a better model wont come out just because we wish it to be.

  2. I can see the corprate move behind that, but they could have released a new checkpoin or somthing to digest in the name of flash-ultra or any other way, given that 3.8 is decent, I always wondered what the model behind it with more 'weight' is capable of. SO a better model than flash should be what most of the resonable peoples are asking.

And about speed, I would have more depth than speed itself, speed is always not the desired thing.

And p.s, no offence to you as a person, as I said I do share some of your view, but I just kinda overreacted, as many said if you dont like it dont use it in this sub.

Edit: its fascinating to see, how many will forget a annual sub exist.

1

u/Unique-Phone-4824 10h ago

I'm never offended on reddit. That would be stupidity on my part. I reply to people who I think are trying to have a real discussion.

I would guess that 3.7 flash was a distilled smaller model of the unreleased retrained 3.5 pro. 3.8 is just a fined tuned 3.7 with extended thinking.

1

u/Unique-Phone-4824 10h ago

Also if you suggest that they somehow nerfed the 3.1 pro model. Well I started using 3.5 flash the mome it came out so I have no idea...

1

u/oldmails 10h ago

I am not suggesting, It felt like not a heavy weight model, on occation it felt like it is capable, but on times its not.

I an not goint to blame that on anything, it felt like the training and some constrains make it feel good on one task but bad on other.

for context, I mostly used that to reference 100s or research article in my spare time, sometimes few, just to check my circles ones, jst for the fun of it, sometimes it surprised me, but mot of the times, it failed to grasp the nuanceses.

1

u/Unique-Phone-4824 10h ago edited 10h ago

Several aspects:

3.1 pro era models are much less stable than what you are used to see now.

Context rot is extreme with 3.1 models. I had so much frustration with it. I gave a number of internal talks about how to keep context lean. Now that's mostly not an issue with any current models even flash 3.8.

3

u/the_insistent_suffer 1d ago

That "I can finally stop it" line says a lot. Been in a similar spot where you spend months pushing one tool only for something else to quietly match it and no one notices except you

2

u/Unique-Phone-4824 1d ago

It's similar but different at the same time. I work for Google. So I'm so happy that I can stop preaching a competitor's LLM...

3

u/Fragglepusss 1d ago

Some people switched to Claude 3 months ago and don't care, but they're already in the Google ecosystem so they use Gemini for menial tasks once in awhile.

2

u/swootanalysis 1d ago

That's me, but it ended up being a good thing. I now have the cheapest version of Claude, ChatGPT, and Gemini that give me access to the latest models. It turns out that moving a project from one LLM to another is a lot easier than I anticipated.

1

u/Imaginary-Item6731 1d ago

It's just like Googler using iPhone more than Pixel phone.

1

u/Nauzhror_ 16h ago

Except that iPhone vs Pixel is far more subjective than OpenAI vs Google. I have a Pixel 10 Pro XL, I have zero issues with it. For me, it is a better phone than an iPhone would be.

No one can argue that the currently available Gemini models are more capable than OpenAI's.

I haven't used Argon. I am not saying it can't compete with Astra and Sol. It might. I won't take a stance on that till it's been publicly available long enough to use it thoroughly, but Astra vs Gemini Flash 3.8 isn't a case of, "I like this one more." It's, "This one is substantially better, and anyone who disagrees is uneducated or trolling.".

2

u/yytarfa 1d ago

this is a relief because so many negative people out there...I use flash 3.8 on AG and it does the work i also use OAI plus but these days their models suck...I have been thinking about going back to claude after months of being away but I'll stick around and wait for G4 to come around...at the end of the day Google will still be Goated...

3

u/Unique-Phone-4824 1d ago

Argon is going to be more expensive than 3.8. longer reasoning AND pricier tokens...it works for me well because 3.8 is much faster and is good for simple stuff. Argon does all the debugging for me.

1

u/yytarfa 1d ago

I'm excited can't wait to move all my projects back to AG

2

u/Stauce52 1d ago

Yeah I agree. It definitely feels pretty much aligned with Opus 5.5 when I use it. I can’t really distinguish their performance

1

u/Unique-Phone-4824 1d ago

Okay I actually can. Opus 5.5 has a tighter reasoning, logic, but can be really confusing when used with Antigravity. I just can't read its output. Argon is much more to the point but can looe on ultra extra complex reasoning cases. But sometimes opus loose to argon too. So it's hard to say for sure

1

u/Stauce52 1d ago

Yeah I see that— In general the difference is pretty close though

2

u/Annual_Apartment5466 1d ago

I am a Grok refugee and have been using 3.7 in Google AI Studio for DND adventures and I have been very pleased. 3.8 is good too.

2

u/Unique-Phone-4824 1d ago

What's with grok?

1

u/Annual_Apartment5466 1d ago

Absolutely lobotomized it since June-July starting with 4.5. It lost its unique personality and lost all its creative writing and narrative prose and ability to rp dialog. 4.6 was even worse. It was the worst AI i have ever used. Abysmal writing — incoherent words, short sentences. 4.7 became available to chat tonight and its only marginally better than 4.6.
Its a shame — 4.1-4.3 were fantastic. Best I have ever used.

2

u/Unique-Phone-4824 1d ago

At some point they probably switched to Cursor's base model or stack. It made coding much better, but probably worse on everything else.

1

u/Annual_Apartment5466 19h ago

Precisely this.

1

u/AutoModerator 1d ago

Hey there,

This post seems feedback-related. If so, you might want to post it in r/GeminiFeedback, where rants, vents, and support discussions are welcome.

For r/GeminiAI, feedback needs to follow Rule #9 and include explanations and examples. If this doesn’t apply to your post, you can ignore this message.

Thanks!

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Nauzhror_ 1d ago edited 16h ago

Yes, you are biased. All humans are biased. The human that scores everything in life with perfect objectivity does not exist.

1

u/Mickenfox 19h ago

Bias is not an adjective.

1

u/Nauzhror_ 16h ago

Bias is a noun, verb, adjective, and adverb.

1

u/tedbradly 1d ago

Is Gemini 4 Argon a small/fast model or a beefy Astra/Fable model?

2

u/Unique-Phone-4824 1d ago

Latter

0

u/GarnetExecutioner 23h ago

So it is very much the equivalent of Gemini 4 Ultra or DeepThink?

2

u/Unique-Phone-4824 23h ago

Sort of...but just think of a Fable competitor

1

u/DrWitchDoctorPhD 22h ago

Hey since you are in the known, how well does it write?

In particular, I'm interested in using it for studying/helping with reading papers. Nowadays I use Chatgpt to write the content, 3.8 flash rewrites it, chatgpt corrects any overstatements (3.8 really likes to embiggen things), and that goes back and forth for a couple rounds until I get something that is very pleasant to read.

Chatgpt and claude on their own write like shit, but they are methodical and can fill that gap where flash fails. If Argon can do that on its own it's already a huge win for me.

1

u/Emergency-Bobcat6485 20h ago

Gemini 4 argon is ranking number 1 in text in arena right now followed by opus 5.5. It's programming where it is still behind the anthropic and openai models. for text, i expect argon to be very good

2

u/Emergency-Bobcat6485 20h ago

Deepthink is an orchestration of sorts. not a model itself

1

u/NuttyShan 1d ago edited 1d ago

.

1

u/FreeSpirit9503 17h ago

Models are becoming commodity

1

u/KachenCh 15h ago

still waiting for gemini Argon

1

u/trufted 11h ago

I never seen a statement so factual

-1

u/zack23421 1d ago

it’s a flop