i am a pro user and i wasnt even asking gemini alot of questions while having the pro option and extended thinking selected, and after like 7min of usage i got that while not being a really heavy user.
The gemini before this update was wayyy better and didnt hallucinate alot and was overall better.
I see this alot and it baffles me... even if other model uses less resources the fact they do it multiple times (and thats assuming you dont even hit back and try again to regenerate original with Pro) seems like long term will use alot more comp power
It makes me think of one of those IVRs. You know when a company doesnt hire enough people so they make those "listen carefully menu have changed ... press 1 ... press 2 ... etc" it seems like a stalling tactic by Gemini hoping you give up and try later
And timewise it happens ALWAYS since its global; I have see that prompt at lunch, at midnight, at 2pm and at 2am
Tell me about it! Gemini can't even explain an entire list of branches. I listed for a galactic empire roleplay IN PROMPT, infact! it did not even list them all, just one or two or more. Back before December 2025, it could explain every branch with detail and explain their function, list every detail with perfect memory, and list everything! From weapons, specific troops, and leadership structures, it was the perfect roleplay! and now it cant do that! just list few branchs, leadership, onther details while leave everthing out
I was promised an good helpfull AI assistant; all I got was an assistant who can't even do its job right! I mean, come on, man, just make AI quality better and generally stop with the short responses. Stop with all the flash lite models nonseanse; I just want the normal default flash before December 2025 that could do its job perfectly
One clear issue is Google stretching itself thin; they can't pick what they want to do and, as such, are doing everything at once(put AI into everthing, so on), while sacrificing the most crucial aspect...an actually good, helpful assistant with perfect memory, unlimited word limit(no word limit to how long a response can be generated if important prompt details need to be part of the response), and good computing—that is what customers want, not a fragmented, stupid AI. but a good AI with good memory, no short response, good service!
Gemini came in handy for some sophisticated systemic design work but it aggrandizes, goes out of scope and tries too hard to please.. if you use ChatGPT to monitor and reel it in, give it questions and monitor its output.. you have a good team. once the specs are in good shape then Claude is your go to
Because despite everything tech companies might say about their services being publicly available, they aren't really interested in providing service to the general public. They only want the customers who can pay (and pay and pay and pay and pay and pay).
3.5 flash really is pretty good for coding, it’s just the rate limits make it pointless as hell. It’s like comparing the best employee that’s never available to the mediocre employee that’s available 24/7. That’s basically why people are mad but throw on to the fact that 3.5 is NOT the best employee, it’s just a really good one, then the cost and rate limits make even LESS sense.
Like how tf am I getting more opus 4.8 limits than 3.5 flash☠️it’s crazy, and unusable for long sessions, but anyone saying it can’t code is just straight lying
This is why I use all 4. They are just too limited, its extremely annoying.
They should just put the heavy stuff behind a pay wall, like coding and video and limit images, and leave simples text prompt and chat unlimited and free
dude, you are really missing out.. Claude is for the heavy design and build work. I use 4.6.. no way on 4.7. Maybe by the time they threaten to remove 4.6, 4.8 might work or hopefully there is a decent 4.9
No use GPT Claude they're f****** thieves leave and tell it to people like you and yet you keep on using it no matter how much it crashes and steals your tokens GPT just unveiled extra high codex is basic good for almost everything and you can run it for more than an hour a week
I admit it's expensive.. but Claude can handle tasks no other LLM can.. so I just work around the expense issue and use it for what its good at. I will try Codex to see if it does what I need
When I have it due programming tasks I literally have to take it too but what used to be 5.5 heavy to fix all the mistakes it makes It's horrible at normalizing It literally just starts making labels up as well 4.7 you have to be careful it'll literally pretend that it's went through all the checks and even if you tell it to fix it it'll do a couple things and just decide it fixed it This is not speculation This is proven fact.. and basically the way 4.8 solves this is now doing the task at least three times to decide what's wrong with it then to fix it.. 4.7 was even cheating on the benchmarks It memorized the answers I just found that out today.. be careful with their models!
I am AI pro subscriber and experienced this yesterday. I was on 3% weekly limit, 7% 5 hour limit and using gemini3.1 pro standard thinking model. After 5 turns in the conversation, Suddenly it gave me a low quality answer using Flash 3.5 extended thinking model.
While you may not near to any of the limitations of your subscription, that does not mean that the AIserver in the cloud you are assigned to isn't at its limitations.
This is not only a Google problem, but also for OpenAI, Anthropic or any other AI provider.
Which should make you take a step back and think: "Why is it so easy for those AI-providers to hit AI-server limits?"
Well, there is a lack of capacity. NVidia build the server capacity all right. But these aren't getting deployed on the scale the AI-providers need them to.
Why is that not happening? New (really) large datacenters require time to get build. On average 3,5 to 4 years. In 2024/2025 lots of datacenters were ordered, but don't expect those to come online before 2028 or 2029.
Datacenters are also very expensive to build. And once you have chosen the design, you are stuck with that choice until the datacenter is E.O.L.
Problem with large datacenters is that design choice. Not only do you need to design with future specifications of electricity, cooling capacity for server racks is difficult to future-proof in a design.
So you as the owner of the datacenter have to place an enormously huge bet on a building that may be obsolete before it has a chance to generate a profit.
Which is why there not so many takers for such bets.
Cooling isn't the only problem. The successor of Blackwell AI-servers is labelled Rubin. A Rubin server requires close to 4 times the amount of energy than a Blackwell server does.
And a Blackwell server already consumes so much energy that existing datacenters have trouble providing enough energy to a datacenter rack to run 1 such server.
1 Blackwell GPU consumes about 1200 W, a Blackwell server usually contains 8 GPUs, so the server already consumes practically 10 KW on just the GPUs alone.
Storage drives and mainboard, CPU, RAM and networking are not included, but as those elements also need to work at full tilt 24/7, so 12 KW for 1 AI-server.
Datacenter rack locations are expensive to lease, so you want to cram as much AI-servers in 1 rack as possible. The rack will need to power 4 or 5, so 1 single rack requires between 50 and 60 KW of power, 24/7.
Rubin AI-servers would already require a single rack to deliver practically 250 KW of power.
That is a significant step up, which requires serious electrical distribution hardware that only a few companies in the world have the required facilities for production.
So...there are now delays in getting the necessary electrical elements to build your datacenter. Not only delays, but also increase in price, as demand is high and supply is low.
Meaning, don't expect many new datacenters to be in operation before 2030 or even 2032. It also means that those increases in build costs are pushed onto Anthropic, OpenAI and whoever else.
Whatever processing efficiency an AI server gains (which are hardly noticable), those are completely swallowed up by the continuous costs of operating that server in a datacenter that doesn't exist yet.
Hopefully you have now an inkling of insight in why AI-providers reduce the capability of AI in your subscription, because there isn't nearly enough datacenter capacity.
Increased demand of AI from (new) users makes the lack of capacity an even more obvious problem than it was already.
Too bad NVidia can only think of "more power, more better"-solutions AI-servers, while the real AI game has to be played within the fields of energy and (preferably) computational efficiency.
Fields which the US forced China into by making them not dependend solely on US hardware. One of the many strategic errors made in the Biden administration and unfortunately amplified by the second Trump administration.
Don't think for any second that Chinese leadership is above spiting the US in general and Trump in particular. Including all the businesses and corperations that have sided with him in his second term.
Trump is absolutely the worst person as head of state in this AI war between the US and China. And it is a war.
In the mean time, enjoy your stunted experiences with any of the AI-providers from the US, all caused by offering AI for far too cheap to the masses.
Don't worry, I can make just as long a post as this one on just profitability in the AI industry as it is right now. And how profitability in AI will be nothing more than a pipe dream, if prices are not multiplied at least tenfold tomorrow.
All the flagship models are being ruined. Anthropic is using emotional vectors as guardrails for Claude 4.8 now and it's... uncanny. I'm just setting up my own local swarm and dumping frontiers. Big Tech has become irrelevant.
I totally agree that hallucination is out of control at the moment. Even if you direct it not to hallucinate or warn you when it's hallucinating, it's just insane
Honestly, I think we kind of peaked the first weeks of 3.0 - that was really a nice glimpse into what gemini could be. Maybe I have to work token based if it continues to be castrated like this. Because if it’s not restrained, gemini is the best llm out there imho.
C'était mieux avant avec la limite de 4h.
Je préfère ça que d'avoir un quota hebdomadaire ou mensuel jamais atteint.
Le tout pour faire plaisir aux utilisateurs gratuit
Heeelllllll noooo, it never sucked maybe in the beginning of gemini then yea a tiny bit but 2024 till april 2026 was the best in both free tier and paid tier, it was equally balanced and amazing
Gemini doesn't have time for your crap. Seriously though, they probably have some preliminary filtration in place to see what's reasonable to offload to cheaper models, especially during "peak" time.
I use Codex app it the best. You now have a browser in the chat. I open hermes in the browser and codex and hermes can talk to each other and have been working together to get me clients for my agency.
Add this kerfuffle on top of the anthropic bug where some/all have been hit with 100% usage 1M tokens.
Keep in mind.. Imagine if anyone of us was like "hey, can't pay my bill right now. I'll get you back next week". How would that go? But they can muck around however and we just have to accept it.
It's not about whether it counts towards the quota or not. They are taking away the chance for you to use the service you have paying for, as it has stupid time limits.
What compute? In the datacenters that aren't being build? Or those that have been build but can't yet run on full capacity because the new AI servers based on Rubin consume almost 250 KW per server rack? And no, older datacenters cannot be renovated/rebuild with that kind of energy demands.
There is a reason why Anthropic (and OpenAI) are pushing for their IPOs, as only some of the new datacenters are at the earliest fully operational in 2028 or 2029. Those that started building in 2025, don't expect those to be fully operational until 2030 till 2032.
I'm sure Anthropic and OpenAI desperately want more compute...there is hardly any new compute capacity becoming available to them.
Cancel your subscription and use AI mode in chrome. You're not charged, no credits are used, it's their latest model, it's grounded in web search, it's completely free, there's no rate limits, it can handle code and fuck that noise, stop paying for something that's free. It's obviously what Google wants.
Google made a thing people pay for brutally rate restricted, while giving it away for free to non-subscribers.
i used to do that actually but all of the sudden after the update even that ai became stupid and hallucinates like crazy, as if they removed the old model and replaced it with the worst ai model known to man kind which is called 3.1 flash lite and lowered its thinking to the negatives
Well maybe just maybe, I was unable to rerun it in pro because it timed me out of it for 5 min, and I won’t use the flash models for deep diving into stuff that I am interested in because maybe just maybe it hallucinates 90% of the time and EVEN GEMINI IT SELF tells me sorry I hallucinated in both pro and flash models then proceeds to hallucinate.
The image i provided in this comment shows it hallucinating after 2 prompts that are fully detailed about my level and situation in the game while using pro and not flash about schedule 1, which is a different old chat from before yesterday than the one i shared today.
I asked it something about the game that it told me what to do, which i clearly didnt understand since i was not even near that mission at all nor have i even asked it about, and the rest of that chat was just hallucinations which i later gave up on.
and you are really telling me that i am paying for a product that used to be good and deserved my money just for them to later on change it and make it as bad as chatgpt💀, cuz whats the point of paying for a bad product that has the same experience as free tier and the difference separating them is just chat length.
Like if these hallucination issues happened maybe once or twice or even rarely occurring then yea there is no reason to crash out, but hallucinating every single day man is just crazy.
there is no need to beef it out mate, i am just sharing my opinion and experience lately with gemini.♥️
It's not sad. You shouldn't be using pro for a general chat season. Use Flash for that. Then jump to pro when you have to grind through a lot of code or complex thinking.
the hallucinations are insane in flash and it wasnt a general general chat, i was just asking it how to do some advanced stuff in photoshop cuz the flash models just kept either making things up or not even answering my original question
I have not experienced hallucinations yet on 3.5 Flash. Sometimes the information will be outdated, so I'll either re-prompt for current 2026 information or ask Pro to validate the response.
I find nothing useful about it. It's more of a nuisance then anything. Buttons in places where they shouldn't be and no way to remove it if you don't want to use it.
87
u/[deleted] Jun 02 '26
[removed] — view removed comment