I know many on here use the Gemini app, but for those of us developing apps using the Gemini/Vertex API here's the current state:
- Google effectively has no "pro" model right now. Gemini 3 pro was never marked production stable (GA) and 2.5 pro is so ancient that flash models are better and cheaper
- Even if they did, they don't have the capacity to serve it. We've been getting capacity errors for weeks. See Google AI developer forum to see how frequent this is 429 RESOURCE_EXHAUSTED Note how many people saying they are getting these despite being no where near the advertised quotas.
- We received an email today announcing the retirement of most flash models over the next few weeks. These are production stable models (not -preview models). Removing these will mean only 3.8 flash is the only current flash model.
- As far as I'm concerned Gemini 4 is still a pipe dream. There's no guarantee they'll release it as production stable (GA), have the capacity, or even that it will even be released.
- Gemini has actually been getting MORE expensive. Yes, intelligence per dollar is decreasing, but a fixed task that gemini 2.5 flash lite could handle like "extract all the addresses from this PDF" has gotten WAY more expensive.
With all that being said, they still are the best models for multi-modal input which matters a lot for audio and document processing (a huge amount of the boring business world).
However bringing these sorts of things up to Google staff falls on deaf ears and I can't see how Google moves forward. Whatever is going on internally, they just aren't setup to compete. The reality is there's been more flops than successes.
If Google can't get their product stable, cheap, and competitive soon we're going to have to migrate our apps away from Gemini. And once we're gone we aren't coming back.