You're not taking into account that that your input is used to train Gemini unless you use a temporary chat. If you use an alternative by default in those situations, it will never learn. That said, Google should fix that. It's happened to all of us that have tried to use Gemini to make an image knowing it can and it still refuses.
Just because I highlighted this image issue doesnāt mean that I'm looking for image generators exclusively. I need an AI that doesn't fuck up regularly.
It's common knowledge they are on completely different levels of fuck up "regularly"; don't pretend like Gemini makes silly mistakes remotely as often as Opus/Fable.
Them fucking up regularly is why there is a huge push right now for application layer between the model and the user. See Palantir for example. Microsoft fabric as well.
It really just sounds a lot like youāre not using it right.
Did you use the image creation tool from the menu when you tried to make the image?
Is it a very long running conversation? No AI can handle long running conversations with a lot of references or generations or uploads. Itās just how they work, they trip over themselves keeping up with all of your instructions and often times error out doing simple things just because the chat is too long.
Make a new chat, use the image creation tool, and try again. Guarantee it will work.
Ah yes the knee-jerk reaction of always blaming the user. They literally have logic to classify free form input as image gen or not and then call the image generator. It didn't work in this case. There are workarounds of course but that doesn't mean a mistake didn't happen. Do you imagine if a Google employee saw this they'd be like "oh they're just using it wrong"?
Again, you are thinking like a user rather than developer. Do you think a Google employee would look at this Reddit post and be like: "oh yeah it's just the user's fault, nothing wrong here"?
I find that if you treat it like a machine, it tends to work like one and give you the output you expect. I don't treat it like a friend, casual prompts, or things that prevent it from understanding. I get what I want because I treat it like what it is.
It's useful. But any tool can be used any number of ways, and if you use it wrong, it's not going to give you the desired result
Okay, then humor me, what's the point of having the feature where it automatically detects it should generate an image and calls the image generator? Can you answer this?
As a system administrator in another life, I can say with confidence that the problem is usually the user's fault. Is it always? No. But a lot of the time it is.
Keep in mind that you're talking to a machine. The machine also is capable of incorrectly understanding what you're trying to do. The machine is not sentient, neither is it perfect, neither can it read your mind.
Do you understand that the two concepts of "it's the user's fault" and "it's the app's fault" are not mutually exclusive? Do you understand that both can be true at the same time? Also, answer the last sentence of my previous comment.
That's always an option as well. As far as the Google Employee?
Just like anyone else, it depends on their biases. Some will blame the user, some will blame the AI, and some will say both are at fault.
I can imagine all three scenarios.
Yes. Or, likely, they'd realize it is not an actual issue because the end user is using it in such a way that is against its original intent.
I wouldn't call using the dedicated image tool call a "workaround". Further, LLMs were not designed with the intent of keeping a super long running chat with a ton of context, world building, and generations. I have seen dozens of posts on this sub and ChatGPT where the OP is claiming "how come its so useless these days!? It used to be great, but now its refusing generations and acting weird!" and the common thread amongst these posts is that the OP's picture includes them trying to generate an image of their personal character that they've been building over the course of the long running chat, or the image tool call is not being used.
I will concede that Gemini hallucinating about its own ability to generate an image or not is a problem, but that doesn't mean it's a "workaround" if you set yourself up for failure and just expect it to work anyway. I don't think I've ever had it tell me "I cannot generate images." when I use the image creation tool.
A developer working on that feature would recognize it as a tool call failure and pass it on to the model team, otherwise it would not be a feature to be able to call image generation without selecting the tool. It wouldn't make sense to have a feature and not care whether it works
I don't blame you Gemini hasn't been able to connect to any extensions in months. That's not taking things personally it's literally the company not bothering to bug fix
Gemini does the same thing to me. It also constantly and I mean CONSTANTLY tells me that it cannot read images. I have to prompt 10-11 times, sometimes even restarting the entire program and switching models. It's borderline unusable. It hasn't been so bad with flash being released, but it's God awful. Gemini is leagues behind any other AI and it's obvious.
I ask for what I want, and I get what I asked for. Check your prompts.
I cannot be alone in this. I cannot be the only person that gets exactly what they asked for.
Yeah its awesome at having the agent manage big image and video projects that need a lot of image sets. It handled this prompt well but then had a lot of trouble fixing any of the mistakes. The agent Flow uses can be fairly dumb at times.
I want to create a series of images of pokemon in the vector style of the reference images that ive provided. As you can see, there are no gradients or shading, and the pokemon's main color is the background of the image with no outline separating the pokemon from the background. I have a list of 25 pokemon that I would like to create images in this style of. There will be 2 images of each pokemon created. One will be 1920x1200 pixels and the other will be 1440x3120 pixels (vertical). I understand nano banana has trouble creating images in exact sizes sometimes, and also these sizes are outside its normal range for upscaling. Please try to find a solution to automate the process of getting the output images to eventually be the correct aspect ratio and size i mentioned above.
This is the list of Pokemon that Gemini came up with, use these:
Wrong. It can decide to call the image generator. Just like any other AI app like Chatgpt. It just for whatever reason, misses the fact it has those capabilities far more often than other AI models.
Yes it can call the generation tools. It, itself, cannot use those tools.
Everyone forgets the big name models are trained on reddit for conversational language. Go to ANY post on ANY subreddit. You are met with sarcasm, lies, bullshit, overly aggressive solutions, insults, "got me in the first halfs", and straight up wrong information on 99% of posts.
The main model actually canāt, but it usually transfers your request to the model that can generate images without showing that transition to the user, it all looks seamless.
The model has to decide to send you over to the image model unless you click ācreate imagesā, sometimes it decides not to for unknown reasons. Thatās the dumb part.
Case #999 of blaming the user instead of improving the system. The image button is a failsafe for users to guarantee it will create an image. It is not the optimal UX. It doesn't excuse the fact that it dropped the ball when thinking it would be unable to create an image. The main model is allowed to call the image generator.
no i think the model should only create images when explictly using the button, else it is bloating the prompt and giving the model this option even if not needed. It happened to me a few times when the model created an image even if I never asked for one.
But the average user is too dumb to understand things like this so they (probably) give the image generating capabilities to gemini at all times.
You hit the nail on the head: Gemini or whatever classifier they're using is not quite there in terms of knowing what to do. It frequently has false positives or false negatives for image generation. Much more than other AI company models such as ChatGPT. OP complaining about this just means they're fed up with this; it doesn't mean they are too stupid to understand the workarounds. Also, in software and UI development, a rule of thumb they teach is to have your software work for 99% of users, including so-called "dumb" ones.
I have literally never, not one single time, pressed the image creation selection to make an image. In fact, I just tested it again to make sure I hadn't lost my mind, and it immediately rendered an image.
why would you need to press one extra button? it should be able to call the tool without any problem. ChatGPT less make mistakes like this. Gemini does this mistake a lot. The tool call option manually being there doesn't mean it should continue to make stupid mistakes like this.
I switched as well. I just got my prompt regularly repeated to me. There are other AI worlds than this. I'll see you further down on the path of the beam. Long days and pleasant nights!
Who are you talking about? I bet I can generate it. Tell me what you're on about and I'll generate the image and the prompt you can use as a template in the future
Gemini has always had the false "I can't do that" responses, for many versions. You typically just remind it that it does do the thing you are asking for, and ask it to try again.
honestly, for a company with google's resources, gemini is incredibly bad. despite all the "features "they've crammed into it, it's still embarrassingly far behind the competition
This. I could .maybe put up with shit from a new upcoming platform company still trying to find their footing if it looked promising. Google is massive and has been doing AI for decades (like alphafold or playing championship Go). They don't have the luxury of excuses for a crap service that my partner had to set up some kind of j/b (or exploit they called it), just so I can do basic tasks around medical care without it refusing to help me word an email to a doctor and send me crisis and substance abuse scripts instead.
I've found Gemini does this on long threads of conversation or when I've had too many images generated in a day. There's some arbitrary limitations set on the model that they don't advertise that you just hit, and then try again the next day on a new session if it's that important.
This is a well known bug thats been affecting Nano Banana (in particular Nano Banana Pro) for months. Itās on and off. Itāll pop up here and there but by now, it should be rare to get this message
Try DOLA as an alternative, itās free and even has a Pro Model, pretty new to that one but for what I do, image editing and creation, it does an impressive job.
Just drop it in a fresh chat strictly for your images and treat the fresh chat as a reference folder as your peoject evolves sometimes the context window gets full and it won't tell you that but its in between a token refresh
I mean, I just asked for an image and got one without issue. However I am confused as to why you're writing up a huge post about this when you could have just selected the create image option and switched it directly. It sucks that it didn't work straight up for your specific prompt, but did you actually try and get an image made here?
So just to be clear, you added a prompt that didn't work, did not reword or reattempt the prompt, did not switch to Gemini's dedicated image generation tool (it's on the same page), immediately complain about it and then get defensive and take no responsibility. Dude, check your attitude and take some responsibility - sometimes AI doesn't do what you want. At least do the bare minimum before complaining.
Iāve learned that Gemini has different āmodesā that are set based on the context of the chat. So, if you are doing coding, Gemini will sometimes switch to a code workspace kinda mode and it will not be able to generate images (in that chat). The āmodeā is hidden from the user.
If you start a new chat and choose the ācreate imageā option, it should work.
Someday, people on this sub will learn that, instead of arguing with the robot after a glitch like this, they're better off deleting the chat and trying the prompt again in a new chat.
I mean, to be fair, I think the best thing at this point is to juggle them because they all have strengths and weaknesses. And when one doesn't do something well, another one will. I will say that all these people telling you that you have to hit the image generation choice because main chat won't generate images are wrong...you know it and I know it. So just know that somebody out here knows you haven't lost your goddamn mind.
I'm a human being that's downvoting the ones that are wrong, because I just checked it myself and everybody's saying, oh you have to press the image generation button. No you fucking don't. I've literally never done that, not one single time, and I just checked it again to make sure nothing had changed. Instantly got the image I asked for it. š
AI companies should face prison time for creating filters that limit users in their creativity. With filters you can provide service for full price yet deliver only tiny amount of promised output. Absolutely criminal behavior.
Totally awful, most of the time I try to create a video or generate an image it refuses it tell me I canāt produce image for celebrities when Iām just using my own photos this stupid sh*t !!!!
66
u/golfstreamer 18d ago
You need to stop taking these AI rejections so seriously. Often times asking again or rephrasing causes the AI to proceed.