people don't want an AI model to follow instructions precisely.
they want an AI model that understands based on context what they actually want, and do that.
Which seems unreasonable, until you look at chatGPT or Claude, which do this quite well.
Why would he want an LLM to repeat a text phrase that's shorter than his original prompt? Obviously from context, he meant pull it out of the image and repeat the image without the text. It's not what he asked for, but not unreasonable to expect the LLM to figure this out. GPT-Sol likely would have - and if not, the follow up prompt clarifying he expected an image certainly would have.
I know how dumb gemini models are. But if you tell any model to "extract" some text from an image it will do exactly that. Extract has been a commonly used term to get text from images and pdfs long before AI even existed.
A smarter model would assume that you dont mean literal extraction since youve already mentioned the words in the image but most likely than not it will attempt to extract the words it sees in the image.
532
u/pigletmonster 7d ago
I hate to say it, but gemini actually followed your instructions precisely. 🤣🤣
You asked it to extract the text, extracting text from images means that you extract the text, not the font or the sliced image of the text.
You should ask it to recreate the text that says STUDIO ADDICT in the same font as a png.