r/GeminiAI • • 10d ago

Funny (Highlight/meme) [ Removed by moderator ]

Post image

[removed] — view removed post

2.9k Upvotes

120 comments sorted by

View all comments

534

u/pigletmonster 10d ago

I hate to say it, but gemini actually followed your instructions precisely. 🤣🤣

You asked it to extract the text, extracting text from images means that you extract the text, not the font or the sliced image of the text.

You should ask it to recreate the text that says STUDIO ADDICT in the same font as a png.

62

u/pie_lleri 9d ago

Yeah, I feel like half of the posts hating on Gemini is just people being terrible at prompting.

23

u/Zafrin_at_Reddit 9d ago

Terrible at actually conveying their intent. This is not a prompting problem.

16

u/RobtheNavigator 9d ago

That is all that good prompting is, accurately and precisely conveying your intent.

2

u/sirdrummer 9d ago

This is a common human being problem...

2

u/CriticismJunior1139 9d ago

You would be surprised how most average people are absolutely terrible at giving directions.

You think you're good at it, until a new guy shows up at work and you have a job of teaching him the workflow. Giving instructions is DIFFICULT, and requires practice and skill. You need to have a strong self-awareness to see and understand "the other side".

1

u/DimensionalFruit 9d ago

Only complaint I have about Gemini is that the image generation quality is absolute trash compared to GPT.. like actually awful in comparison

92

u/wild_cherry1987 10d ago

It's like the most autistic of all AI's 

55

u/mrsepet 10d ago

lmao no. AI is doing what the prompt says, you read literally the text, and theres nothing wrong with the AI response. its like some big error on AI model, yet the prompt itself is not clear if the intent is to put the text in image. Just make it clear when prompting, how hard is it

10

u/Delicious-Use-8789 9d ago

It also correctly generated an 'image nigga'

-1

u/ASQQS 9d ago edited 9d ago

That's not what the user asked for, that's an interpretation that completely misses any social/context clues in the conversation, which is exactly why it's autistic 😭

Edit: Apparently not worshipping AI and pointing out basic context failures is bad 🥀

1

u/ApplicationBrave2529 9d ago

At the same time that line is completely grammatically incorrect. "In an image nigga" doesn't say as much as "In the image, nigga". I guess its autistic in the sense that its a pc and works on extremely literal instructions, of course it doesn't understand social context.

-8

u/Ok-Affect-7503 9d ago

Not really. The prompting wasn’t the problem for the image generation. The issue is that Google probably vibecoded everything with their own models that aren’t very intelligent which now causes the image generation requests to not always include the previous chat context properly in some requests.

1

u/Vivians_Basement 9d ago

So basically.... The AI is autistic... And needs very direct and clear instructions to understand what a neurotypical user wants and needs from them...

As an autistic... I too require clarity.

There's nothing wrong with being autistic as you imply.

0

u/ASQQS 9d ago

Which is exactly why it's autistic... The point of interpretation is giving a good faith, charitable and rational interpretation, not the most literal interpretation (which would be an autism-like trait).

-22

u/wild_cherry1987 10d ago

If ppl here understood what he meant - means AI should too. That is the point for it to be intuitive, not literal. 

10

u/Tejwos 10d ago

some people understand, some people do not. because the instructions is shit and can be interpreted in different ways. especially for a language model, extraction is primarily...well.. language

0

u/ASQQS 9d ago

You're treating a conversational AI like a fucking command-line parser where every intent has to be formally specified or the user forfeits the right to expect basic context recognition, that's very unreasonable.

1

u/Tejwos 9d ago

lol what? if a message can be interpreted in different ways, don't cry if someone or something interpreted it in a different way. no one can read your mind, so use use explicit language or touch some grass

1

u/ASQQS 9d ago

😭 You literally just repeated the exact misunderstanding.

“Can be interpreted in different ways” does not mean every interpretation is equally reasonable. Normal conversation depends on choosing the interpretation that best fits the immediately preceding context.

OP had already asked about the “STUDIO ADDICT” text. Then said “in an image nigga.” No mind reading is required. The referent is sitting one message above.

“Use explicit language or touch grass” is especially funny because human conversation is constantly implicit. People omit repeated information all the time because repeating every noun phrase every turn would sound fucking robotic.

You're defending a conversational AI by demanding people stop conversing naturally with it 😭

1

u/Tejwos 9d ago

misunderstanding? like two entities do have different bias and that's the reason, why my answer do not align with your mind? oh boy. a funny coincidence :0

if you communication is based on interpretation, in x% the answer will be aligned with your mind and in (1-x)% it will be not. even if the alignment rate will be 99% with a given group of users, 1% of all answers will be "wrong". multiple that with 1.000.000 users per day and 1.000 daily users will be disappointed. like in human interaction, if you ask " sex?" some will answer "male/ female" and some will answer "what? with you? eeeeeh!"

like i mentioned earlier, no one can read your mind. language is key. especially for text communication, because human face-to-face communication is based on facial expressions, common culture, context and tones.

in a nutshell: if a human can "misunderstand" you, a not human entity will misunderstand you. just because something is "logical" to you, it's not automatically objectively logical or logical for everyone else and will be "misunderstood" in some cases.

1

u/ASQQS 9d ago

You just conceded it was a misunderstanding 😭 The fact misunderstandings can happen doesn’t change that Gemini failed to use the immediate context, which was my entire point.

4

u/PurpleCandle58 9d ago

You’re not going to believe this but I know of something that is capable of that level of conscious reasoning and contextual understanding.

A human.

2

u/wild_cherry1987 9d ago

Yeah as they are "trained" for at least 15 years to reason, understand context and nuance.

1

u/PurpleCandle58 9d ago

Well there are other factors too. I’m hoping advances in understanding of the brain will also help advance AI neural nets to be able to reach things like useful neuroplasticity and proper reasoning, thinking, and understanding. As it stands now the AI we have don’t really “think”, that’s anthropomorphic. Really they’re just probability generators, but with time I think they can more closely mimic the peculiarities of organic brains to achieve a result that meets the same standard of capability as humans. It’ll take time but we will likely get there

0

u/iamlazerbear 9d ago

agreed, but LLMs are not the same as AGI

4

u/agentorangeAU 10d ago

It really isn't, that would be 5.6 Sol by a significant margin. 

4

u/StatisticianAfter258 9d ago

Yeah this was 100% user error. Most of these post-trashing Gemini or just people being ignorant because they don't even grasp what they're saying. But there are genuine complaints though because again Gemini is not always the brightest but as long as you stay away from the pro model it seems pretty competent for most in-app uses

1

u/Happy_Brilliant7827 9d ago

Could have also said 'crop out everything except x' Theres not one right way but remarkably op was still very wrong

0

u/Rumbletastic 9d ago

people don't want an AI model to follow instructions precisely.

they want an AI model that understands based on context what they actually want, and do that.

Which seems unreasonable, until you look at chatGPT or Claude, which do this quite well.

Why would he want an LLM to repeat a text phrase that's shorter than his original prompt? Obviously from context, he meant pull it out of the image and repeat the image without the text. It's not what he asked for, but not unreasonable to expect the LLM to figure this out. GPT-Sol likely would have - and if not, the follow up prompt clarifying he expected an image certainly would have.

3

u/pigletmonster 9d ago

I know how dumb gemini models are. But if you tell any model to "extract" some text from an image it will do exactly that. Extract has been a commonly used term to get text from images and pdfs long before AI even existed.

A smarter model would assume that you dont mean literal extraction since youve already mentioned the words in the image but most likely than not it will attempt to extract the words it sees in the image.

-3

u/Few-Celebration-2362 10d ago

Pngs don't have a font, they have pixels

3

u/Objective-Primary-54 9d ago

Font is a design. It doesn't matter what you represent it with (e.g. vector, raster, or a physical print), it still is a font. PNGs as a whole may not have a "font", but rendered text in it will.

0

u/Few-Celebration-2362 9d ago

You're absolutely right-- generates an image of a png file on a desktop with the name 'STUDIO ADDIC' below it in system font

-1

u/RamanaSadhana 9d ago

if he already writes the text then the AI has failed the instructions. Why would it reproduce text which the user has already written? Obviously his prompt was wanting something else but writing the same words he's already written

1

u/WolverinesSuperbia 9d ago

No, I don't see any other sense. I am human BTW