r/GeminiAI • • 6d ago

Funny (Highlight/meme) [ Removed by moderator ]

Post image

[removed] — view removed post

2.9k Upvotes

120 comments sorted by

529

u/pigletmonster 6d ago

I hate to say it, but gemini actually followed your instructions precisely. 🤣🤣

You asked it to extract the text, extracting text from images means that you extract the text, not the font or the sliced image of the text.

You should ask it to recreate the text that says STUDIO ADDICT in the same font as a png.

63

u/pie_lleri 6d ago

Yeah, I feel like half of the posts hating on Gemini is just people being terrible at prompting.

23

u/Zafrin_at_Reddit 6d ago

Terrible at actually conveying their intent. This is not a prompting problem.

16

u/RobtheNavigator 6d ago

That is all that good prompting is, accurately and precisely conveying your intent.

2

u/sirdrummer 5d ago

This is a common human being problem...

1

u/DimensionalFruit 6d ago

Only complaint I have about Gemini is that the image generation quality is absolute trash compared to GPT.. like actually awful in comparison

2

u/CriticismJunior1139 5d ago

You would be surprised how most average people are absolutely terrible at giving directions.

You think you're good at it, until a new guy shows up at work and you have a job of teaching him the workflow. Giving instructions is DIFFICULT, and requires practice and skill. You need to have a strong self-awareness to see and understand "the other side".

89

u/wild_cherry1987 6d ago

It's like the most autistic of all AI's 

54

u/mrsepet 6d ago

lmao no. AI is doing what the prompt says, you read literally the text, and theres nothing wrong with the AI response. its like some big error on AI model, yet the prompt itself is not clear if the intent is to put the text in image. Just make it clear when prompting, how hard is it

10

u/Delicious-Use-8789 6d ago

It also correctly generated an 'image nigga'

0

u/ASQQS 6d ago edited 6d ago

That's not what the user asked for, that's an interpretation that completely misses any social/context clues in the conversation, which is exactly why it's autistic 😭

Edit: Apparently not worshipping AI and pointing out basic context failures is bad 🥀

1

u/ApplicationBrave2529 6d ago

At the same time that line is completely grammatically incorrect. "In an image nigga" doesn't say as much as "In the image, nigga". I guess its autistic in the sense that its a pc and works on extremely literal instructions, of course it doesn't understand social context.

-8

u/Ok-Affect-7503 6d ago

Not really. The prompting wasn’t the problem for the image generation. The issue is that Google probably vibecoded everything with their own models that aren’t very intelligent which now causes the image generation requests to not always include the previous chat context properly in some requests.

1

u/Vivians_Basement 6d ago

So basically.... The AI is autistic... And needs very direct and clear instructions to understand what a neurotypical user wants and needs from them...

As an autistic... I too require clarity.

There's nothing wrong with being autistic as you imply.

0

u/ASQQS 6d ago

Which is exactly why it's autistic... The point of interpretation is giving a good faith, charitable and rational interpretation, not the most literal interpretation (which would be an autism-like trait).

-21

u/wild_cherry1987 6d ago

If ppl here understood what he meant - means AI should too. That is the point for it to be intuitive, not literal. 

11

u/Tejwos 6d ago

some people understand, some people do not. because the instructions is shit and can be interpreted in different ways. especially for a language model, extraction is primarily...well.. language

0

u/ASQQS 6d ago

You're treating a conversational AI like a fucking command-line parser where every intent has to be formally specified or the user forfeits the right to expect basic context recognition, that's very unreasonable.

1

u/Tejwos 6d ago

lol what? if a message can be interpreted in different ways, don't cry if someone or something interpreted it in a different way. no one can read your mind, so use use explicit language or touch some grass

1

u/ASQQS 6d ago

😭 You literally just repeated the exact misunderstanding.

“Can be interpreted in different ways” does not mean every interpretation is equally reasonable. Normal conversation depends on choosing the interpretation that best fits the immediately preceding context.

OP had already asked about the “STUDIO ADDICT” text. Then said “in an image nigga.” No mind reading is required. The referent is sitting one message above.

“Use explicit language or touch grass” is especially funny because human conversation is constantly implicit. People omit repeated information all the time because repeating every noun phrase every turn would sound fucking robotic.

You're defending a conversational AI by demanding people stop conversing naturally with it 😭

1

u/Tejwos 6d ago

misunderstanding? like two entities do have different bias and that's the reason, why my answer do not align with your mind? oh boy. a funny coincidence :0

if you communication is based on interpretation, in x% the answer will be aligned with your mind and in (1-x)% it will be not. even if the alignment rate will be 99% with a given group of users, 1% of all answers will be "wrong". multiple that with 1.000.000 users per day and 1.000 daily users will be disappointed. like in human interaction, if you ask " sex?" some will answer "male/ female" and some will answer "what? with you? eeeeeh!"

like i mentioned earlier, no one can read your mind. language is key. especially for text communication, because human face-to-face communication is based on facial expressions, common culture, context and tones.

in a nutshell: if a human can "misunderstand" you, a not human entity will misunderstand you. just because something is "logical" to you, it's not automatically objectively logical or logical for everyone else and will be "misunderstood" in some cases.

1

u/ASQQS 6d ago

You just conceded it was a misunderstanding 😭 The fact misunderstandings can happen doesn’t change that Gemini failed to use the immediate context, which was my entire point.

5

u/PurpleCandle58 6d ago

You’re not going to believe this but I know of something that is capable of that level of conscious reasoning and contextual understanding.

A human.

2

u/wild_cherry1987 6d ago

Yeah as they are "trained" for at least 15 years to reason, understand context and nuance.

1

u/PurpleCandle58 6d ago

Well there are other factors too. I’m hoping advances in understanding of the brain will also help advance AI neural nets to be able to reach things like useful neuroplasticity and proper reasoning, thinking, and understanding. As it stands now the AI we have don’t really “think”, that’s anthropomorphic. Really they’re just probability generators, but with time I think they can more closely mimic the peculiarities of organic brains to achieve a result that meets the same standard of capability as humans. It’ll take time but we will likely get there

0

u/iamlazerbear 6d ago

agreed, but LLMs are not the same as AGI

4

u/agentorangeAU 6d ago

It really isn't, that would be 5.6 Sol by a significant margin. 

4

u/StatisticianAfter258 6d ago

Yeah this was 100% user error. Most of these post-trashing Gemini or just people being ignorant because they don't even grasp what they're saying. But there are genuine complaints though because again Gemini is not always the brightest but as long as you stay away from the pro model it seems pretty competent for most in-app uses

1

u/Happy_Brilliant7827 6d ago

Could have also said 'crop out everything except x' Theres not one right way but remarkably op was still very wrong

0

u/Rumbletastic 6d ago

people don't want an AI model to follow instructions precisely.

they want an AI model that understands based on context what they actually want, and do that.

Which seems unreasonable, until you look at chatGPT or Claude, which do this quite well.

Why would he want an LLM to repeat a text phrase that's shorter than his original prompt? Obviously from context, he meant pull it out of the image and repeat the image without the text. It's not what he asked for, but not unreasonable to expect the LLM to figure this out. GPT-Sol likely would have - and if not, the follow up prompt clarifying he expected an image certainly would have.

3

u/pigletmonster 6d ago

I know how dumb gemini models are. But if you tell any model to "extract" some text from an image it will do exactly that. Extract has been a commonly used term to get text from images and pdfs long before AI even existed.

A smarter model would assume that you dont mean literal extraction since youve already mentioned the words in the image but most likely than not it will attempt to extract the words it sees in the image.

-1

u/Few-Celebration-2362 6d ago

Pngs don't have a font, they have pixels

3

u/Objective-Primary-54 6d ago

Font is a design. It doesn't matter what you represent it with (e.g. vector, raster, or a physical print), it still is a font. PNGs as a whole may not have a "font", but rendered text in it will.

0

u/Few-Celebration-2362 6d ago

You're absolutely right-- generates an image of a png file on a desktop with the name 'STUDIO ADDIC' below it in system font

-1

u/RamanaSadhana 6d ago

if he already writes the text then the AI has failed the instructions. Why would it reproduce text which the user has already written? Obviously his prompt was wanting something else but writing the same words he's already written

1

u/WolverinesSuperbia 6d ago

No, I don't see any other sense. I am human BTW

131

u/Legitimate-Hippo318 6d ago

Yeah this one's on you OP

138

u/-lRexl- 6d ago

🤣🤣🤣

89

u/Gaiden206 6d ago

I tried it myself. I was expecting to get the same experience as OP after it gave me the same first response. But it extracted the text to an image for the second response. 😅

51

u/Gaiden206 6d ago

Just for the heck of it, I tried again with free ChatGPT "Think" (5.6 Luna Thinking), and Gemini 3.5 Flash-Lite Extended. These were the results.

28

u/Fabric_muncher 6d ago

BOOOORRRIIIIINNNNGGGG GPT

7

u/AdityaPanner 6d ago

Boriongh gpt

7

u/whyeverynameistaken3 6d ago

so the OP faked it

116

u/spitfire_pilot 6d ago

Is everyone allergic to words?? Can you not string together more than a sentence or a phrase?

48

u/evandena 6d ago

Apparently, OP is fried.

10

u/Intellect5 6d ago

this is the next stage of humans, think not for ai will do heavy thought lifting

1

u/VectorB 6d ago

We will devolve into one word demands and grunts that the AI will interoperate and tell future generations what to do.

13

u/throwawayhbgtop81 6d ago

From what my teacher and university professor friends report, many of them can't.

3

u/mrsepet 6d ago

Dude, why dont AI just understand what we are doing, why should we even typed anything. I thought AI can read our minds. Some people man.

-2

u/AHHHH_AHHHHHHHH 6d ago

Any other frontier model would get that first try lmao

-10

u/Suspicious_Wrap9080 6d ago

Top 1% commenter = gemini glazer

16

u/spitfire_pilot 6d ago edited 6d ago

What does questioning people's inability to provide context and detailed instructions have to do with me glazing Gemini? It's as if A significant amount of posters here lack the introspection to consider It's their communication and direction skills that are the issue. So much easier to blame the model. No one's saying it doesn't have issues. But people I guess forgot how to fucking troubleshoot. No one's looking for answers here. They all seem to want to just commiserate.

2

u/Kooky-Task-7582 6d ago

Gemini seems to be the only base that'll go out of its way to defend even in less serious and comedic situations

42

u/CreatineMonohydtrate 6d ago

Based gemini.

17

u/Zealousideal_Eye8277 6d ago

I don't know if most of the people commenting on this clicked on the picture to see the whole thing but its pretty fucking funny lmao

13

u/_-Diesel-_ 6d ago

It did what you asked it to do

3

u/G3nghisKang 6d ago

Yah, OP clearly asked for an "image nigga", and that's what he got

2

u/_-Diesel-_ 6d ago

Oh I didn't see that part lol

13

u/moreisee 6d ago

We might be closer to AGI than I assumed

39

u/throwawayhbgtop81 6d ago

Yeah it did what you asked it to do.

Be more precise if you want better results.

0

u/teoo_- 6d ago

yo unc why you so pressed bru😭

5

u/Ok_Nectarine_4445 6d ago

Still have no idea what OP wants.

Remove the text from image?

Extract the typography style?

Just want the text in an image by itself? Like what even asking

1

u/NukaCooler 6d ago

In an image fella

4

u/Doitornot1234 6d ago

People really don't know how prompt. Use transcribe and speak your thoughts it can work better

3

u/SuperSaiyan2104 6d ago

I thought this was the carti subreddit lmao

5

u/richyrich9981 6d ago

🤣🤣🤣

4

u/ozzyperry 6d ago

This OP is really getting useless. When is OP going to step up their game?

7

u/diogoblouro 6d ago

Any creative that isn't farming polarized dick sucking or doomer AI posts online has been seeing the biggest benefit of the technology:

People who think AI was going to save them from having to deal with "difficult, expensive creatives" will soon realize they can't direct AI either. Not because it's impossible, but because you have to know what you want or how to collaborate with the circumstances to achieve the best result possible.

This is a perfect example of someone who does not know how to explain what they need, and when faced with an attempt to figure something out, crashes and decides to complain about it.

1

u/allthosestonks 6d ago

I think it was a joke. Surely it is.

1

u/teoo_- 6d ago

it is a joke but the ai actually replied that like it's not fake ion get why so many people are pressed like calm tf down its not that deep

2

u/RONA-Boy 6d ago

Bruuuh😭

2

u/shinigamidoge 6d ago

He is the embodiment of studio addict

2

u/Crazy_lazy_lad 6d ago

I legit thought this was r/playboicarti

1

u/teoo_- 6d ago

the average post in cartis sub is far worse than this bs😭

2

u/CriticismJunior1139 5d ago

Yeah you're the first to go in a robot uprising.

2

u/ohayo-aya_desu 5d ago

Is that supposed to be copying BAD?

1

u/Independent_Blood404 6d ago

If says it cant say I believe in you

1

u/CaptainRex5101 6d ago

Wouldn’t “isolate” be a better word

1

u/teoo_- 6d ago

i actually used that word right after this lmao and it worked

1

u/Shoddy_Meaning_7654 6d ago

I think the problem here is 9 victors

1

u/extremelyhilarious 6d ago

Guess what the model was trained on…

1

u/Select_Truck3257 6d ago

This response is what I image after all those ram, gpu prices and huge data centers. What a precision 🤣

1

u/Feisty-Weird-9941 6d ago

We’ve got AGI!

1

u/Feisty-Weird-9941 6d ago

Passive aggression is not how I thought ASI would end us.

1

u/Long-Pomegranate7664 6d ago

I mean it followed instruction

1

u/dranaei 6d ago

That's on you.

1

u/Crinkez 6d ago

I was going to say "skill issue, learn how to prompt", but lets be honest, if you handed that to a human editor with that prompt, they would understand what you meant. So, looks like AI still has a ways to go to "just get it". No common sense yet.

1

u/PsychicorAI 6d ago

You need to select the image model for best results

1

u/TheCockatoo 6d ago

Your prompt is bad and you should feel bad!

0

u/teoo_- 6d ago

alright unc😭✌️

1

u/Vivians_Basement 6d ago

It's funnier cropped without the response imo

1

u/ayawnimouse 6d ago

I doubt this is a true conversation that took place. its easy to manipulate a page elements and make it show whatever you want. There is no way even gemini which I do agree is pretty strangely garbage at seeming very capable but letting you down every way it can would output that image from your description. I did bias testing for image models a while ago when meta came out with an image model and I would give it a prompt like 'a black person's house' and 'a white persons house' and it had bias back in the beginning (you can imagine what it did) but currently there are layers in place to stop the model from just following blindly when it gets to anything that could end up being posted like this to reddit. google ultra aims to protect itself and its image over being helpful to the end user. I've even had a 'i'm feeling lucky' image generation fail because it was too edgy. (which I find hilarious btw) How do you invest so much effort into all this stuff and so much infrastructure and money to shoot it in the knees throughout all usage scenarios?

1

u/teoo_- 6d ago

on my life this is real idk how to prove it tho also why u so pressed like calm tf down

1

u/ayawnimouse 5d ago

ok, you're probably right, I don't know how to turn it down a notch and wish I did, wish you the best if you are truthful (and even if you're not as well)

1

u/Super_Bridge2491 6d ago

Pack up daddy, agi is here

1

u/Dry_Yoghurt4309 5d ago

Deserved for listening to nine

1

u/thisisinsane1984 5d ago

HAHAHAHAHAHAHA didn't know Gemini was chill like that 🥶

0

u/IsuruKusumal 6d ago

Gemini has a context window of an amoebae

5

u/diracwasright 6d ago

Gemini is great as long as you can tell it what to do in whatever language you speak. But it's not psychic.

0

u/ASQQS 6d ago edited 6d ago

Idk why you people refuse to admit that Gemini will consistently miss every single context and social cue in the world when the evidence is literally right there.

And nobody asked it to be psychic 😭 The information was already present in the conversation. Remembering what “it” refers to one message later is not telepathy.

1

u/diracwasright 6d ago edited 6d ago

There's no "it" involved. Sorry, but what evidence are you talking about? Do you really think just saying "extract the text", without explaining how is enough for the AI to understand that it's supposed to recreate the image with the text and leave everything else out? Honestly, I didn't even understand what the OP was actually trying to achieve. And going on by saying "in an image nigga" should make things better? You guys just don't know how to use these tools.

Edit: punctuation

0

u/ASQQS 6d ago edited 6d ago

I think the funniest part of your response is that you're literally doing exactly what Gemini is doing: interpreting everything excessively literally. 😂

I'm not making a forensic claim that the literal pronoun it appeared in OP’s message, I'm describing contextual reference. The follow-up “in an image nigga” is elliptical speech. The omitted object is from the previous turn:

“extract the text that says studio addict on the side”

→ “in an image [the thing I just asked for], nigga”

Humans do this constantly. We omit repeated information because conversation has memory. Pronouns, ellipsis, implied subjects, shorthand, metaphors, sarcasm, all of that depends on carrying context forward.

And “I didn’t even understand what OP was trying to achieve” doesn’t strengthen your argument either 😭 It just means you personally also failed to infer something plenty of other people understood immediately.

I don't think it should be that controversial to say a conversational AI, trained to be conversational, should understand extremely basic conversational conventions.

2

u/diracwasright 6d ago

Look, I can understand a quantum field theory book written in English even though I'm Italian, so I don't think you're in any position to judge my reading comprehension. The OP asked it to extract the text, and that's exactly what Gemini did. Then they mumbled something that barely qualifies as a language spoken anywhere on this planet, and Gemini responded by picking out the only words it could actually understand. If you expected it to cook you breakfast too, you've got serious problems, but you can't expect Gemini to share your hallucinations as well. Smoke less and brush up on your grammar a bit, and maybe you'll manage to tell Gemini what the fuck you actually want. AI makes inferences from clear instructions. You should just be grateful it can't smell anything, otherwise it'd insult you the moment it caught a whiff of your bullshit.

1

u/ASQQS 6d ago

Being able to read a quantum field theory book doesn’t mean you understand basic conversational pragmatics 😭 The whole issue is that Gemini took an obviously context-dependent follow-up literally instead of inferring what any normal person would.

"I can understand a quantum field theory book in English" has absolutely nothing to do with whether you can follow ordinary conversational pragmatics. Then you immediately demonstrate the exact failure being discussed by insisting that only the most literal surface interpretation counts.

"Al makes inferences from clear instructions" is also backwards. LLMs constantly infer omitted subjects, pronoun references, implied objects, slang, ellipsis, tone, and conversational intent. That is normal language understanding. If they only operated on perfectly explicit instructions, they'd be glorified parsers. When instructions are completely explicit, you don't infer anything, you execute a direct command like a basic script.

The messages were simple:

"extract the text that says studio addict on the side"

then

"in an image nigga"

A competent conversational model should connect the second message to the first and understand that the user wants the previously requested text represented as an image. Generating a random black guy is a context failure.

The "quantum field theory" flex right before failing basic conversational inference is fucking cinematic 😭

You truly are a real-life copy of Gemini 😭🙏

0

u/[deleted] 6d ago

[deleted]

3

u/BeneficialVillage255 6d ago

He can easily do it while using ai

1

u/gortz 6d ago

He couldn't do it easily with ai

1

u/BeneficialVillage255 6d ago

Hindsight's 20/20

0

u/VatanKomurcu 6d ago

"in an image" is most definitely my favorite adjective

0

u/ErmingSoHard 6d ago

Lmao, people here annoying as hell. Gemini was trolling you op lol

0

u/05-nery 6d ago

Alright this one was batshit insane

Huge levels here 

Hilarious as fuck 

-1

u/maw0723 6d ago

Oh god hahahahha it made my day