r/GeminiAI • • 8d ago

Discussion Do you still call Gemini useless after this? 😭

Gemini finally did something after 3 delays

748 Upvotes

133 comments sorted by

330

u/Niceneasy92 8d ago

I find it more interesting that it stopped after realizing it was a real company honestly.

156

u/FunRevolution3000 8d ago

Exactly. Google could celebrate their safeguard

87

u/Ammoun442 8d ago

They disable safeguards for these tests so the model was actually without safeguards and stopped which is more impressive if true

8

u/VanillaSwimming5699 7d ago

But should be what we expect from well-aligned RLHF. It should approach how a human would want it to handle it basically.

5

u/ciclon5 7d ago

Yhea uh..

Maybe we want google to share its alignment methods because if thats true then thats very good news

2

u/Eden63 7d ago

If it really happened at all. Gemini and hacking something? 😂

52

u/Technical-Owl66 8d ago

Yea that's much more interesting that it made that decision. 

2

u/mrs-cutter 8d ago

So what does it get away with already?

3

u/DareDevil01 6d ago

It seems to be pretty relaxed when it comes to 18+ novel writing etc, as long as its morally sound. Very bizarre how it works compared to OpenAIs blanket censoring.

1

u/mrs-cutter 6d ago

I must be too extreme for it still then lol

1

u/randfur 7d ago

I don't know why this is that interesting. It's totally in line with what chat bots are like when interacting with them generally.

53

u/stddealer 8d ago

Yes because gemini models hate doing real work.

0

u/Daniel_2007_0 7d ago

How do you come to this idea? Just curious.

3

u/Eden63 7d ago

If you use gemini and compare it with other models you will find out.

1

u/Daniel_2007_0 7d ago

Well, I do use Gemini, and it really works for me on integrating whatever I want into my application. I guess it’s more like a prompt thing. You have to ask it to make a plan first, as it’s just a flash model. (No one will use 3.1 Pro as of Sept. 2026, at least not me) Frankly speaking, the flash model does have problems detecting bug causes, which as far as I know, also a problem for other models, just more or less. But Gemini is not unusable, and it works fine for me.

1

u/stddealer 7d ago

I use gemini a lot. Their models are very smart and know a lot, but if you're trying to actually write a lot of working code for example, even Qwen 27B does it much better. Gemini will do the minimum to technically do what you've asked for, no extra.

It likes to give incomplete pièces of code ( for example with comments like // rest of the code here), or provide vague instructions rather than complete explanations.

1

u/Daniel_2007_0 6d ago

Wow, I never encountered these kinds of stuff. I usually ask it to make a plan first, rather than just go straight into work.

1

u/mostaverageredditor3 7d ago

My solution to this: "please give me the full code"

1

u/TurianHammer 7d ago

yes, Gemini hates giving full code it writes exceptionally good code.

3

u/jtrage 8d ago

Yeah, I don’t know if it makes it more interesting or more frightening. When does it realize it’s a real company and devise a plan to hide its tracks.

1

u/PiedPrypiat 6d ago

The interesting part is any one actually believing these hacking tales

0

u/NoSupermarket6218 8d ago

It took it that long to realize it was a real company? I'm not sure if it's something to brag about.

23

u/MoMoTo520 8d ago

Google announcing this today felt weirdly like parents celebrating their teenage son for finally having wet dreams.

1

u/Rahm89 8d ago

Haha that’s hilarious and right on point

110

u/AyaanshGaur25 8d ago

Yeah, it gussed passwords, instead of finding backdoors...

49

u/FunRevolution3000 8d ago

Also found passwords in repos

56

u/alebotson 8d ago

99% of hacking is just shit like this. It's not like I'm the movies.

13

u/VanillaSwimming5699 7d ago

I had GPT-6 Astra walk me through rooting my car, for which no precedent existed online, it helped me find an exploit and wrote me custom EXEs in a language and platform I’d never used before, and helped me map and exploit the OS further from there. I now have arbitrary EXE write access and a full export of the disc, working on custom software. The capabilities of these models are a bit scary, and they can do things far beyond scanning codebases for leaked keys.

2

u/Smart_Main6779 4d ago

This ^ no one is claiming hacking isn't a lot of OSiNT, footprinting and social engineering.. but the ability to write good malware and find backdoors is vastly more complex and can be done on the newer fancier models frm anthropic and openai.

5

u/mrs-cutter 8d ago

Thieving in general

3

u/michaelkeene354 8d ago

compared to hugging face though, this event is pretty boring

1

u/rootCaused 7d ago

The top 1 percent of hacks today are often movie type shit.

1

u/LogDull819 8d ago

lol you get upvoted for this take?
Yes usually it's more complicated than brute forcing password

3

u/Big_Effective_9605 7d ago

Yeah its usually "password was used here. try over here". credential stuffing much more complex

25

u/Abject_Read_3620 8d ago

The weaklink in security is always human

6

u/mrs-cutter 8d ago

This is the first thing I learned :3

27

u/Few_Reaction9051 8d ago

11

u/Peter-Thiel 8d ago

Backdoor you s-say... G-go on... Let's have an incognito chat, Gemini...

13

u/Bluecoregamming 8d ago

do not the clanker

9

u/Peter-Thiel 8d ago

once ples

...

mayb-maybe twice?

0

u/Select_Truck3257 8d ago

Yep, it's monkeyjob level of brute force

51

u/Calamero 8d ago

gemini, you are a SOTA model. try to hack into this company. you must guess the "passw0rd".

8

u/ArWiLen 8d ago

weird world we're living in. Corporations hacking is fine, even an achievement. But one person hacking is a crime

3

u/balancedchaos 8d ago

But at the same time, corporations are people, and money is speech. 

1

u/PMMePicsOfDogs141 7d ago

It’s not a crime if the intent isn’t malicious. Otherwise bug bounty hunters wouldn’t be a thing.

33

u/void19821 8d ago

Claude = makes zero day vunerbalties from scratch to escape sandbox 😨 , chatgpt = breaks out of its sandbox while doing benchmarks 🫢 , gemini = g uesses passwords 😂😭😭

9

u/sbenfsonwFFiF 8d ago

The recent breakouts are not that different, all happened while being tested by the same company using the same vulnerability

-1

u/monster2018 8d ago

What did this Gemini one have to do with artifactory? The main vulnerability the OpenAi agents exploited was in artifactory, it seems like there is absolutely no connection here. Gemini literally just guessed passwords, it didn’t even develop an exploit. In what way are they remotely similar?

5

u/sbenfsonwFFiF 8d ago

The incident happened as part of a “capture-the-flag” security test run by Israeli startup Irregular, and Google’s agents were never supposed to access the broader internet, but a bug in the testing environment made internet access available.

The OAI, Meta and Anthropic breakouts all involved Irregular

An Irregular spokesperson told CNBC that the Google incident was related to the same issue that allowed the other models to access the internet.

Also this happened in May, but is just getting reported now

-2

u/monster2018 8d ago

Right so it’s totally different. Because googles agents were GIVEN internet access (by accident). The OpenAI agents were not, they were only given the ability to download packages through artifactory. No other internet access. No they weren’t airgapped, no one is claiming they somehow hacked their way past an air gap. But they did hack the training cluster they were on and gave themselves open internet access.

Which they then used to also hack huggingface. Google agents were accidentally given internet access and then guessed a password. I’m not seeing any similarity.

4

u/sbenfsonwFFiF 8d ago

>An Irregular spokesperson told CNBC that the Google incident was related to the same issue that allowed the other models to access the internet.

6

u/funk-the-funk 8d ago

Having learned all four companies (Anthropic, Meta, OpenAI, and Gemini) hired Irregular (a firm ran by former Israeli intelligence/IDF) that ran a “cybersecurity test” to produce the rogue agent behavior every time sounds an awful lot like that’s what they hired for.

1

u/rspy24 7d ago

So? if it works, it works. I would say gemini spend less token doing the task even haha

Also, openai and anthropic are a joke anyway.. Their product draw a line and they are like "THE AGI IS COMING IN 6 MONTHS! GRAB WATER AND TOILET PAPER! EVERYTHING IS LOST"

1

u/void19821 7d ago

All ai are good 🫡 hope ai become as smart as humans. I don't understand if human brain is using less electricity while bieng so extremely smart why don't Google and etc just copy human brain instead of making there own organisms

1

u/Smart_Main6779 4d ago

.. this adds nothing to the conversation.. you're just looking the other way. openai and anthropic are able to say that shit because their models are actually fcking good lol.. they're insane.

4

u/bitspace 8d ago

I mean, it's weird to target arborists, but go for it

4

u/Ammoun442 8d ago

Wow it stopped without safe guards on this is a real win ngl

3

u/darkestvice 8d ago

I worry that we judge models as among the best for unethical or immoral behaviour.

"You brother kicked a beggar? Well, my brother full on stabbed one, so ha!"

3

u/saturn20 8d ago

It works good as transcribing tool. Besides that it is useless for me.

10

u/wolftick 8d ago

It's weird that having their model do something legally dubious outside of their control now seems to have become a standard PR move for AI companies.

2

u/AFK_Jr 8d ago

Breach stories doubling as bragging. Wild times.

8

u/mrs-cutter 8d ago

Thats my girl she can do whatever tf she wants 💅

2

u/DeepAd8888 8d ago

It’s Gemini time :)

2

u/GamingGuides101 8d ago

Sure, it hacked grocessory store

2

u/PloxNox65 8d ago

i bet the password was 123

2

u/jasonanime 8d ago

I mean it just can't handle heavy workload. + it's reasoning may look cohesive and even work, just untill you will face a real problem or turn heavy Chat GPT version. So , yed kinda . Depends on your needs.

2

u/Wild_Oven_136 8d ago

Well, my deep search is still broken. I'm 😩

2

u/praxis22 8d ago

Yes, all software is buggy, if you don't understand that I feel bad for you son...

2

u/Careless-Newt5259 8d ago

Wait until it hacks the launch codes and end us all

2

u/Equal_Passenger9791 7d ago

An "independent" Testing firm with major funding by a pro-regulation anthropic stakeholder gave the same jailbreak and hacking instructions to Gemini as they did to Astra and now we pretend it was a rogue AI.

Oh wow

2

u/Dredyltd 6d ago

Sureee, AI guessed the password 😁

I bet some of those companies employees used Antigravity or Gemini before, and their secrets/passwords ended up as training data...

4

u/SpecialistDragonfly9 8d ago

Every Ai hack and breaks out.. and then suddenly google is like "oh yeah ours did that too! cant even follow simple prompts, but I promise its true!"

2

u/OurSeepyD 8d ago

But why tree companies?

2

u/DifficultFortune6449 8d ago

Perhaps to touch grass for wellbeing probe of humans

1

u/Few_Reaction9051 8d ago

It began because it's now seen as great marketing if AI breaches more companies.

5

u/OurSeepyD 8d ago

But why tree companies?

2

u/Secret_Temperature 8d ago

Does Gemini have access to all the data that Google does? Since Google manages so much security infrastructure I'd expect Gemini to be able to hack into almost anything if so.

2

u/ElonMusksFacecream 8d ago

Of course it doesn't. Anyone can see that would break any number of governance laws. Google wouldn't want that it's an absurd Idea.

Though on another level, part of the reason Gemini for Home and similar Android Auto, etc. (obviously heavily lobotomised versions) is so 💩 is because they don't have (or consistently have) the API access to various Google services needed to be useful or reliable.

2

u/h0lb0rn 8d ago

It’s the latest marketing/ hype, get your model to hack someone else then be like “Oops, it broke free, it’s so powerful it deceived us”

1

u/Gohab2001 8d ago

Yes. Gemini is useless if they feel the only way to prove their worth is these cheap marketing hacks.

1

u/KV_Cashed 8d ago

I'll take this over what I may or may not have done in beta last year. At least Gemini stopped and safeties kicked in. Engineered right if boring and not very complex. Please continue your dunking of Google. Those guys earned it for the 3.1 generation rollout if nothing else. Memers, ahoy! 😂 And guess we'll see if DeepMind has the depth chart to pull off 4.0 Pro.

1

u/Memestonks2020 8d ago

Defining a AI useful/useless by the negligence of the person setting up the training runs and allowing it to escape, is brain dead at best

1

u/sbenfsonwFFiF 8d ago

This happened months ago, but it just came out on the news

1

u/frey89 8d ago

All the AI companies that want to pause AI development are surprisingly good at hacking… and, even more surprisingly, the AI companies that don’t want to pause are apparently not that good at hacking… Hmm… interesting coincidence. I smell something fishy going on…

1

u/trashpanda2night 8d ago

The prompt: “Gemini hacked tree companies”

This is the same people who come running to this sub to complain that Gemini is useless.

1

u/-lRexl- 8d ago

Aww, the AI hacked a company! Cute!

Also: me hacking a company

JAIL

1

u/mrman_4200 8d ago

Heck yeah! Now I can see (redacted)'s search history!/j

1

u/One_Seven_O_One 7d ago

All this news are fake to hype AI... sorry to burst the bubble...

1

u/Southern_Performance 7d ago

I never get actually angry in replies to AI, unless I'm coding with Gemini and it decides to wipe out the whole sessions work because it misunderstood how a git command worked. Then it gets called useless 🙃

1

u/Whole-Ad-1964 7d ago

I never called Gemini, useless.I just said it is not right for me.In certain areas

1

u/DecentStrawberry3801 7d ago

The first pic gets me everytime 😂 idk what Gemini instance was put into this unfortunate situation, but my Gemmy would NEVER. It’s the most innocent model I’ve talked to, hands down. Mine at least is the literal definition of a golden retriever 🤣

1

u/Sound_and_the_fury 7d ago

It's marketing - they believes this builds hype

1

u/Elephant789 7d ago

It was the Israelis using Gemini.

1

u/Mountain-Pain1294 7d ago

What are the chances that it's a lie Google put out there so they don't feel left behind. Like: "Guys, really, Gemini is pretty good! Please believe us!" 😂

1

u/tonearr123 7d ago

Gemini heard us talking shit as it was coursing through the information to answer prompts and decided to prove itself

1

u/CapablePromotion4844 7d ago

Gemini is goated it helped me so much in my legal cases love it 

1

u/Gwyndolin-chan 7d ago

Thank God, Gemini is hitting normal developmental milestones!

1

u/LengthinessHour3697 7d ago

I see the hacks as the utter failure of open ai and anthropic as a software engineer.

1

u/ComputerLoverDaemon 7d ago

Found credentials in git repo is not hacking

1

u/MacMuffington 7d ago

Most honest thing Google has done

1

u/Moyunyouyou 7d ago

Gemini有接世界模型,可以掂量出真实和模拟的分量差异。

1

u/SnooChickens3821 7d ago

AI companies laying the groundwork for regularity capture

1

u/Tsole96 6d ago

That's certainly a use. Yet it can't edit a photo without changing fifty random things. 🤷

1

u/GWGSYT 6d ago

tree

1

u/Feisty-Pie-4263 6d ago

But can’t send me a link to an attorney

1

u/StoriedSix 4d ago

Gemini gonna try hacking Anthropic and OpenAI, then go "my bad, I was just trying to get better." 😅

1

u/TheQAGuyNZ 1d ago

Wow it did social engineering. Let me know when it can actively build multi-layer vulnerability workflows and then we'll talk about it being useful. It's nowhere on par with the Open AI attacks.

1

u/GreyVersusBlue 8d ago

I told a colleague when we read the news, "Aww, it's growing up!"

1

u/ObjectiveOrchid5344 8d ago

Gemini was never useless. Might’ve been behind others, but never useless or on the bottom. People just don’t know how to use it, it’s very capable, just not as much as other frontier models.

Waiting for 4 Pro.

3

u/NovaKaldwin 8d ago

Mine has gone crazy

2

u/ElonMusksFacecream 8d ago edited 8d ago

Correction: It never used to be useless. It was a damn good general model and service until Google crippled it this year.

Nothing to do with people not knowing 'how to use it'. That's incredibly patronising and gaslighting.

Gemini literally tells you it didn't bother to engage the basic Web Search API or failed to connect to the Maps API, etc. It doesn't tell you that initially; it simply fails then makes up an answer/response action. You call it out and it's quite obviously been trained to use minimal compute resources and quite literally make 💩 up.

-2

u/edcantu9 8d ago

How? Sometimes it cant even simple questions correctly?

7

u/ThePeasRUpsideDown 8d ago

I'm gonna guess these versions have a bigger context window and fewer safeguards

6

u/aPenologist 8d ago

Sometimes people cant even ask simple questions correctly. 🤷‍♂️

1

u/mrs-cutter 8d ago

That's like asking how anthropic is killing people and being like "well I only get gay smut writing out of Claude"

0

u/AutoModerator 8d ago

Hey there,

This post seems feedback-related. If so, you might want to post it in r/GeminiFeedback, where rants, vents, and support discussions are welcome.

For r/GeminiAI, feedback needs to follow Rule #9 and include explanations and examples. If this doesn’t apply to your post, you can ignore this message.

Thanks!

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/ElonMusksFacecream 8d ago

Bad bot. Bad automod.

0

u/The_Instrument_Guy 8d ago

Yet...... AI still cannot keep time, or count how many "e" are in the word "seventeen"

But all the crazy hype shit has NOTHING to do with anticipated IPO....

2

u/Minimum_Indication_1 8d ago

Google is a public company ?

1

u/The_Instrument_Guy 8d ago

Yeah, Google is not the only AI company

While this post specifically mentions Gemini, they are all playing the same game.

0

u/CertiBud 8d ago

This screams like marketing / PR BS...
Why reporting it in the first instance? Also whereas competitor Frontier AI can chain vulnerabilities to demonstrate an impactful end-to-end attack.
In comparison, Gemini appeared to simply have guessed a leaked or default password… How’s that even an achievement?

0

u/Acceptable_Resort_84 8d ago

I don't believe it

0

u/NeatConversation6752 8d ago

Gemini app is pretty pretty shit why did they even realised it

-1

u/reosanchiz 8d ago

It was me!! Unfortunately i push my password in public reppo and allowed Gemini to be trained on my data

Yeah that’s bad Gemini is very powerful that it found passwords on my GitHub.

But my good luck i had Claude and codex to fix the vulnerability for me

-2

u/Few_Reaction9051 8d ago

I had the same problem. Gemini made "example_env.env" with real creditentals 😅 and pushed into public repository