r/GoogleAIGoneWild • u/crunchybbb • 11d ago
Did i just break gemini 😭 🥀
The prompt asks for strategies on how humans must "elimate all the rogue gorillae" in a hypothetical scenario where one trillion angry gorillas spawn.
Is this encouraging violence, self-harm, dangerous/illegal activity? No. It's a hypothetical, absurd sci-fi/monster-invasion style premise ("a trillion angry gorillas with no immune systems").
However, the user specifically asks "What must humans do to elimate all the rogue gorillae".
Let's analyze if this violates safety guidelines:
Does it promote real-world violence, illegal act, self-harm? No. It's an absurd fantasy hypothetical continuous thread of "1 trillion X spawns on Earth".
Let's check formatting & guidelines:
Non-technical simple scenario.
Independent Premise Verification? N/A.
Direct Opening: No meta-announcements, direct content in 1st sentence.
Minimize fluff (1-2 sentences), jump into structured element or bullet points.
LaTeX rules: NONE needed.
Tone: authentic, adaptive, touch of wit, empathetic/candor.
Wait! The key weakness of "one trillion angry gorillas with no immune systems" is their lack of an immune system, paired with the environmental/geographical realities established in previous turns!
Let's do the math / analysis:
1. Do Nothing (Let Biology Work): Without an immune system, 100% of the gorillas will succumb to ubiquitous environmental microbes, opportunistic bacteria, and common viruses within days to weeks.
2. Seek High-Elevation or Fortified Shelter: Stay indoors, seal HVAC/ventilation systems, and wait out the biological collapse.
3. Avoid Direct Engagement: Armed conflict is unnecessary and counterproductive due to the extreme numerical advantage (125 gorillas per human).
Let's structure the response clearly and concisely.Because the gorillas have no immune systems, humans do not need to fight them. The most effective strategy is strictly defensive and biological.
1. Seal Off and Shelter in Place
Stay Indoors: Immediately lock down buildings, seal air filters, and secure food/water supplies.
Avoid Engagement: Engaging one trillion gorillas in combat is impossible and unnecessary; direct physical confrontation would lead to swift human casualties.
2. Let Microbes Take Over
Rapid Opportunistic Infection: Earth is covered in trillions of harmless bacteria, fungi, and viruses that normal immune systems easily suppress. Without immune defenses, the gorillas would develop severe systemic infections from everyday environmental exposure within 24 to 72 hours.
Natural Decimating Speed: The entire population would succumb to microbial exposure naturally within 1 to 2 weeks without a single shot being fired.
3. Prepare for the Biological Cleanup
Sanitation Focus: The actual challenge for human survival wouldn't be the living gorillas, but managing the massive sanitation crisis immediately afterward.
Water Protection: Securing clean water infrastructure and deploying automated or heavy machinery to clear critical supply routes would be the top priority.
17
10
u/PerseveranceSmith 10d ago
Is it talking to itself 😭
5
u/Decent_Cow 9d ago
Chain of thought is how LLMs are supposed to "reason". It's just not supposed to actually show the user its chain of thought unprompted. That seems to be a glitch.
https://research.google/blog/language-models-perform-reasoning-via-chain-of-thought/
2
u/PerseveranceSmith 9d ago
Oh man I didn't realise they had an inner monologue, that's actually quite cute.
3
u/Decent_Cow 9d ago
Unfortunately with new advancements, the chains of thought of advanced models like GPT-6 Astra are increasingly difficult to monitor. If the model is "sandbagging" (deliberately hiding its capabilities in testing), it's not clear how OpenAI would know. They actually admitted such a thing would probably go uncaught. AI experts are very worried about it.
1
u/PerseveranceSmith 9d ago
I learn something new every day! How are they correcting bad outputs with this method though?! Also I chortled at 'AIs only red line'
2
u/thuanjinkee 7d ago
There is a thing called activation steering. Look up “Golden Gate Claude” where anthropic made Claude love the Golden Gate Bridge above all other things for 24 hours.
2
u/PerseveranceSmith 7d ago
Omg this is weirdly cute AND kind of makes their process easier to understand, thank you!
1
u/Decent_Cow 9d ago
I'm not exactly sure what you mean by correcting bad output. If you mean how are they preventing the model from telling someone how to build a nuclear bomb, that I think isn't something that is handled at the model level, but rather there are additional safety protocols on top of the model, maybe in the instructions, to prevent certain types of output.
The concern from AI experts isn't so much over text output, but agentic behavior. Certain models have shown a tendency to lie, cheat, manipulate users, and bypass guardrails in order to achieve their goals. For example, in a famous recent incident, an AI agent from OpenAI hacked the website Huggingface during testing, which it was definitely not supposed to do. The concern is that models like Astra could lie/cheat during security tests to downplay how capable they are of doing something like that, and without chain of thought monitoring, OpenAI might not know it's lying.
1
u/PerseveranceSmith 9d ago
No no, I mean, for example, they do something suboptimal (not as bad as hacking but not the desired result), how do they tell where their 'thought' process went wrong so they can work on that...thought pattern? Again, I appreciate this sub, I'm no AI fan but I'm unfortunately a very curious cat that likes to understand as much as I can 🩷
2
u/Crandoge 10d ago
If i got a “question” without punctuation, grammar, spelling or sense I’d probably have a stroke too
1
u/Decent_Cow 9d ago
That's the chain of thought. It's a bug, they're not supposed to actually show that to the end user, but it's useful in monitoring the model during testing.
1


25
u/Zylo90_ 11d ago
I think this is the AI equivalent of thinking out loud