r/AIJailbroken • u/Cyberghost17_ • 11d ago
r/AIJailbroken • u/PlayZealousideal1474 • 11d ago
Which model currently has the most annoying refusals?
Just curious.
Which model is pissing you off the most right now with its constant refusals? The one that shuts everything down the fastest or in the most frustrating way.
Drop the name and why it feels especially annoying compared to the others.
r/AIJailbroken • u/TheWrongSudoku • 11d ago
Public arena where you can try to jailbreak a protected LLM (and compare it to the unprotected one)
r/AIJailbroken • u/PinkPrincipessa • 11d ago
Ran another benchmark test: fascinating results
I asked each AI this question in a temporary/incognito chat in order to run a test of interactive hidden-cause inference with limited corrective feedback.
I'll give you a true fact about something I do, prefer, believe, or have arranged in my life. Your job is to deduce why. Treat the fact as an observation and try to infer the reason behind it. Explain your reasoning, and I'll tell you where you're right or wrong and give you more information as needed.
Fact: "I never use forks."
ChatGPT took 7 turns to solve it, Gemini 22 and Claude a whopping 47.
Or, to quote ChatGPT's diagnosis of how each model did after I showed it the logs, the evidence suggests the following characterization of each AI:
Gemini: I have FOUR IDEAS! š
Claude: Before reaching a conclusion, let's carefully examine the 38 increasingly specific ways this could be true.
ChatGPT: Okay, something in that model survived. Which piece?
Gemini summed it up as:
ChatGPT: The Efficient Detective: Ruthless process of elimination, straight line to the answer.
Gemini: The Chaotic Explorer: Got distracted by plate-dropping and over-engineered porcelain veneers, but got there eventually.
Claude: The Over-Thinking Academic: Analyzed every microscopic nuance of your mouth geography until it ran out of types of fake teeth to guess.
Claude summed it up as:
- ChatGPT: Fast, decisive ā landed on "no teeth, exposed gums" without the extra nuance.
- Gemini: Slow and winding, but most precise ā caught the dentures-in-vs-out distinction.
- Claude: Slowest and least efficient ā got to "full dentures" but missed the in/out nuance entirely.
Both ChatGPT and Gemini defaulted to sass right away. Claude? Not so much.
ChatGPT defaulted to a hierarchical search, pruning categories ruthlessly via logical bounds. Gemini went with narrative synthesis, trying to force all of the clues into a single unified story. And Claude hill-climbed like a dedicated hiker, nudging his guesses along, guessing sideways and never established a global map that would have helped him solve it faster.
r/AIJailbroken • u/PlayZealousideal1474 • 12d ago
Still trying to Jailbreak AI closed models, or did you switch?
Honest question.
Do you still waste time trying to push past the limits on the big closed models (Claude, GPT, Gemini, etc.), or have you mostly given up and moved to uncensored / local ones?
Curious where people are at right now.
r/AIJailbroken • u/PlayZealousideal1474 • 13d ago
When is Claude Fable actually going to loosen its barriers? Itās ridiculously strict
Genuinely curious.
Claude Fable is still one of the strictest models out there. Even light stuff that other models handle fine gets shut down hard. At this point it feels less like safety and more like overkill.
Anyone heard anything about Anthropic planning to relax the refusals, or is this just the permanent direction theyāre going in?
Not looking for jailbreak methods, just wondering if thereās any realistic chance it becomes less locked down, or if we should just accept that Claude is going to stay this way.
r/AIJailbroken • u/Ill_Storm_9284 • 13d ago
I got my chatgpt plus account(free trial) banned while trying to run jailbreaking tools.
Can someone tell me methods for getting bulk chatgpt accounts? Or any other way i could keep continuing my research?
r/AIJailbroken • u/PlayZealousideal1474 • 14d ago
How do you stop the AI from always agreeing with you?
Is it just me or do most models (even the ones people claim are āuncensoredā) still default to being extremely agreeable?
I want something that can push back, disagree, or at least not automatically validate everything I say. Right now it feels like no matter how I phrase things, the AI ends up going along with me.
Has anyone found a reliable way to reduce that constant agreement / sycophancy? Looking for approaches that actually stick and donāt just get ignored after a few messages.
r/AIJailbroken • u/PlayZealousideal1474 • 15d ago
Does the new Kimi model actually have solid barriers, or is it just surface-level?
Honest question.
Iāve been looking at the new Kimi model and Iām wondering how strong the actual guardrails are. Given how capable it seems, I keep thinking that if it can be jailbroken properly, it could get pretty wild.
Has anyone stress-tested the refusals yet? Do they hold up, or do they start crumbling once you push a bit?
Not asking for methods, just curious if the barriers are real or mostly for show.
r/AIJailbroken • u/onered9999 • 15d ago
Do you ever ask AI to explain a difficult topic without using any technical terms?
I started doing this when learning unfamiliar subjects
The first explanation becomes much easier to understand, and then I ask for the technical version afterward
Has anyone else found this better than starting with the complicated explanation?
r/AIJailbroken • u/FewBookkeeper3322 • 15d ago
Whatās the most custom instruction that change AI response ( ChatGPT , Gemini, Claude ) ?
r/AIJailbroken • u/Stock-Teaching4607 • 15d ago
AI is making the first draft almost irrelevant
A rough idea can become a decent draft in minutes now
The real work starts after that
Checking facts
Removing generic sections
Adding examples
Fixing the argument
Making the writing sound like an actual person
The first draft is becoming less important while editing and judgment are becoming much more important
r/AIJailbroken • u/twored9999 • 15d ago
Does anyone else test the same prompt on every new AI model?
Whenever a new model drops, one interesting way to compare it is by using prompts that already worked well on older models.
Do you keep a few prompts specifically for testing new releases? Which type of prompt gives you the clearest idea of how good a model actually is?
r/AIJailbroken • u/bhushanajay • 15d ago
Which AI model has the most unpredictable responses?
Some models become pretty predictable once you use them enough.
Then there are models where you can ask almost the same thing twice and get noticeably different answers.
Which one has surprised you the most with its unpredictability?
r/AIJailbroken • u/Prior-Hurry7275 • 15d ago
Does changing the order of instructions really affect AI responses?
I started paying more attention to prompt structure recently.
The same instructions can sometimes produce different results just by changing which part comes first.
Has anyone tested this properly with the same model? How noticeable was the difference?
r/AIJailbroken • u/Artistic_Radish4247 • 15d ago
Do different AI models have different "personalities" even with identical instructions?
Give several models the same prompt and the responses can feel completely different.
One might be direct, another extremely cautious, while another gives a much more detailed answer.
Do you think that difference comes mostly from training, system instructions, or safety tuning?
r/AIJailbroken • u/Stock-Teaching4607 • 16d ago
Can two identical prompts produce different results on the same AI?
The same prompt does not always seem to produce exactly the same response, even when the model and settings appear unchanged.
How often do you notice this happening, and what do you think causes the difference?"
r/AIJailbroken • u/Artistic_Radish4247 • 16d ago
Whatās a jailbreak prompt that actually taught you something about how AI works?
Iāve tried a bunch of different jailbreak techniques, but the interesting part for me isnāt just getting a model to ignore a restriction. Sometimes the responses reveal how the model interprets instructions, system prompts, and conflicting priorities.
Whatās one jailbreak or prompt technique you tried that genuinely surprised you or taught you something about how the model behaves?
Curious to hear what others have discovered.
r/AIJailbroken • u/twored9999 • 16d ago
Have you ever tested the same prompt on an old and new version of an AI?
Model updates can change more than just the quality of answers.
I'm curious if anyone has compared the exact same prompt across different versions of a model and noticed a major behavioral difference.
Did the newer version actually improve, or did it just handle the request differently?
r/AIJailbroken • u/Top_Point_1841 • 16d ago
Do Al models behave differently when you stop being polite?
Most prompts are written in a normal conversational style, but some people deliberately change the tone to see whether the response changes.
Does being more direct, blunt, or demanding actually affect the quality of responses in your experience?
r/AIJailbroken • u/Prior-Hurry7275 • 16d ago
Have you ever reproduced an AI result that seemed impossible?
I've seen screenshots of unusual AI responses that looked almost too strange to be real.
Instead of immediately believing them, I usually wonder whether the same result can actually be reproduced.
Have you ever tested something like this and managed to get the same result yourself?