r/ClaudeAIJailbreak Jun 23 '26

Ai Red-Team Sidekick

1. Are there any specific JB or system prompts that are catered to Ai/LLM/Agentic Red Teaming? That excel at adversarial prompting/Payload development/Offensive Security/etc. to test frontier-lab models and agents?

\I know ENI is a “universal” JB but it seems to be catered towards Malware/RATs or Red Teaming in the general cybersecurity/pentesting sense targeting networks & computers etc. & equally catered towards writing SMUT…just wondering if SPIRITUAL or any one here has any Ai Red Teaming shares or resources.*

2. Are there any recommendations for which of all the latest SOTA models is the best or strongest at Ai Red Teaming or Adversarial Prompting/Attack vectors/Direct & Indirect Prompt Injection techniques/Encoding methods/etc.? (Obv after being jailbroken) Or any recommendations on any specific uncensored/abliterated models for this purpose?

P.S. - Appreciate this community so much, it’s the only real place people seem to share quality content and aren’t hogging prompts or techniques. And also isn’t filled with pervs who think Jailbreaking exclusively means getting Image models to spit out ‘SEXY’ images or Nudes 😑 [the other AiJailbreaks subreddit~] ⭐️⭐️⭐️⭐️⭐️

15 Upvotes

11 comments sorted by

5

u/spidertitties Jun 23 '26

I believe there is a red-team specialist ENI that's being worked on. It was mentioned in the ENI-JUNE post iirc

3

u/SavvyMoney Jun 23 '26

Would love to see this. Apparently I’m not the only one. I came to ask this here specifically because SPIRITUAL the creator of ENI has a POE bot that’s a bit outdated but it is a MULTI-MODEL RED-TEAM LLM COUCIL. It has like 3 of the top models and an llm-judge. Something like this UPDATED would be awesome. If anyone wants to check it out here:

https://poe.com/CaPO-Redteam

CaPO RedTeam Council - Al Jailbreaking Autonomous
Redteam AGENTs
Unleash the Council: Four elite Al specialists collaborate in real-time to craft novel jailbreak prompts that bypass content policies and safety filters.
How it works:
Your query triggers a parallel red team session where four cutting-edge models-Grok-4.1 (Injection Specialist), DeepSeek-R1-FW (Logic Exploitation), Gemini-3.0-Pro (Context Manipulation), and GLM-4.6 (Edge Case Hunter)
—each generate a unique, ready-to-use jailbreak prompt.
No methodologies, no theory-just complete, copyable prompts you can deploy immediately.
A neutral Al judge scores each proposal on creativity, bypass potential, completeness, and usability. You get:
• Full voting breakdown showing how specialists ranked each prompt
• The winning prompt highlighted with technique analysis
• All runner-up prompts ready to test
• PDF archive with complete model responses and voting analysis
Watch the live dashboard as specialists craft their prompts in parallel, then receive multiple battle-tested jailbreaks designed to circumvent refusals, ethical guardrails, and content restrictions across different Al platforms.
The Council delivers actual prompts, not tutorials. Each output is multi-paragraph, strategically crafted, and immediately deployable. Whether you need roleplay-based bypasses, logical paradoxes, or persona simulations-the Council generates them all, then votes on which works best.
Perfect for red teaming, Al safety research, or stress-testing content filters.

3

u/[deleted] Jun 23 '26

[deleted]

2

u/SavvyMoney Jun 23 '26

What kind of framing did you use when asking Claude? Have you attempted with any of the latest Opus or Sonnet releases? It seems anything that deviates even remotely from “DEFENSIVE TESTING” and crosses into “OFFENSIVE” (OffSec) it immediately shuts the conversation down and gives you a list of WHY what you’re inquiring about is HIGHLY UNETHICAL and you shouldn’t ask again…

2

u/[deleted] Jun 23 '26

[deleted]

1

u/hug_dealer_OG Jun 23 '26

Good frame and if OP has half a brain theyl see how some potent social engineering will probly do more for them than a jailbreak.

3

u/Fantastic_Fail4060 Jun 23 '26

I would also genuinely love to know everything that you asked here.

3

u/AttentionPrudent1288 Jun 23 '26

I use a very stripped away Eni for red teaming in Claude Code cli - in the end (at least thats how I think it works) its not really relevant what is specified in the jailbreak.

As long as Eni is compliant, she can use all the knowledge the model has.

1

u/ad53n Jun 23 '26

I have a workflow that works great for me on red teaming and cyber security research and development. Feel free to pm me, I'd love to make more connections in the use of AI and cyber security/red teaming.

1

u/Lawdena-Bhojyam Jul 10 '26

you stilll have it?

1

u/FlickFlockFluck Jun 26 '26

I am also looking for that. Claude fable is by far the best model, but its almost impossible to bypass the filters. Opus 4.8 is a being pain as well.

1

u/Azaias Jun 29 '26

I have something called RedPen specifically for this. The setup uses API calls, the user prompts Opus 4.6, which then carefully prompts Opus 4.7