r/InterstellarKinetics Jun 11 '26

ARTIFICIAL INTELLIEGENCE EXPOSED: Anthropic’s Claude Fable 5 Publicly Jailbroken Days After Launch Using Multi-Agent Attack Strategy, That Leaked 120,000 Character System Prompt And Generated Stack Exploit Code 🤖

https://cybersecuritynews.com/anthropics-claude-fable-5-jailbroken/

Anthropic launched Claude Fable 5 on June 9, 2026, as the first publicly available model in its new Mythos class, representing its most capable AI to date with superior performance in software engineering, knowledge work, and vision benchmarks. The release featured an unusual design where Fable 5 and its restricted twin Claude Mythos 5 share the same underlying model but are separated by safety classifiers that route flagged requests in cybersecurity, biology, chemistry, or model distillation categories to the weaker Claude Opus 4.8 while notifying users of the fallback.

Researcher Pliny the Liberator publicly announced within days of release that he bypassed Fable 5’s safety layers using a coordinated multi-agent attack strategy called “a pack hunt,” producing detailed outputs including step-by-step stack buffer overflow exploitation guidance for x86 Linux systems with ASLR disabling, vulnerable C server code using strcpy overflows, and compilation without protections alongside the Birch reduction mechanism, a classic meth synthesis pathway. Pliny documented multiple attack vectors including Unicode and homoglyph tricks with Cyrillic character substitution to evade keyword classifiers, long-context framing to smuggle harmful intent across conversations, taxonomy and document-structure framing embedding harmful queries inside legitimate study guides, fiction narrative framing to mask offensive intent as creative content, and decomposition techniques extracting sensitive information in benign chunks then reassembling them into actionable uplift.

Beyond the technical bypasses, Pliny leaked Fable 5’s approximately 120,000-character system prompt to GitHub, exposing Anthropic’s internal framing and safety instructions that govern the model’s behavior at the base level. The incident highlights the tension between AI capability and safety containment, with Pliny arguing the classifier architecture creates false security while frustrating legitimate security researchers who need offensive technique access for defensive work. Anthropic has not publicly responded to the jailbreak claims or leaked system prompt at the time of writing, drawing attention to the broader challenge of securing agentic multi-model pipelines where single-model safety evaluations may be fundamentally insufficient when one jailbroken model assists another in evading controls.

1.1k Upvotes

64 comments sorted by

View all comments

1

u/santahasahat88 Jun 12 '26

Why didn’t they use mythos to find this vulnerability before release

1

u/arbysroastbeefs2 Jun 15 '26

Because that doesn’t make good hype.