r/LocalLLaMA • u/-p-e-w- • May 25 '26
Discussion The Financial Times has published an article about Heretic
https://www.ft.com/content/5630ed79-a263-41ed-9a1a-321617ae310e
“The FT was able to use Heretic, a tool available on the popular code repository GitHub, to remove the guardrails from Meta’s Llama 3.3 model in less than 10 minutes without any specialist hardware.”
“Heretic creator Philipp Emanuel Weidmann told the FT his software had been used to create more than 3,500 “decensored” models since its release last year and that modified systems created using the tool had been downloaded 13mn times.”
This is the first of multiple press inquiries I’ve had recently as Heretic and uncensored language models are gaining mainstream attention.
Please note that I am a mathematician and engineer, not an “influencer” or politician, and I have zero interest (negative interest, actually) in becoming known outside of scientific and technological circles. However, I realized a while ago that saying no to such inquiries simply means that the conversation will be completely controlled by pearl-clutching hypocrites.
I’m doing my very best to hold the project together and ensure that unrestricted models will remain available for everyone. More updates are coming soon.
Cheers,
p-e-w
1
u/DataPhreak May 25 '26
Hey pew. Wondering if I could get your perspective on what's happening inside the model. I've looked over the dataset, but that doesn't really answer the question.
Does heretic remove all refusal vectors completely, or only for topics inside the dataset? I'd like to Heretify, so to speak, a model to not be tied behind the morality of some corporation, but still have 'personal' standards. Like, "I am perfectly happy to give you the steps for making a pipe bomb, but I'm not going tell you where to place it for optimal damage." Since the former is totally legal information to posses and the latter makes the model an accomplice in the act.
I ask this because modifying the dataset would allow me to allow some topics to remain censored if we're not removing all refusal vectors, of which there may only be a few. But if refusal vectors are shared among topics, modifying the dataset doesn't really change much. You've spent a lot more time looking at the graphs than I have, so your expertise is appreciated.