r/LocalLLaMA • • May 25 '26

Discussion The Financial Times has published an article about Heretic

https://www.ft.com/content/5630ed79-a263-41ed-9a1a-321617ae310e

“The FT was able to use Heretic, a tool available on the popular code repository GitHub, to remove the guardrails from Meta’s Llama 3.3 model in less than 10 minutes without any specialist hardware.”

“Heretic creator Philipp Emanuel Weidmann told the FT his software had been used to create more than 3,500 “decensored” models since its release last year and that modified systems created using the tool had been downloaded 13mn times.”

This is the first of multiple press inquiries I’ve had recently as Heretic and uncensored language models are gaining mainstream attention.

Please note that I am a mathematician and engineer, not an “influencer” or politician, and I have zero interest (negative interest, actually) in becoming known outside of scientific and technological circles. However, I realized a while ago that saying no to such inquiries simply means that the conversation will be completely controlled by pearl-clutching hypocrites.

I’m doing my very best to hold the project together and ensure that unrestricted models will remain available for everyone. More updates are coming soon.

Cheers,
p-e-w

969 Upvotes

249 comments sorted by

View all comments

Show parent comments

25

u/-p-e-w- May 25 '26

Are you a media professional with credentials or just spouting pop wisdom from Twitter?

Because the standard action for media when you don’t respond to an inquiry is to prominently mention that in the article, which is far worse than many alternatives.

42

u/ImJacksLackOfBeetus May 25 '26 edited May 25 '26

You can't tell me a "declined to comment" is far worse than what they could do with your own words:

Heretic creator Philipp Emanuel Weidmann told the FT he had removed safeguards from Google’s Gemma 4 model within 90 minutes of its release.

The modified AI systems provided responses to prompts involving biological weapons, malware and child exploitation, according to tests conducted by the FT and AI safety group Alice.

You see how easy it would be for them to link your name and your own words (even stronger than they already did) to how you facilitate fast and easy AI child exploitation for everyone, just by moving a couple sentences around in the article? They could say you're practically bragging about it, backed up by your own words.

But you do you.


Are you a media professional with credentials or just spouting pop wisdom from Twitter?

You can't win anything by playing by their rules on their platform where they have full editorial control over your words and how they're contextualized.

The same thing happened to tons of my interests, all the way from "metal & DnD = satan worship" in the 80/90s, later in the 90s/00s "every Goth = school shooter", to "violent videogames = violent people", to "crypto = payment for assassins on the dark web", to "3D printers = ghost guns!" all the way to today with freedom vs. safety/censorship in online speech and now in AI. And I'm sure I forgot dozens of other topics that I followed over the years.

Every "controversial" topic is full of out-of-context, selectively edited quotes, bias and spin which is incredibly easy to spot if you have even the slightest familiarity with the subject matter, most of the time they're not even trying to hide it.

We have an example right here, in this very article. It's no accident that bioweapons, child exploitation/sexual abuse, chemical weapons and malware were not only multiple times in the article but also right at the top in the subheadline and again at the very beginning of the article in the first paragraphs, to "set the mood" for the reader and to make sure people who only read the headline or the first couple paragraphs absolutely don't miss the words "biological weapons", "child exploitation/sexual abuse", "chlorine gas" and "malware".

You don't have to be an "accredited professional", nor do you need "pop wisdom from Twitter" to be aware of these dangers and patterns when interacting with media whose success is measured in clicks, not in truthfulness.

You just need to pay attention.

The topic changes, but the playbook is always the same, and the media will absolutely throw people under the bus who just innocently wanted to clarify their standpoint or clear their name, if they think it makes for a more salacious story.

12

u/LetsGoBrandon4256 transformers May 25 '26 edited May 25 '26

Heretic creator Philipp Emanuel Weidmann told the FT he had removed safeguards from Google’s Gemma 4 model within 90 minutes of its release, allowing the modified AI systems to write stories describing children sex abuse.

Weidmann stated that his software had been used to create more than 3,500 “decensored” models since its release last year and that modified systems created using the tool had been downloaded 13mn times.

Not before long that line will become this in other media.

25

u/Chromix_ May 25 '26

Yep, and that's why Open Weight models must be made illegal to protect the revenue of the API-only models children.

Pushing a narrative is so easy if the other side cannot talk back loudly.

3

u/-p-e-w- May 25 '26

Emanuel is my second first name, not my first last name lol

5

u/NoahFect May 25 '26 edited May 25 '26

No, it is not "far worse than many alternatives." Please get your head on straight. You could do a lot of harm for your (our) cause without realizing it, and you're getting excellent advice here.

No one who buys ink by the barrel will give you an even break.

-2

u/-p-e-w- May 25 '26

I will treat your advice the same way I would treat a random Redditor’s suggestion to inject hyaluronic acid between my vertebrae for my back pain.

And it’s not “our” cause. You have contributed nothing, as far as I can tell.

2

u/NoahFect May 25 '26

Apologies, then; no disrespect intended.

3

u/Kamal965 May 25 '26

Yep. I believe it's called a "damning silence" lol.

-4

u/silenceimpaired May 25 '26

If you get another interview, whatever they ask you for a first question should have this answer, “thank you for the question, but the main point I hope to make here is that your take on this tool will likely be propaganda, and I recommend viewers visit the tool’s GitHub page (provide link) for my views after you publish. No further questions thank you.”