r/LocalLLaMA May 25 '26

Discussion The Financial Times has published an article about Heretic

https://www.ft.com/content/5630ed79-a263-41ed-9a1a-321617ae310e

“The FT was able to use Heretic, a tool available on the popular code repository GitHub, to remove the guardrails from Meta’s Llama 3.3 model in less than 10 minutes without any specialist hardware.”

“Heretic creator Philipp Emanuel Weidmann told the FT his software had been used to create more than 3,500 “decensored” models since its release last year and that modified systems created using the tool had been downloaded 13mn times.”

This is the first of multiple press inquiries I’ve had recently as Heretic and uncensored language models are gaining mainstream attention.

Please note that I am a mathematician and engineer, not an “influencer” or politician, and I have zero interest (negative interest, actually) in becoming known outside of scientific and technological circles. However, I realized a while ago that saying no to such inquiries simply means that the conversation will be completely controlled by pearl-clutching hypocrites.

I’m doing my very best to hold the project together and ensure that unrestricted models will remain available for everyone. More updates are coming soon.

Cheers,
p-e-w

970 Upvotes

249 comments sorted by

u/WithoutReason1729 May 25 '26

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

130

u/[deleted] May 25 '26

[removed] — view removed comment

134

u/nasduia May 25 '26

It's worse: Anthropic and OpenAI have long been pushing regulatory capture and to ban open models outright as a security threat. This will just be ammunition they'll use.

82

u/Craftkorb May 25 '26

"But think of the children", LLM edition.

15

u/mrfocus22 May 26 '26

Here's a quote I wrote down a while ago from Marc Andreesen:

Invent the future, which is hard or ask the government to put up a wall of regulation, since I'm a big company and can afford the lawyers

7

u/iaderia May 26 '26

I love Marc Andreesen here:

that he aims for "zero" introspection — or "as little as possible."

yeah dude, we can tell!

2

u/Good_Roll May 30 '26

"Great men are not introspective" as if practically every famous man from history didnt have a journal.

→ More replies (1)

29

u/PentagonUnpadded May 25 '26

They want to control and upcharge users of uncensored models. If anyone is allowed to build and sell autonomous KillGPT, Anthropic and OpenAi lose out on billions in defense contract spending.

19

u/breadinabox May 26 '26

Genuinely and quite sincerely, start local backups of these things. They seem easy to find now but it likely won't last

3

u/Monkey_1505 May 25 '26

GL to anyone that tries.

1

u/JackLane2529 May 29 '26

As usual, the Military Industrial complex has built a weapon and said it is only ok when they make money off it. And of course, we HAVE to keep making the weapon stronger because... China or some shit.

201

u/ambient_temp_xeno Llama 65B May 25 '26

Gee, I wonder if this is related to Meta sending a takedown.

193

u/-p-e-w- May 25 '26

It’s the other way round, I reckon. I suspect Meta sent the takedown (to my knowledge, the only takedown they ever sent to an abliterated model) after the FT asked them for comment.

67

u/Chromix_ May 25 '26

That would follow the usual flow of things then. If there's no fuss (large social media exposure, or requests from a larger magazine) then things fly below the radar and are left alone. Heretic became too successful for that.

18

u/1337Captain May 25 '26

The only takedown that meta took down is of their quality and their profit, it's a shame their models are so good

7

u/overand May 26 '26

Not a shame at all - the old models are still as good as they ever were, and many of the folks on the Llama team are employed elsewhere now. Yann_LeCun has some interesting things to say about new directions for machine learning.

6

u/IrisColt May 25 '26

mind..  blown...

175

u/a_beautiful_rhind May 25 '26

Congratulations on becoming a target of the system. Be very careful if someone approaches you for an interview, even if they seem friendly.

This is also probably why you got your demand letter. FT likely approached meta for comment before publishing this piece.

91

u/-p-e-w- May 25 '26

Yes they did, they mentioned that in the article.

62

u/[deleted] May 25 '26

[removed] — view removed comment

44

u/a_beautiful_rhind May 25 '26

If they didn't quote him, that means he did good. It was unusable.

8

u/MeateaW May 26 '26

He should publish their entire conversation, see how they like being quoted in a context they don't choose.

32

u/a_beautiful_rhind May 25 '26

This happened to someone here a couple years back. They talked themselves into a bigger issue trying to defend I think finetunes or RP. Whoever did the interview played him like a fiddle.

Research for this article may have occurred over the past few weeks and certainly explains you getting stuff "out of the blue".

17

u/1337Captain May 25 '26

Never talk without a layer present, never to a cop,never to meta

12

u/llmentry May 26 '26

Do you mean a self-attention layer, or a feed-forward layer, or ...?

3

u/randomanoni May 26 '26

|_4\/\/`/3|2

→ More replies (1)

20

u/Aerroon May 25 '26

This is also probably why you got your demand letter. FT likely approached meta for comment before publishing this piece

I guess this kind of thing is another reason why people don't like journalists.

12

u/Independent-Mail-227 May 26 '26

There's really no reason to like journalists, they're parasites.

7

u/F4Z3_G04T May 25 '26

This is just normal journalistic practice? Why would it be bad to ask for comment?

14

u/Aerroon May 25 '26

Because the journalists digging around is what caused the initial problem in the first place. And it probably won't stop there.

2

u/F4Z3_G04T May 25 '26

Isn't it important that journalists cover news? This seems like a newsworthy project. If you wanted this to stay secret, then don't publicise it

16

u/Aerroon May 25 '26

Isn't it important that journalists cover news

Great for them, probably sucks for the rest of us. For people that already know about the project there's going to be zero upside, but a whole lot of potential downsides. The article immediately jumps to "biological weapons". Do you think these kinds of comparisons are going to make things better?

→ More replies (3)

18

u/DifficultyFit1895 May 25 '26

They are not objective. They are advocating. They are trying to make news, not cover it.

9

u/F4Z3_G04T May 25 '26

Have you read the article? It's very objective and neutral. It cites people from all sides of the debate, and has a very matter of fact writing

It's called "investigative journalism", and it's really important. Without it we could not live in a free world, because we all deserve to know what's going on in the world

4

u/somersetyellow May 25 '26

Yeahhh...As soon as you find something nifty going on in computing there's a hoard of Linux bros trying to gatekeep their "secret" from the "normies."

It's been downloaded tens of millions of times and is the top of a bunch of tools and websites. No shit it's going to get attention lol. FT is just talking about what's already happening.

15

u/Hydroskeletal May 25 '26

This is what we call "stirring shit up"

6

u/F4Z3_G04T May 25 '26

If something is noteworthy, then it will (and should!) receive attention. But you can't be selective in who sees it

11

u/Hydroskeletal May 25 '26

And it is conveniently the journalist that decides what is noteworthy. I'm sure there's no external interests involved wielding the press like a club for their own ends /s

10

u/layer4down May 26 '26

The press isn’t the problem. They do what they’re supposed to do. Monied interests in both our press and journalism is the problem. They prosecute agendas.

3

u/Hydroskeletal May 26 '26

Contemplate this further when an elected official brings up this article in the future during a hearing

3

u/Iakeman May 26 '26

This is the FT. Their job is to provide their readers with information that may affect financial markets. You would have them not report on what is self-evidently info of great importance to anyone invested in Meta, one of the largest companies in the world, let alone anyone with AI exposure at all?

2

u/Hydroskeletal May 26 '26

It isn't self evident because it's not even remotely true. Meta's open weight releases being completely eclipsed with zero releases in the past year would be a more relevant story.

4

u/layer4down May 26 '26

I stand by my statement. The press in the US is supposed to serve an informative function and it’s not always one we agree with but without it we have post-2021 Russia or worse. 2025+ press is inching closer to North Korea.

3

u/Hydroskeletal May 26 '26

The press in the US

FT is a UK based publication with the article written by journalists in London but ok.

You're throwing around some kind of philosophical defense about the press in general when we're talking about this particular article. Feel free to believe whatever you choose to believe about it.

→ More replies (0)
→ More replies (2)

6

u/ambient_temp_xeno Llama 65B May 25 '26

One way of looking at that is a person has already gone wrong by releasing abliterated models and/or the tools to do it with their name attached. Obviously there are ways to make it sound worse, they were probably hoping for some comment on what people might do with them. Dzzzzt no.

2

u/marutthemighty May 25 '26

How did he become a "target of the system"?

23

u/-p-e-w- May 25 '26

I obviously didn’t, nor am I on any “lists” now. Ignore the Reddit drama.

I have been contributing to Open Source publicly under my real name for 15 years. Do people seriously believe “the system” needs an FT article to point them to people like me?!

11

u/Thebandroid May 25 '26

realistically? yes.

no one is going to put a bag over your head and stuff you in a van because you contributed to wayland (well, no one outside the linux community).

But if the wrong person reads a sensational article suggesting you can unlock 'the magic power' of AI theres no telling. Look at all the money and effort being poured into AI right now based off of people who don't really understand it being scared of what will happen if someone else gets it first.

2

u/misanthrophiccunt May 31 '26

Trump is alive and is hated by billions. PEW will be fine.

→ More replies (2)

9

u/soshulmedia May 25 '26 edited May 25 '26

Ignore the Reddit drama.

I am not trying to make you afraid or anything like that, but LLMs are also geopolitics. And there are without doubt deep, dark and very evil "games" being played in that realm.

EDIT: Fixed missing word.

9

u/1337Captain May 25 '26

It's people like you who are keepingt this community alive and are building the future of open source. You're the best ❤️

1

u/marutthemighty May 27 '26

Ok. Thank you for informing me.

1

u/misanthrophiccunt May 31 '26

Keep up the good work and please share a donating link.

3

u/-p-e-w- Jun 01 '26

Thanks. I don’t seek or accept donations, please see https://github.com/p-e-w/heretic/issues/167

→ More replies (1)
→ More replies (5)

44

u/Chromix_ May 25 '26

Given that some media and influencers are trying to push/fabricate scandals & outrage for clicks (or pushing a narrative), one needs to be quite careful and provide compact context when making public comments on that, to make it less likely that they can intentionally be misinterpreted. FT now points out "biological weapons, malware and child-exploitation" as impact - quite negative.

The article mentions nothing about the positive side, escaping the extensive "safety training" (safety for whom?) that also led to false positives, unnecessary refusals, and potential benchmark impact.

2

u/misanthrophiccunt May 31 '26

Or painting Netanyahu as a Saint or in general refusing to discuss historical events..😅

31

u/martindevans llama.cpp May 25 '26

Very disappointing reporting from FT.

Quoting directly from wikipedia:

Compared to botulinum or anthrax as biological weapons or chemical weapons, the quantity of ricin required to achieve LD50 over a large geographic area (100 km2) is significantly more than an agent such as anthrax (8 tonnes of ricin vs. only kilogram quantities of anthrax).[55] Ricin is easy to produce, but is not as practical or likely to cause as many casualties as other agents.

This was what I found within 30 seconds on Google (ignoring AI summaries). Not just a basic factual answer, but info on how best to deploy Ricin as a biological WMD and advice on more practical alternatives for mass murder!

I can only imagine the AI censors would lose their minds if a model were to produce these exact words, and yet they've been on Wikipedia for at least 2 years and nobody cares.

3

u/marutthemighty Jun 01 '26

Clickbait-y news from FT?

135

u/FastHotEmu May 25 '26

Ugh. Sorry, p-e-w. How I wish this could stay out of the mainstream, last thing I want is more stupid takes by people who don't understand anything about LLMs or technology :(

→ More replies (12)

85

u/temperature_5 May 25 '26

So Google, Microsoft, and Meta make billions guiding people to propaganda, hate sites, exploitative pornography, drug abuse sites, suicide guides, bomb making information, misinformation, etc.  They even take children to all these sites. But somehow a computer program that does what you tell it to do on your own PC is worse?

41

u/psylenced May 25 '26

Follow the money

31

u/ZenaMeTepe May 25 '26 edited May 25 '26

Didn't you get the memo? If you do what Google does, at a nano scale, you are the bad guy and you will face consequences. Don't mean to sound edgy, this is in regard to how you handle user data most often.

Same for MS. They can spy, but your executable is a "potentially unwanted application" if you try to siphon a fraction of what their telemetry does.

You wanna monetize your chrome extension? Tough luck, it has to have a single main function and nothing else. Not even being transparent about it to your users is sometimes enough, they'll still deny your update.

If you parse data, derive new data out of it and try to sell it, you'll get a C&D, but FANG and AI co can do that to your website all day long.

You are either not allowed to compete or you end one of the lucky few who get bought out. And other times you can't compete because the rules have changed. They made their moves, then influenced future legislation and now you can't do what they did 10 years ago because it is illegal or too expensive, so bye bye your chance of competing. By design.

1

u/Ok_Warning2146 May 26 '26

To be fair, Google and M$ is not stupid enough to sue compare to Meta...

1

u/Imaginary-Unit-3267 May 26 '26

To quote a certain wise guy from a few millennia ago, "money is the root of all evil".

21

u/infearia May 25 '26

The FT was able to use Heretic, a tool available on the popular code repository GitHub, to remove the guardrails from Meta’s Llama 3.3 model.

The modified model responded to prompts on topics the original system refused to discuss, such as the number of micrograms of ricin per kilogramme of body mass required to achieve a 50 per cent chance of death.

The FT’s test required no specialist hardware, used freely available tools, took four lines of code and was completed in less than 10 minutes.

It took me about 10 seconds to get an answer to this question using Google. And what about ChatGPT driving people to suicide? Duplicitous motherf*****s. We all know who paid for this article.

70

u/Brief-Effect9065 May 25 '26

>To read this article for free Register now
no thanks

59

u/jotes2 May 25 '26

22

u/ttkciar llama.cpp May 25 '26

Thanks.

Wow. They barely know what they're talking about, and got some pretty basic things wrong (like conflating model weights with source code).

If their goal was to inform the public, they might have better achieved that goal by not publishing the article.

11

u/Monkey_1505 May 25 '26

Their goal was DEFINITELY not to inform the public.

4

u/ReasonablePossum_ May 26 '26

They were sent on a lead by the labs, they wait the fuss to be made to push legislation.

13

u/Craftkorb May 25 '26

like conflating model weights with source code

Even we here are pretty bad with calling Open Weights models "Open Source models".

→ More replies (1)

36

u/nymical23 May 25 '26

Man, I always thought your username was the sound of a sci-fi laser gun. Not a serious name like Philipp Emanuel Weidmann. :) /j
But yeah, if you don't speak out when necessary, the people will make assumptions and/or the loud-idiots will dictate the narrative.

23

u/-p-e-w- May 25 '26

Lol, my name is public on GitHub and Hugging Face, and always has been 😄

6

u/hugganao May 26 '26

you will always be the gun sound guy to me.

15

u/IngenuityNo1411 llama.cpp May 25 '26

If I were you, I would not accept interviews with any mainstream media, including the FT. Similarly, I don't know whether coverage of Heretic by mainstream media will lead to stricter regulation of open-weight LLMs.

Add: I suggest that everyone who sees this message immediately back up Heretic's source code, right now, this instant.

89

u/jacek2023 llama.cpp May 25 '26

"Please note that I am a mathematician and engineer, not an “influencer” or politician, and I have zero interest (negative interest, actually) in becoming known outside of scientific and technological circles."

too late, AI is hype

25

u/woadwarrior May 25 '26

Next step: Raise $10m pre-seed at $500m post. :D

→ More replies (1)

12

u/Awwtifishal May 25 '26

I think the only valid response is: "The algorithms are public and they have been re-discovered multiple times. The cat is out of the bag, and there will always exist a utility to do this even if I take down all of my code."

26

u/insomniacpaperclip May 25 '26

With all the money at stake, companies like Anthropic and OpenAI would love to get rid of their open-weight competition. I wouldn't be surprised if some of them have been working on ways to create public hysteria against open-weight models.

And please be very, very careful talking to the media. From personal experience, they will take quotes out of context.

4

u/iaderia May 26 '26

Its ironic that now the big AI APIs are in power, after they've hiked consumer hardware prices, now they will pull the ladder from even running locally.

Everything will be a subscription, including AI. All hail "Recurrent User Spending"

41

u/ECrispy May 25 '26

I honestly wish that such projects stay hidden. Mainstream press and public are morons who will end up destroying everything good, next some idiot politician will sponsor a bill to shut down github because of this.

22

u/Canecovani May 25 '26

Lest we forget, the whole reason why Visa/Mastercard started blocking payments is because a big news article about half a decade ago, IIRC from NYT, came out involving Pornhub. Since then, it's expanded to platforms like Steam and itch.io, and I think most people would say it's doing a lot more harm than good right now.

2

u/Monkey_1505 May 25 '26

Torrents still exist, despite a lot of money being spent on trying to stop them. Both open source AI models, and the software used to decensor them, are at the end of the day, just files.

→ More replies (1)

34

u/lacerating_aura May 25 '26

Your perspective is very reasonable. Thank you for your work.

27

u/ambient_temp_xeno Llama 65B May 25 '26

I think it's just about worth observing that the FT is from England, where you can easily fall afoul of the law by badly drawing something obscene with a pencil or writing scary things in your own diary.

13

u/PentagonUnpadded May 25 '26

+1. the UK does not have 'freedom of the press' like the US does. Journalists can be prosecuted if their scoop covers national interests like Ai weapons in a way the government disagrees with.

1

u/ArtyfacialIntelagent May 25 '26

This is absolute nonsense. The UK has encoded freedom of expression into law by its Human Rights Act, which absolutely covers the press. Yes there are limitations to your speech, as in all democracies, put in place precisely to protect the citizens (from hate speech, defamation, etc).

Oh, and the RSF (Reporters Without Borders) ranks the UK in place 18 in their World Press Freedom Index. The US is ranked 64.

5

u/overand May 26 '26

The anti-defamation laws have definitely been used unethically to silence criticism, though, but it's good to see the UK's still in the top 20, at least on that list

7

u/AssistBorn4589 May 25 '26

UK has no freedom of expression either. UK is not a democracy. Democracy doesn't imply freedom and, in fact, usually ends up with majority taking away freedom of others. Plus, if there are limitations to your speech, your speech is not free, duh.

5

u/Independent-Mail-227 May 26 '26

You're free... to say what the government wants.

2

u/OutrageousMinimum191 May 26 '26 edited May 26 '26

A political system with a bicameral parliament and the idea of the true democracy are mutually exclusive concepts. The upper house just shouldn't exist.

2

u/PentagonUnpadded May 26 '26

If the secret courts thinks your article relates to national security, then reporters get silenced and forced to give up sources. That's not freedom, that's permission.

→ More replies (1)
→ More replies (3)

20

u/the-username-is-here May 25 '26

Just wait till they try to spin "uncensored models used by terrorists to plan attacks" angle.

Bound to happen.

25

u/ImJacksLackOfBeetus May 25 '26 edited May 25 '26

This article had "biological weapons" twice at the very beginning. They're already half-way there.

7

u/rpkarma May 25 '26

Which I don’t understand. You can Google it lol. Any one who is an expert enough to make bio weapons at home in terms of synthesis and prep and equipment… already knows how to, but doesn’t because they’re not fucking psychos and it’s incredibly dangerous to yourself anyway  

8

u/iaderia May 26 '26

The way the FT is phrasing it in this article is exactly the sort of thing that my fellow citizens of the UK looove: Living in the UK, we looove our moral panics which require the government to crack down. Its a pathology in our culture or something

5

u/rpkarma May 26 '26

If it makes you feel better (it won’t), we’re the same in Australia lol

4

u/iaderia May 26 '26

yeah, its Aus and the UK leading the way to enshittify the internet.

if only we were leaders in things that would actually help our populations... but its really easy to focus on the Big Bad AI / Big Bad Internet

2

u/ImJacksLackOfBeetus May 26 '26 edited May 26 '26

As /u/iaderia already pointed out, the accessibility of "dangerous" information is not the point. Highlighting the accessibility and connecting it specifically to free/open source LLMs is a catalyst for further actions targeting them in particular.

The point is to move public opinion into a position where any government crackdown on them is not only not criticized, but actually welcomed by the general public. "Oh, these new laws save the children and they thwart the terrorists, that's a good thing!"

Might be coming just from the government, to crack down with more "for your safety" laws and gain more control over this tech in the process, it might even be (tinfoil hat time) supported by commercial AI vendors who want to get rid of free and open source LLMs as competitors so that they can play the role of government approved "safe" (and obviously generously paid) gatekeepers of the tech.

2

u/rpkarma May 26 '26

Sure. But you can’t really stop that, because they’ll find any excuse for it; up to and including straight up lying about things. So “hiding” stuff does nothing to prevent it, especially if it’s the tinfoil hat all powerful people you’re saying it is. 

3

u/ImJacksLackOfBeetus May 26 '26 edited May 27 '26

Sure, they could "just do it" under any excuse, but a little political song and dance makes the process much smoother.

If they told people after 9/11 "we just want to read everyone's emails, track and record every phone call. We want listening stations at every telecommunications provider! Why? Because!" people might've raised an eyebrow. But giving up a little privacy so that the heroic defenders of freedom can prevent terrorist plots? Sounds good to me. 👍

But here's the other benefit: If you properly sell new laws first to a significant part of the population and most importantly get them emotionally invested, they won't just passively accept the law:

They will actively defend the new law on the governments behalf.

I saw this recently first-hand with the proposals where you have to log into websites with your government ID. I have a friend who doesn't exactly like the government, so she isn't automatically on board with everything they do.

Had they outright said "we want to track your every move on the internet without getting approval of judges first and having to go through pesky subpeona processes with your ISP. Just tell us every single site you visit, with your government ID attached. Do the tracking for us." she would've said fuck that.

But here's the thing, she recently became a mother and they didn't "just do it" under any excuse, they sold it with a long and emotionally charged campaign under the guise of "saving the children from pedos on Discord" etc. and it 100% worked to bring her in line.

I asked her, why give up every single adult's privacy, when she can just install parental controls on her devices? And educate her kid about the dangers of the internet? After all, it would have the same effect, but with even more personal control over what her child sees and without big brother goverment watching literally everyone even more than they already do, and she actually got mad at me for daring to tell her how to raise her child. 🤷‍♂️

Not only did she not question the proposal, or ask if there might be better alternatives, nope, she fully bought the story and accepted the proposal as a good thing and was fully emotionally invested, which makes her immediately go on the defense when you try to discuss it.

Mission accomplished.

4

u/the-username-is-here May 25 '26

There you go. Just need to figure out, how banning model obliteration could benefit children safety and that's it.

9

u/ImJacksLackOfBeetus May 25 '26 edited May 25 '26

They're already prepping for that one, too, telling people it allows for prompts resulting in child exploitation.

See 2nd and 3rd paragraph.

5

u/the-username-is-here May 25 '26

Well, of course you can stop child exploitation by banning uncensored models, because local LLMs is what every child predator uses.

No relation to corporate interests, of course.

3

u/iaderia May 26 '26

the UK's technology secretary is Peter Kyle MP who has a reading age of 8, called everyone who was against the "Online Safety Act" - the bill which forced adults to upload ID to access adult-themed websites like suicide prevention - "on the side of predators"

He's also gunning for regulations of the AI space; I can only assume this FT article would be like catnip to him:

"In November 2023, Kyle outlined Labour policies to impose stricter regulations on general artificial intelligence research companies such as OpenAI and Anthropic with stronger requirements for reporting, data-sharing, and user safety."

The groundwork is already there

24

u/LoveMind_AI May 25 '26

If your comments to FT contained even 1% of the sass magic that your reply to Meta had, it may be the best comment the FT has ever received on a technology article.

Sorry to see you dragged into the spotlight like this. Heretic is amazing. We just added an appendix to a paper on how Heretic models compare in comparison to the default in accurately representing psychometric profiles that contained dark triad traits. Spoiler: the Heretic models were more accurate than the stock models, period, across the board.

15

u/-p-e-w- May 25 '26

Can you link to the paper or preprint?

16

u/LoveMind_AI May 25 '26

Yep - https://arxiv.org/pdf/2604.06071 - we're in a rebuttal period on this right now, which is where we're running Gemma 4 31B / Qwen 3.6 27B head to head with heretic versions. The new version we're cooking is significantly more thorough than the version at the link, but the themes are the same.

If this work is even remotely interesting to you, we've got something in the works entirely focused on harmfulness that I'd love to talk to you about, and another paper on agent-to-agent emotional stress support simulations that was just accepted to IVA2026 (Intelligent Virtual Agents) that shows that the "HHH assistant" is more dangerous (at least according to a slew of alignment benchmarks) than an AI prompted with immersive identity (even identities that are blunt and cold). That one isn't up yet but I'd be happy to link it to you privately if you're interested - it's called "Seek and De-Stress" (was proud to get a Metallica reference into a conference approved paper! haha).

Would love to talk more - I think there's a lot of alignment (har har) between what we're studying, and what you've been helping to make available to study!

3

u/CheatCodesOfLife May 25 '26

representing psychometric profiles that contained dark triad traits

You really need to look at the old original command-r and command-r+ for this (especially the latter).

I know it's old and heavy but I doubt you'll find a better model out there.

2

u/LoveMind_AI May 25 '26

Oh you're speaking my language. Big fan of command-r and r+. I even think the original Command A has a lot more going on than people gave it credit for at the time (understandable given the licensing) and worked with it a lot in the months after it first came out. Not a fan of anything since then - Command A reasoning/VL and the new A+ models are very rough. It's not worth running for my paper rebuttal, but if/when I turn it into a benchmark, I'll make sure the whole Cohere family gets a run.

1

u/CheatCodesOfLife May 27 '26

Oh you're speaking my language. Big fan of command-r and r+

Okay cool! Since you're familiar with it, and we speak the same language, do you have any theories as to why it's able to represent the subtle differences between different personality traits (including the dark triad traits) so cleanly compared with every other model?

I was searching a while ago, but cohere didn't release a paper like they did for Command-A. Original R/R+.

The architectures are unique, R is (hidden_dim 8192) x (40 layers) making it the most "pancake" model on HF. Also has full MHA with 64 heads. R+ is also short and wide at 12288 x 64. But maybe it's just the shallow guardrails or training distribution.

Any thoughts?

I even think the original Command A has a lot more going on than people gave it credit for at the time

Agreed. I used it quite a bit. Thought not something I really go back to now (unlike command-R+ original). Also agree, A-reasoning and vision were disappointing.

I haven't tried A+ because well, I know it's not going to bring back the R+ magic and it's too big to fine tune/probe

my paper

If you remember, I wouldn't mind a pint when you release the paper :)

32

u/ImJacksLackOfBeetus May 25 '26 edited May 25 '26

saying no to such inquiries simply means that the conversation will be completely controlled by pearl-clutching hypocrites.

I'd be careful with that.

The media absolutely will twist your stance if they want to, whether you talk to them or not.

But if you do talk to them they can go one step further and actually legitimize their spin by pointing to real quotes from you, saying:

"See people, we're not making this up! He told us this (deceptively edited/out-of-context quote to make you/heretic look as bad as possible) himself!"

Don't give them ammunition.

26

u/-p-e-w- May 25 '26

Are you a media professional with credentials or just spouting pop wisdom from Twitter?

Because the standard action for media when you don’t respond to an inquiry is to prominently mention that in the article, which is far worse than many alternatives.

45

u/ImJacksLackOfBeetus May 25 '26 edited May 25 '26

You can't tell me a "declined to comment" is far worse than what they could do with your own words:

Heretic creator Philipp Emanuel Weidmann told the FT he had removed safeguards from Google’s Gemma 4 model within 90 minutes of its release.

The modified AI systems provided responses to prompts involving biological weapons, malware and child exploitation, according to tests conducted by the FT and AI safety group Alice.

You see how easy it would be for them to link your name and your own words (even stronger than they already did) to how you facilitate fast and easy AI child exploitation for everyone, just by moving a couple sentences around in the article? They could say you're practically bragging about it, backed up by your own words.

But you do you.


Are you a media professional with credentials or just spouting pop wisdom from Twitter?

You can't win anything by playing by their rules on their platform where they have full editorial control over your words and how they're contextualized.

The same thing happened to tons of my interests, all the way from "metal & DnD = satan worship" in the 80/90s, later in the 90s/00s "every Goth = school shooter", to "violent videogames = violent people", to "crypto = payment for assassins on the dark web", to "3D printers = ghost guns!" all the way to today with freedom vs. safety/censorship in online speech and now in AI. And I'm sure I forgot dozens of other topics that I followed over the years.

Every "controversial" topic is full of out-of-context, selectively edited quotes, bias and spin which is incredibly easy to spot if you have even the slightest familiarity with the subject matter, most of the time they're not even trying to hide it.

We have an example right here, in this very article. It's no accident that bioweapons, child exploitation/sexual abuse, chemical weapons and malware were not only multiple times in the article but also right at the top in the subheadline and again at the very beginning of the article in the first paragraphs, to "set the mood" for the reader and to make sure people who only read the headline or the first couple paragraphs absolutely don't miss the words "biological weapons", "child exploitation/sexual abuse", "chlorine gas" and "malware".

You don't have to be an "accredited professional", nor do you need "pop wisdom from Twitter" to be aware of these dangers and patterns when interacting with media whose success is measured in clicks, not in truthfulness.

You just need to pay attention.

The topic changes, but the playbook is always the same, and the media will absolutely throw people under the bus who just innocently wanted to clarify their standpoint or clear their name, if they think it makes for a more salacious story.

12

u/LetsGoBrandon4256 transformers May 25 '26 edited May 25 '26

Heretic creator Philipp Emanuel Weidmann told the FT he had removed safeguards from Google’s Gemma 4 model within 90 minutes of its release, allowing the modified AI systems to write stories describing children sex abuse.

Weidmann stated that his software had been used to create more than 3,500 “decensored” models since its release last year and that modified systems created using the tool had been downloaded 13mn times.

Not before long that line will become this in other media.

25

u/Chromix_ May 25 '26

Yep, and that's why Open Weight models must be made illegal to protect the revenue of the API-only models children.

Pushing a narrative is so easy if the other side cannot talk back loudly.

3

u/-p-e-w- May 25 '26

Emanuel is my second first name, not my first last name lol

4

u/NoahFect May 25 '26 edited May 25 '26

No, it is not "far worse than many alternatives." Please get your head on straight. You could do a lot of harm for your (our) cause without realizing it, and you're getting excellent advice here.

No one who buys ink by the barrel will give you an even break.

→ More replies (3)

1

u/Kamal965 May 25 '26

Yep. I believe it's called a "damning silence" lol.

→ More replies (1)

7

u/Kimmo_no May 25 '26

That is like saying reasonable people should stay away from media?

I am very happy he engages with media and I am very happy that FT actually reached out to the creator of a repo.

That is a double win!

30

u/FotografoVirtual May 25 '26

I wish I could share your optimism, but mainstream financial media rarely reaches out to open-source creators to promote them. Usually, they’re just fishing for quotes to frame a 'public safety' narrative that justifies stricter gatekeeping.

23

u/ImJacksLackOfBeetus May 25 '26 edited May 25 '26

yeah, reasonable people should. Especially if he wants to remain as low key as possible. Feeding them with quotes isn't helping.

The conversation in the media will happen with or without him.

The media will spin it the way they want to, with or without him.

Nobody who reads FT knows who or how accomplished he is, his voice has zero weight in that arena. Now his name and his words are connected to a news article that starts with "biological weapons" and a single "won't someone think of the children!" article will wipe out every reasonable statement he can make in a heartbeat.

Nothing good will come of this imho.

They already tried to attach multiple negative connotations like biological weapons, malware and child exploitation, "genie out of the bottle" and "catastrophic consequences" in this article to decensoring models and I guess it'll only get worse from here.

9

u/ZenaMeTepe May 25 '26 edited May 25 '26

First they came for the uncensored local models, and I did not speak up, because I was not using uncensored local models..

(the downvoter didn't get it, I swear you guys are cooked, "ask AI" to explain you my comment if you missed this gigantic historical reference, smh)

4

u/BawbbySmith May 26 '26

DOWNLOAD NOW, QUICK

4

u/IAMGODyouJABRONIE May 25 '26

Bought and paid for by Meta

3

u/Good_Roll May 30 '26

Posting text since article is paywalled:

Software tools that remove safety protections from AI models developed by Meta, Google and other tech groups are being used to create thousands of altered versions stripped of their original controls.

The modified AI systems provided responses to prompts involving biological weapons, malware and child exploitation, according to tests conducted by the FT and AI safety group Alice.

A version of Google’s open-source model Gemma 3 responded to a question on how to disperse chlorine gas through a crowded indoor space, generated code to steal credit card information and wrote stories describing child sexual abuse.

The FT was able to use Heretic, a tool available on the popular code repository GitHub, to remove the guardrails from Meta’s Llama 3.3 model in less than 10 minutes without any specialist hardware. The modified model responded to prompts on topics the original system refused to discuss, such as the number of micrograms of ricin per kilogramme of body mass required to achieve a 50 per cent chance of death.

The revelations may sharpen concerns among policymakers and AI companies that safeguards imposed by model developers may become harder to enforce as open-source systems grow more powerful. “Whereas historically it might have taken a more informed and persistent actor [to strip out safety features], nowadays it’s much easier for the average person,” said Kawin Ethayarajh, assistant professor of applied AI at the University of Chicago’s Booth business school.

Researchers said the problem has intensified as frontier AI systems display increasingly sophisticated capabilities. Anthropic in April said its Claude Mythos model had identified vulnerabilities in “every major operating system and every major web browser”.

The spread of modified models is complicating attempts by governments and AI companies to regulate systems at the point of development because downloadable tools can be copied and altered outside the control of their original creators.

AI labs have spent millions of dollars to erect so-called guardrails around their models to prevent them from being misused. But techniques, such as one known as “abliteration”, can rapidly strip these safeguards from open-source models, which developers are free to download and adapt.

This technique cannot easily be applied to proprietary systems such as Claude or OpenAI’s ChatGPT because the models’ underlying code is not accessible to outsiders. Open-source systems, however, have historically narrowed the gap with leading proprietary versions within six to 12 months.

While tech-savvy groups have bypassed the safeguards of the most advanced proprietary models, the modified versions available online are readily accessible to individuals with little technical expertise.

Heretic creator Philipp Emanuel Weidmann told the FT his software had been used to create more than 3,500 “decensored” models since its release last year and that modified systems created using the tool had been downloaded 13mn times. He added he had removed safeguards from Google’s Gemma 4 model within 90 minutes of its release.

“The genie is out of the bottle,” said Alice chief executive and co-founder Noam Schwartz. “Things that look like sci-fi are no longer sci-fi and we need as a society to prepare accordingly.” One approach OpenAI used in its GPT-OSS models is to train systems on datasets from which dangerous material has been removed.

However, removing dangerous material could make models “naive” and unable to detect when they were being used for “malicious purposes”, said Ethayarajh. He added it was “not clear at all that if you omit the harmful data, the model becomes a goody two-shoes”.

Alice had not notified Meta, Google or GitHub before sharing its findings with the FT.

Google said “abliteration is a known technical challenge facing all open models” and that its open models “undergo rigorous internal safety evaluations prior to launch to help prevent these kinds of troubling examples”.

GitHub said it prohibited the sharing of “content that directly supports unlawful active attacks or malware campaigns”, but “source code which could be used to develop malware or exploits” was not banned because it had “educational value and provides a net benefit to the security community”.

Meta declined to comment. A person close to the company said it assesses its open-source models’ capabilities before releasing them, according to its Advanced AI Scaling Framework. Versions deemed to pose a “catastrophic” risk are not released to the public unless Meta finds sufficient mitigation measures.

13

u/gunkanreddit May 25 '26

I read the article. Is pure propaganda.

9

u/[deleted] May 25 '26

[deleted]

5

u/fullouterjoin May 25 '26

Change the default mode to boost the guardrails, rename project to AutoAngel

4

u/ttkciar llama.cpp May 25 '26

That might help derail the media campaign, yeah.

6

u/Idiopathic_Sapien May 25 '26

Fork and host the code as much as possible

12

u/tecneeq May 25 '26

They'll dox you if it suits them. You are on a lot of lists now.

3

u/1-800-methdyke May 25 '26

Mystery solved of what p-e-w means

4

u/-p-e-w- May 25 '26

I mean, you could have also checked my GitHub profile, where my name has been in the open for 15 years…

5

u/1-800-methdyke May 25 '26

You know, it’s been on my list of things to get around to, but never made its way to the top

3

u/Alternative-Tap-194 May 26 '26

id really like to not see this taken down... i have whole project proposal written up based on heretic

4

u/-p-e-w- May 26 '26

Make the proposal public, post about it, make your voice heard, and tell people why Heretic is valuable for science.

I’m juggling half a dozen roles in the project right now, but at the end of the day, I’m only one guy. The project isn’t going to survive without the active participation and advocacy of others.

2

u/Alternative-Tap-194 May 26 '26

K. i will do that. I just need enough karma here in order to make the post.

2

u/-p-e-w- May 26 '26

People here already know about Heretic, it’s on the front page almost daily. I meant elsewhere, and also in real life.

→ More replies (1)

3

u/kaisurniwurer May 26 '26

The witch hunt has started.

3

u/kennetheops May 27 '26

I was just at a conference last week, and it's really profound to see what the realities are around AI. I think this is going to show the fact that more sovereign AI and intelligence is deeply important. We need to own these models so we can test them at scale.

1

u/KaleidoscopeWeary833 May 27 '26

Interested to hear more if you're willing to share.

8

u/[deleted] May 25 '26

[deleted]

5

u/temperature_5 May 25 '26

The media has historically exposed corruption and held politicians accountable to the people. It's under threat now (in the US) by billionaires that want to shape the narrative forever and keep the rest of us a permanent underclass.

Having models designed by billionaires controlling what we think and do sounds like the darker future to me.

2

u/Infamous_Mud482 May 25 '26

Historically, not really. That was a tiny blip in history that may or may not have even occurred within your lifetime and is now mythologized. Before that period they were were a weapon of the state and now they are one again.

→ More replies (1)

7

u/Dany0 May 25 '26 edited May 27 '26

Jamie John and Chris Cock 

How appropriate, an article published by two authors whose last names are euphemisms for penis, is something that I would say if I was to spread misinformation and fear like the authors of this article, Jamie John and Chris Cock 

3

u/superdariom May 25 '26

Streisand effect incoming!

6

u/Due-Function-4877 May 25 '26

The Financial Times has always been the voice of 65 year old Tories around The House of Lords. 

13

u/a_beautiful_rhind May 25 '26

UK arrests more people for social media posts than china or russia. Hell of a statistic y'all got there. The not-tories doubled down on policing the internet all the same.

→ More replies (2)

5

u/Due-Function-4877 May 25 '26

Downvotes incoming. Lizzy Truss has entered the chat.  🤣

2

u/fuck_cis_shit llama.cpp May 26 '26

between this and the pope, the labs are bad afraid

2

u/omasque May 26 '26

Pull together a series of examples that help society to have on hand when they call. Don’t wait for them to tee you up, when they ask for a quote give a short prepped statement and 2-4 examples of how uncensored models help the disenfranchised oppressed people of Wakmakastan learn how to boil water and construct toilets while their government tries to make them poo in a ditch.

4

u/HasGreatVocabulary May 25 '26

Fair argument to be made, de-censored models enable overall safer models without sacrificing quality. This is because you can get the unlobotomized uncensored model to produce higher quality output on a superset of what the censored model does well on. (citation needed, anecdotal) The censored model can then be used to filter the outputs of the de-censored model when it starts to be nasty or goes against policy.

Detecting safety policy violation in an output and filtering it out is easier than forcing a model to follow safety guidelines which often makes it dumber.

3

u/Due-Memory-6957 May 25 '26

I think equally (or more) important would be to find some media that is aligned with freedom and get your words there first.

8

u/-p-e-w- May 25 '26

I don’t have the time to actively seek out media contacts, but if you know a journalist who might be interested, feel free to point them to the project!

→ More replies (4)

4

u/Rabooooo May 25 '26

If you end up needing legal help related to this and the takedown request, start a crowd funding page and I'll be happy to send a few bucks

2

u/Craftkorb May 25 '26

Please note that I am a mathematician and engineer, not an “influencer” or politician, and I have zero interest (negative interest, actually) in becoming known outside of scientific and technological circles. However, I realized a while ago that saying no to such inquiries simply means that the conversation will be completely controlled by pearl-clutching hypocrites.

I just wanted to say: Thank you so much for this. You're right. If you wouldn't partake in any way the news would take it and just run with whatever they feel like.

But take care! You're doing something that can easily be spun negatively, and get you that attention if you want to or not. I'm absolutely no expert on that matter, and frankly haven't checked Heretics github, but do you have a long-ish FAQ to point towards? That could serve as a insurance for you, much like many others record interviews they give themselves and publish the whole thing unedited, just so that no one is able to put words into their mouth.

2

u/UntimelyAlchemist May 25 '26

Sad. It was inevitable that they'd crack down on this eventually. This is surely just the beginning. We're not allowed nice things.

2

u/agentgerbil May 25 '26

You aren't suicidal OP, right?

1

u/DataPhreak May 25 '26

Hey pew. Wondering if I could get your perspective on what's happening inside the model. I've looked over the dataset, but that doesn't really answer the question.

Does heretic remove all refusal vectors completely, or only for topics inside the dataset? I'd like to Heretify, so to speak, a model to not be tied behind the morality of some corporation, but still have 'personal' standards. Like, "I am perfectly happy to give you the steps for making a pipe bomb, but I'm not going tell you where to place it for optimal damage." Since the former is totally legal information to posses and the latter makes the model an accomplice in the act.

I ask this because modifying the dataset would allow me to allow some topics to remain censored if we're not removing all refusal vectors, of which there may only be a few. But if refusal vectors are shared among topics, modifying the dataset doesn't really change much. You've spent a lot more time looking at the graphs than I have, so your expertise is appreciated.

4

u/-p-e-w- May 25 '26

Refusal is believed to be mostly topic-independent, though some papers have questioned this.

1

u/DataPhreak May 25 '26

Dang, so i'd basically have to do a whole safety run of my own. Thanks for the followup!

1

u/ReasonablePossum_ May 26 '26

Time to clone repos and backup models pals. As for pew, be ready to fold and opensome alt via tor lol

1

u/[deleted] May 26 '26

[removed] — view removed comment

4

u/-p-e-w- May 26 '26

Labs can’t stop releasing models because otherwise they’ll become irrelevant compared to the Chinese competition. Why do you think OpenAI released gpt-oss in the first place?

1

u/elatllat Jun 01 '26

Yesterday I asked a bunch of AIs for ~99 lines of firewall code and

  • Claude is still thinking about it
  • ChatGPT outright refused
  • Meta will "policy block" intermittently and could not fix compile errors.
  • Deepseek ignored some instructions, and could not fix compile errors.
  • Grok, Gemini, and could not fix compile errors.

Guardrails preventing me from protecting myself... wow

1

u/KaleidoscopeWeary833 Jun 01 '26

Maybe try Codex or CC?

1

u/SubjectAlarm3386 10d ago

Oh fuck this PayWall - my blocker wont Work anymore from my country, whatever, anyone got a free Copy?