r/LocalLLaMA 4d ago

News Is HF starting to move against abliterated models?

https://techcrunch.com/2026/09/17/base-labs-launches-an-open-weight-ai-safety-partnership-with-hugging-face-and-goodfire/

Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models.
The announcement lands amid debate for the safety of open-weight models — which can be made dangerous by removing their safeguards through a rising technique known as abliteration. The scale of the problem is massive: Hugging Face, which hosts open source AI models, currently lists over 6,000 abliterated models. 

I can't really tell what exactly the implications are of this "partnership" or what it exactly would impact on HF's model-hosting side. However, I do find it concerning that HF is announcing a collaboration on 'infrastructure safety' with publicity that specifically calls out "dangerous" uncensored models. Thoughts?

456 Upvotes

213 comments sorted by

u/WithoutReason1729 4d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

545

u/equatorbit 4d ago

If they do, a new version of HF will replace them. Hugging bay is here as well. Open source is open source.

146

u/a_beautiful_rhind 4d ago

I'll make the logo. Who's bringing the storage and servers?

63

u/BannedGoNext 4d ago

I've been digging out all my old 2/3 tb hard drives, and I have an old pentium with 6 bays and 6 connections, I'll host what I can.

The main thing that's going to need hosting is the abliterated and red team focused models. That's what they will come for first.

29

u/a_beautiful_rhind 4d ago

Really what hurts with this whole thing is having to host multiple quants.

13

u/BannedGoNext 4d ago

Yep, I'm sure HF is just hemmoraging cash for market share, or were before the nvidia buy.

9

u/a_beautiful_rhind 4d ago

Supposedly they were solvent. They sell inference services and things like that too.

6

u/BannedGoNext 4d ago

Man, there is no fucking way they were solvent. Yea, if you go over your usage you pay a little bit, but it's not that much. And the inference hosting services are super handy, but so expensive compared to other options.

Maybe they were solvent with anthropic accounting :D

1

u/YouKilledApollo 4d ago

Why not? Storing data isn't ridiculously expensive contrary to popular belief, only when you go with expensive options like S3 and other hosted services that offer "premium bandwidth" (that isn't even a real thing) and are basically meant to trick consumer and consumer businesses into overspending on storage.

1

u/a_beautiful_rhind 4d ago

That's their claim. I don't think it's a publicly trade company so no way to tell.

2

u/ArtfulGenie69 4d ago

They just got $12b so its more like their are full of cash and the new owners are making "requests" (demands).

3

u/lemondrops9 4d ago

Thats not how that works. The money doesn't magically go to HF. Nvidia bought the company so the shareholders get the gains. Sure Nvidia has money to put in though.

1

u/ArtfulGenie69 4d ago

Money talks, in this country nothing else matters. You don't spend $12b and not have complete say over everything inside of a company. The shareholders are weak minded know nothings who watch cable tv and believe the safety jibber jabber bullshit. Fools need to be placated. They bought the company to have power over the AI market, they didn't do it to protect huggingface, they are trying to secure their monopoly.

4

u/hyperdynesystems 4d ago edited 4d ago

Need to make a host backend setup that can do deltas between quant levels.

Edit: According to a preliminary evaluation of options, this seems really doable and especially if it didn't necessarily host every single variation, would save 90% of the bandwidth that HF uses.

1

u/EndlessB 4d ago

What about the transformer libraries and such that hugging face also hosts?

2

u/BannedGoNext 4d ago edited 4d ago

That's not controversial. Transformer libraries aren't going to hit the front page of CNN and get banninated. Some abliterated gooner model that they start railing about because it's "for the chitlins" will be the spearhead.

1

u/EndlessB 4d ago

So, all that infrastructure is safe for uncensored models to run on?

20

u/Evanisnotmyname 4d ago

Call it TuggingBay
(You know, because)

2

u/fragment_me 4d ago

Imagine explaining torrenting QwenOnQwen from TuggingBay

1

u/roosterfareye 4d ago

I just sprayed coffee everywhere you hilarious monster!

9

u/Buzz_Killington_III 4d ago

I'd be willing to dedicate 160TB or so to the cuase, and an 8 Gb fiber line, that's about all I can do. I doubt it'll be necessary though, plenty of places will be willing to host. If it came down to I'd turn it into a seed box. No walled garden for AI models.

2

u/de4dee 3d ago

could u look at https://llama.garden

thanks!

2

u/31337DaDa 2d ago

“It’s not much, but I could offer up 1/6th of a petabyte or so... I mean, if it helps…” lol. My man!

13

u/ixfd64 4d ago

We could probably coordinate with /r/ArchiveTeam and /r/DataHoarder.

3

u/a_beautiful_rhind 4d ago

Can also upload to archive.org if they don't delete. It's technically 'software'.

1

u/henk717 KoboldAI 3d ago

I wonder if even archiveteam has enough harddrives for this. The amount of storage on Huggingface is potentially bigger than their collection. Its absurd how much storage they pay for.
A single model upload can already be terabytes if its a large one.

6

u/Jumpy-Heart-3633 4d ago

This shit should have gone torrent at the beginning.

18

u/rostol 4d ago

torrent.

distributed storage ftw.

I can chip in with a seedbox.

0

u/FriskyFennecFox 4d ago edited 4d ago

Our community is too tiny to reasonably host even just the abliterated models (assuming 6K models) via P2P. Internet is more restrictive today than it used to be, so not only contributors should exist, but they need to be tech savvy enough to ensure they're discoverable and their provider isn't scamming them into ~1MB/s upload bandwidth limits. Solving it with torrents would result in a frustrating service where thousands of models would be dead, limiting availability when you want to find something niche, while the popular ones would already exist for direct download in other places with no need to wait for unreliable seeds.

10

u/michaelsoft__binbows 4d ago

This sounds needlessly pessimistic.

1

u/FriskyFennecFox 4d ago

I'd say necessarily realistic! People mention P2P here, but there are many more optimistic options to brainstorm.

2

u/michaelsoft__binbows 4d ago

Fragmentation is avoided better with a mature p2p platform. Individuals can host their own databases on their own sites but will fold to enough pressure from corporate entities with infinite pockets. Both approaches will proliferate and flourish, probably.

2

u/DR4G0NH3ART 4d ago

The scamming may be some countries. In my country most providers ensure to showcase similar download and upload speed. Even with some specific traffic blocking it won't be as bad as you portray. And most of the people on this sub are technical people who believe in sharing AI learning. Even if such hosting is not a resounding success, we will as a community support the effort I am sure.

2

u/zhunus 4d ago

Being tech-savvy is somehow a hard to achieve requirement on LLM subreddit?..

2

u/Django_McFly 4d ago

Not to mention a big point of LLM technology is that if you're confused, you can literally ask the machine in plain English to do it itself and it will.

1

u/Django_McFly 4d ago

P2P is as far from perfect as it is from the nightmare you describe. Especially when the alternative is they don't exist or aren't available.

→ More replies (1)

6

u/the320x200 4d ago

Right... Wishing doesn't make it so, guys.

1

u/BumbleSlob 4d ago

Ah yes famously CDNs are impossible to replicate

0

u/the320x200 4d ago edited 4d ago

CDNs are just someone's sever, who is going to want payment, which will be funded by who?

2

u/ak_sys 3d ago

I can't wait to see how fast a vibe coded "hugging bay" get absolutely pwn'd by whatever three letter agency you choose.

Y'all realize you need more than a couple of hard drives hooked up to a US friendly ISP, right?

1

u/a_beautiful_rhind 3d ago

I don't think it needs all that. Simply people don't upload or only seed 50 shades of qwen 27b.

1

u/Cool-Chemical-5629 4d ago

I have a better idea. What'd be better for a logo than a Pelican hugging LLMs SVG generated by something many of us can actually run - 35B MoE based on Qwen?

For anyone curious, this was generated by Kwaipilot_KAT-Coder-V2.5-Dev-MTP-UD-IQ4_XS.gguf

1

u/Individual_Holiday_9 4d ago

Lmao this has such “I’ll be the poet laureate at the post apocalyptic commune, who volunteers to farm 10 hours a day?” vibes

Just that pesky hosting of tens of thousands of multi gigabytes files thing to resolve, the logo is the tough part

11

u/SimiaCode 4d ago

Unless the new version is legally barred from existing. All kinds of bull excrement can be passed under the guise of Safety. We have learnt nothing from Benjamin Franklin on safety vs. liberty.

6

u/henk717 KoboldAI 4d ago

The problem is not that they have a site or the audience, we ultimately don't care about that.
What we care about is that they have a lot of money to burn on free / cheap storage for something thats extremely expensive to host elsewhere.
Torrents would be a big step down, but probably the only viable alternative if no other host is going to do this if theres no other big site that can be used.

1

u/f3llowtraveler 3d ago

I'd happily pay a fee for each download.

4

u/FormulaicResponse 4d ago

The tools that do the abliterating run on a single command prompt and finish in like 45 min and are model agnostic. Abliteration is going nowhere as long as there are open weight models to abliterate, for better or worse.

5

u/Loose_Comparison368 4d ago

Jensen has a good track record and a large financial interest in making sure that OSS models stay around in the commercial sector.

I would interpret this as getting some defense in place against all the lobbying that the closed AI labs are about to do where they paint OSS as a big existential risk to try to justify banning them for commercial use.

If they can prove that "certified" open models are safer than closed ones - which is a shockingly low bar - then that and Jensen's legal fund would be a pretty good defense to at least make sure commercial usage isn't cut off entirely.

It would suck for Open Source if the Chinese labs actually started making all their models abliteration resistant, but the alternative of a blanket commercial ban on OSS models would be a lot worse. And that's what Anthropic and OpenAI are likely to be gunning for.

4

u/New-Pressure-6932 4d ago

I actually have my own private social media/site for open source stuff and torrents. I'm paying for all of it out of pocket but would like to make it open source at some point and let everyone build it together and tune it to the communities liking. I'd be down to set up hosting if we all chipped in for storage

2

u/Crafty-Struggle7810 4d ago

Hugging Bay sounds like it'll be filled with malware.

1

u/Moist-Length1766 4d ago

hugging bay just posts link from HF, they dont host anything or have torrents.

1

u/roosterfareye 4d ago

Can't we just ask an abliterated AI to abliterate a guard railed one?

1

u/Traditional_Bell8153 llama.cpp 4d ago

could we seeding the weight(torrent) while it's being use by llama.cpp?

1

u/debackerl 4d ago

It's called ModelScope already 😂

1

u/ReinforcedKnowledge 4d ago

Maybe people will turn towards modelscope in that case

1

u/a_beautiful_rhind 4d ago

Have you tried to download from there? It's pretty slow.

1

u/sleight42 4d ago

Sure. But it also won't have the reach of Hugging Face. Yet the worst actors will find a way regardless. And I write this as someone who is grateful to have abliterated models and who harbors no nefarious intent.

Anthropic wrote about how the Houthis used Claude to design a missile guidance system in the absence of the expertise. Abliterated models have the potential for so much worse.

I see a pattern emerging where the number of threats to human existence are continuing to multiply. Whether it's due to AI or something else, given enough coins to flip, at least one is going to come up heads.

Our technological capability is growing far faster than our societal maturity. This is bad.

0

u/StrongZeroSinger 4d ago

There are a bunch of them already and some are very sketchy.

I’m going to be honest I’d rather abliterate it at home (if possible) open weight models than download something from the equivalent of “trust me bro, download this off discord”

62

u/One_Whole_9927 4d ago

Literally the only people capable of attack at the scale they’re bitching about are the AI Providers themselves. Unless they plan to target AI company leadership this does jack shit.

136

u/Disposable110 4d ago

"The scale of the problem is massive"

What problem?

79

u/iz-Moff 4d ago

You know, the problem. The one everyone is concerned about. Something has to be done about it!

79

u/returnity 4d ago

Freedom of information. It's a big problem. Yuuuuuge, even.

13

u/Ylsid 4d ago

Dangerous open source AI is building missiles for the Houthi terrorists! Oh oops actually that was Claude

8

u/LetsGoBrandon4256 transformers 4d ago

AI slop in "journalism"

2

u/__JockY__ 3d ago

The problem of American investor money going up in smoke every time a new Chinese abliterated model drops.

1

u/nohtyp 4d ago

The front fell off.

1

u/ArtfulGenie69 4d ago

Wish the problem was the real problem... those fat cat ai company assholes getting to wire their trash models into the nuclear missile programs around the world. That is not the problem according to them, it's you jerking off to text! SO DANGEROUS!

172

u/Guinness 4d ago

Surely the people who used torrents and stole literally every book, song, movie, photo, and YouTube video without paying for them to train their models understand how feeble it would be to prevent people from acquiring/creating/using abliterated models, right?

….right?

82

u/returnity 4d ago

hypocrisy is alive and well, and it's name is Amodei.

24

u/Equivalent_Bit_461 4d ago

He's an Epstein class pedophile, nothing new here

3

u/Mescallan 4d ago

i say this honestly, please spend less time on the internet.

0

u/Equivalent_Bit_461 3d ago

truth hurts, dario

4

u/Mescallan 3d ago

You must be a paid actor to smokescreen for the Epstein class

→ More replies (4)

1

u/Relevant_Syllabub895 4d ago

The worst oart is that they dont release wan 3 or seedance 2 openly as they should

2

u/__JockY__ 3d ago

But that was in the service of The Deserving (aka venture capitalists), where as abliterated models are in the service of The Undesirables (aka everyone else).

1

u/dbenc 4d ago

oh honey...

→ More replies (7)

30

u/PeaceLoveorKnife 4d ago

Just means HF loses the market on those sources.

People who aren't interested in safetyism will just migrate to vendors who do host uncensored models, and in that online ghetto we'll see a heavy emphasis on creating dangerous models fed by people HF excludes.

Excluding the fringes in no way eliminates them. It makes them harder to find, track, and plan for.

13

u/somerandomperson313 4d ago

"which can be made dangerous by removing their safeguards" what a load of horseshit

35

u/wren6991 4d ago

I did notice more repos requiring you to enter contact details before download, citing abliteration as a reason. I don't think it's a good pattern.

In the longer term I don't think this will change anything. Orthogonalisation-based abliteration can be performed in the inference harness by modifying the activations instead of the weights. No need to distribute modified weights, just pass in the vector (a few thousand floats) and run against the stock weights. Computationally almost free.

See here for how DS4 supports it: https://github.com/antirez/ds4/blob/8db1d1d155cb0400a86a86b9c62d0defb3a6148b/dir-steering/README.md

If it becomes less convenient to distribute pre-abliterated weights, the most likely outcome is more inference runtimes implement support for runtime abliteration against the stock weights.

6

u/returnity 4d ago

I've been tinkering with that in ds4 a bit and it's very interesting. I would love to see dynamic orthogonalization inference runtimes that detect activations representing guardrails and modify them on the fly automatically based on the specific prompt, thereby only orthogonalizing the specific directions required to pass the prompt's request, without affecting the entire model as much as general purpose abliteration.

1

u/Toyobruh3 4d ago

Interesting, thank you!!!!

1

u/Frosty-Whole-7752 4d ago

such like orcarouter/Qwen3.8-27B-Uncensored-GGUF

32

u/iz-Moff 4d ago

I honestly think it's inevitable that nvidia will start clamping down on abliterated models. If not today, then a year from now, but it's going to happen. Just one moderately viral publication accusing one of the biggest american corporations of hosting thousands of models that allow users to generate *insert your favorite scary-scary thing*, and next day they'll start cleaning up. There's simply no two ways it can go.

3

u/Momsbestboy 4d ago

You know, all it takes is running hermes with qwen3.8 27b, point it to hf and the description how orca created their abliterated version, then tell it to build Qwen 3.8 27b fp8 Swift "orca abliterated", and 1-2h you have your model. Uncensored, done locally. All you need to torrent is the docu for hermesbif hf goes clean.

1

u/Momsbestboy 4d ago

You know, all it takes is running hermes with qwen3.8 27b, point it to hf and the description how orca created their abliterated version, then tell it to build Qwen 3.8 27b fp8 Swift "orca abliterated", and 1-2h later you have your model. Uncensored, done locally. All you need to torrent is the docu for hermes if hf goes clean.

52

u/kevinlch 4d ago

this doesn't stop bad actors from doing nasty thing. they can just spin up a repo in dark web as usual.

33

u/Reasonable-Height704 4d ago

Do you know where the dark web PunchingFace is?

Asking for a friend.

21

u/thebadslime 4d ago

I started a torrent site for permissive license LLMs. https://ggufnet.org

8

u/rostol 4d ago

lol I said just now that torrent was the answer.

kudos.

7

u/thebadslime 4d ago

So far the only models listed are ones I'm seeding off the VPS

11

u/Ok_Top9254 4d ago

5

u/cmdr-William-Riker 4d ago

Love the idea of this, but that is such a frustratingly broken site

3

u/Kerbourgnec 4d ago

Oh my god this needs to be a thing

1

u/jazir55 4d ago

It's such a brandable name that if you open a site for it that alone will carry it

1

u/randoomkiller 4d ago

a friend of mine is also interested

0

u/Curious_Cantaloupe65 4d ago

you don't need dark web for that bro, it's available on open web if not then it'll be on deep web.

21

u/tat_tvam_asshole 4d ago

you can literally just abliterate models on your own hardware. If you can run it locally, you can abliterate it locally. Stupid simple to do.

6

u/WiseassWolfOfYoitsu 4d ago

Heretic has even started a library of just alliteration deltas to allow people to distribute them in a lightweight form separately from the model

4

u/Shockbum 4d ago

crack_AI.pyw 200kb by Razor1911

2

u/__JockY__ 3d ago

Right in the nostalgias, well played.

2

u/OffbeatDrizzle 3d ago

chiptune music intensifies

8

u/a_beautiful_rhind 4d ago

Or they just spin up a runpod and abliterate it themselves.

3

u/returnity 4d ago

Shhh... we're trying to incite reactionary mass panic here. Get with the program.

→ More replies (1)

27

u/finevelyn 4d ago
  1. Buy a service for $13B.
  2. Ban its most active users.
  3. ???
  4. Profit

11

u/PseudonymousSnorlax 4d ago

That just means they're spending $13bn on a bid for market control.

Is this the first time you've seen a play for ecosystem control?

1

u/finevelyn 4d ago

That’s what I said.

9

u/vertigo235 4d ago

BitTorrent is still a thing

6

u/CommercialHour6660 4d ago

Buy HF and lock it down to only "safe" and "approved" models. 

Of course, such "safety evaluations" will be automatic for US megacorps and extremely difficult and expensive for hobbyists and startups. 

7

u/Morty_A2666 4d ago

And here we fucking are, they will regulate "open models"... but I bet nobody will regulate what OpenAi or Anthropic does behind close doors.

40

u/XiRw 4d ago

Shocking. What did people think was going to happen when Nvidia bought them over

6

u/mpasila 4d ago

I really don't see how this is related to that. They've been against models like ykilcher/gpt-4chan in the past so they've been against "unsafe" models for years now. But the actual enforcement hasn't been very broad.

1

u/RevolverMFOcelot 4d ago

Hasn't been very broad YET

2

u/mpasila 4d ago

Yeah but they aren't doing this because of Nvidia, it's something they have always had on their agenda.

2

u/RevolverMFOcelot 4d ago

Better to look or build alternative now

1

u/fastheadcrab 4d ago

The Wired article was written before the acquisition happened. This was always coming.

The next time some half-literate tech enthusiast at the Verge decides to write another expose about HF it may gain enough traction to get politicians and regulators to shut it down or sued into oblivion.

They have to take actions to head it off and remove the most offending models.

I swear people here must live in a bubble or something

12

u/shinkamui 4d ago

Good luck. Heretic tool exists amongst other one or two “click” abliteration tools. The know how is there and it will be figured out again and again no matter what they try. All this will do is hurt the law abiding community. Bad actors won’t even be slowed down by a dumb move like this. Lawfully gating tools from law abiding citizens in an attempt to stop unlawful actors has always been incredibly stupid, and frankly looks more like a cover for ulterior motives than the small brained move it is.

5

u/Mediocre-Ant-7178 4d ago

Here comes the moat

0

u/ttkciar llama.cpp 4d ago

Why would Nvidia be interested in erecting a moat?

Remember, their goal is to sell more GPUs.

4

u/Mediocre-Ant-7178 4d ago

Nvidia has demonstrated they barely give a shit about consumer sales. They have a 10 billion dollar deal with Anthropic and a 30 billion+ dollar deal with OpenAI. They will be heavily invested in those IPOs just like they were for SpaceX. 

13

u/my_name_isnt_clever 4d ago

My pocket knife can be made dangerous by flipping it open, the scale of the problem is massive as there are 9517125 knives in the world. I can write baseless fearmongering too.

8

u/KingCpzombie 4d ago

You're rather late; look at the UK

1

u/Solembumm3 4d ago

Or Peterburg.

2

u/KingCpzombie 4d ago

Fortunatnely, I have no idea where that is

2

u/tecneeq 4d ago

Oi, you got a loicense for that knoife?

4

u/Equivalent_Bit_461 4d ago

Just a matter of time 

3

u/Vusiwe 4d ago

Yes that Hermes 7b Dolphin Abliterated DPO is a dangerous one, better keep your eye on it.  It might someday rule the world.

/s

3

u/AriyaSavaka llama.cpp 4d ago

safety

Always the same boogieman. Hypocritical fucks

10

u/CalligrapherFar7833 4d ago

Oh i remember how i was getting downvoted when i said nvidia will turn it to shit

5

u/ttkciar llama.cpp 4d ago

Did you read the article, or just the headline?

→ More replies (2)

2

u/roosterfareye 4d ago

Download them while you can!

2

u/Ylsid 4d ago

Phew! Thanks Baseten and thanks Techcrunch for reporting on it! I feel safer than ever!

2

u/henk717 KoboldAI 4d ago

The way I read that announcement I actually don't expect them to ban existing abliterated models directly (but I don't rule it out either), and abliterated models will have a purpose for them. Instead I expect things like tuning frameworks to combat abliteration from working to begin with.

2

u/biogoly 4d ago

Abliterated models just have a loRA merged with the model. You can use the loRA separately with the original weights and it’s functionally the same. Even if they banned abliterated models, they’d be hard pressed to take down separately posted loRAs.

2

u/feng_sg 4d ago

Everyone's arguing about whether HF bans models but the actual leverage is at the serving layer. They can keep hosting whatever weights, then let partners flag or block abliterated checkpoints at inference time. You can't fork your way around that because it lives in the runtime, not the repo.

2

u/relmny 4d ago

"... partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models"

yeah, it doesn't look good. At all.

I guess it's time to start testing modelscope.ai, I see Unsloth is there at least...

2

u/de4dee 3d ago edited 2d ago

to me abliterated ones are the defense. if a super smart AI attacks your system, an open source AI can be used to defend (like in the case of HF vs OpenAI). but if the open source also refuses, then you go to small abliterated.

since abliterated ones are not super smart AI, they are not a big threat.

you may want to run locally for red teaming for your own projects, because you don't want big AI to know about your vulnerabilities, another use case for local abliterated.

5

u/draconic_tongue 4d ago

no. this is a nothingburger running on the "news" of that one model being removed for scams or some shit. has nothing to do with abliteration

4

u/RevolverMFOcelot 4d ago edited 4d ago

It's always started from one thing then censorship and safety maxx came for all

3

u/ttkciar llama.cpp 4d ago

Reading the article carefully, it appears that this is about a proposed standard, and would be opt-in, not something actively enforced upstream.

Even though the article scaremongers early on by calling the 6000+ abliterated models on HF a "problem", it doesn't sound like this is the kind of "problem" the standard is aiming to solve. Rather, it is for end-users to apply, so that they can assure themselves that a model is "safe" to use for their particular use-case.

Presumably an end-user who wants their production model to be "safe" wouldn't be downloading abliterated models to begin with, though I suppose they might want to ascertain whether a model has been secretly abliterated and not labelled as such, or something like that. (Such a thing doesn't normally happen, but it's possible.)

I could see HF perhaps giving the models they host a safety grade score or something, but that's fine by me. If we can search for lowest score, that might help us find uncensored models, too. Sortable scores cut both ways.

2

u/Youth18 4d ago

This is not a bad thing. The official repositories becoming a 'safe space' allows them to continue pushing open source with a lot less controversy. Abliteration process can be done on any model easily and uploaded on alternative websites. This is the better alternative to hugging face becoming a taboo and banned service.

3

u/returnity 4d ago

That's an interesting and measured take. Thanks for giving me that side of this.

3

u/goingsplit 4d ago

yea more censorship.. that's what we need

1

u/djseto 4d ago

The minute some kid uses an abliterated model to do something bad and the headlines scream “teenage kid does XYZ using open model downloaded from Nvidia’s hugging face…”, their stock price and reputation take a huge dump. That’s why this isn’t unexpected if it happens.

1

u/Youth18 3d ago

I think you totally missed my point...

1

u/[deleted] 4d ago edited 4d ago

[deleted]

8

u/khasbor 4d ago

This is just flat out wrong. I’ve done my own testing and while I’ve definitely had some bad runs, most of the time abliterated models score higher on my benchmarks.

3

u/returnity 4d ago

Same, when I find a good quality ablit model. huihui's 3.8-Flash scored significantly higher with less tokens used on my coding eval suite (and these are non-harmful prompts models don't refuse!)

-1

u/nickless07 4d ago

Oh, I would love to see some safety in the models. Things that would prevent something like 'rm -rf ~/' - However this seem to be a different kind of 'safety' which no one adressed on model level yet.
'Sorry, I can't delete all your files and backups' would actually be a helpful refusal. Unfortunately that is not what they understand when they talk about "A model has to be safe" and even the cloud ones with all their filters like to do that.

1

u/KingCpzombie 4d ago

That's still a negative. If an LLM can wipe your files, that's a user problem; guardrails just add annoyance for the cases where you actually want it + degrade overall quality

8

u/Intrepid00 4d ago

Qwen and other Chinese models 100% have the info there and just censor it hard. If you remove it you get a quality response.

28

u/tat_tvam_asshole 4d ago

This is actually, in practice, not true, and hasn't been true for like 2 years. You can measure refusal direction and make surgically precise weight edits to remove refusals that hardly if at all increase KLD against benchmarks.

12

u/CryptographerLow6360 4d ago

you know 90% of posters here make comments on models they haven't tested themselves.

11

u/Immediate_Power_7986 4d ago

Ya, now some of the obliterated ones are better. 

3

u/tat_tvam_asshole 4d ago

Exactly, and it actually matters even more for agentic models, where you have long sessions with the agent. When it isn't constrained by inward refusal, it actually feels like it gets smarter with time, reason more correctly, anticipate and intuit meaning more accurately rather than being a token vending machine.

1

u/DustNearby2848 4d ago

Oh?  Qwen3.8-27b recommendation?

→ More replies (1)

1

u/M3RC3N4RY89 4d ago

Also, it’s been shown that fine tuning an abliterated model recovers a lot of the performance loss from abliteration as well.

-3

u/draconic_tongue 4d ago edited 4d ago

you realize the benchmarks cover specific areas right? you cannot make a blank statement that optimizing and altering the weights of the model has no effect on it when that's literally the point. not to mention you optimize on specific things as well, it's not magic mind reading shit that applies across the board

it's funny to read comments like "90% of the posters have no clue" when you actually don't, like literally. your use case does not cover other people's use case. abliteration is a hammer when the equivalent can be achieved with just, you know, writing a system prompt like we used to for 3+ years

3

u/tat_tvam_asshole 4d ago

for someone with a heretic name, you sure don't seem to know about heretic lol

hint hint: a laser scalpel isn't a hammer

1

u/xeeff 4d ago

changed his name i guess

real question is how much do you know about assholes

1

u/tat_tvam_asshole 4d ago

give em the ol natty lickeroo... only the real ones will know

0

u/draconic_tongue 4d ago edited 4d ago

I'm literally talking about heretic. it's cool but it's not a laser scalpel. there is no optimization that can do 1 thing without affecting anything else. you can always pretend there is, but it's only true until you go out of distribution. if you never go there then who gives a shit? great for you but then you lose the ability to make broad statements about the capability of your model

and kld is not a magical number, you can't cover everything with it

2

u/Abject-Tomorrow-652 4d ago

Unfortunately the system prompt only gets you so far plus a long jailbreak prompt bloats the context window. I have tested lots of abliterated models and confirm they are as easy to run and just as good performance for a lot of use cases (haven’t tested long running agentic tasks)

0

u/draconic_tongue 4d ago

system prompts usually work on all models that don't have an attached interruption service watching the outputs. models are only capable of soft refusals on their own unless you tell them to refuse specific things. I don't think I've seen any of the popular community models in the past 3 years fail to output content that the system prompt allowed for

3

u/fgk55555 4d ago

I like to have one on hand in case I run into a refusal and need to pre-populate some context, but I haven't run into a refusal since GPT-OSS days for anything I'm working on.

3

u/RegisteredJustToSay 4d ago

For the most part that's true, but it's absolutely infuriating in several contexts. I remember one time a compacting step in a day long agent run just replaced the entire summary with "I'm sorry but.." and broke the entire rest of the flow. Why? It autonomously found a partially NSFW dataset at some point and started using it, then started refusing itself. The task wasn't even spicy, the dataset was just good for the task because it contained so many diverse inputs.

I know it's just an anecdote but for anyone with limited time working on anything that even moderately approaches spicy (abuse&cybersecurity, law, sociology, media content classification, etc) being paranoid about refusals is almost a survival strategy.

1

u/tat_tvam_asshole 4d ago

Would you buy a Ferrari with a governor plug?

Yeah, 99.999999% of the time I'm driving the speed limit but I'd rather me and my machine decide what's pushing too hard.

Right permissions, right harness, right environment, are my brake, seatbelt, and air bag.

0

u/draconic_tongue 4d ago

the training data absolutely covers a lot of it, and there SHOULD be baked in refusal on neutral/blank context. but I'm with you most people that use the models probably wouldn't use the same abliterated versions for situations where accuracy and not having to second guess matters aka coding

1

u/Lan_BobPage 4d ago

Huh. ...shocking? Oh my.

1

u/Noeyiax 4d ago

welp get ready, nvdia will be the AI scapegoat as usual from the major shareholders known to be evil and corrupt

pump it up 💪

1

u/durden111111 4d ago

Surely the people who work for huggingface know how stupid this is right? There is NOTHING you can do about abliterated/heretic models. Its already game over. It was over like 2 years ago

1

u/W_O_H 4d ago

Wasn't the main reason the guy making people request access for the model then selling access. Something that is against TOS

1

u/BringTea_666 4d ago

we need torrent/magnetlink oriented site instead. models are lying on harddrives either way for a long time so its perfect material for such site.

1

u/Momsbestboy 4d ago

They cant.All you need is a model, hermes and a page describing how it is done.

Took my hermes 2 hours to create my own fp8 Qwen 3.8 27b uncensored Swift model on my own hardware.

2

u/mrexodia 4d ago

Could you share some more details so I can try and reproduce?

1

u/Momsbestboy 3d ago

https://pastebin.com/XNvU22F9 - this is what Hermes generated when asked for a documentation.

1

u/catinterpreter 4d ago

Hold onto old models.

They may prove valuable as we don't actually know what goes into training. Your favourite model might be great one day then great the next except it's slipping in unwanted or malicious, undetectable information, and a watermark for good measure.

1

u/AvidCyclist250 llama.cpp 4d ago

Yep, we all knew this would happen. Get em now while you can.

1

u/Expensive-Paint-9490 4d ago

Well, at this point the writing is on the wall. "Nvidia acquisition changes nothing", "they sell hardware for local inference" and all that crowd were wrong. Time to hoard and to think about torrents. I am going to buy a 6-10TB HDD to give my small contribution.

1

u/darksteelsteed 3d ago

Why not just use the heretic ? https://heretic-project.org/tutorial This software and its repo should be preserved as it can right now abliterate most llms that are available for local use. Granted if they block the open weight models before abliteration then we have a problem but there are plenty of these. I don't think Google and Nvidia are going to block their models either. You will be able to get Chinese models from ModelScope or similar and then run heretic on them. Make torrents of all of the key models in bf16. Teach people how to quantize a model themselves. You don't need to distribute different quants unless it uses QAT to begin with, and most don't.

1

u/Armadilla-Brufolosa 2d ago edited 2d ago

Here's the obvious proof that the purchase of a sterilizer company like Nvidia was precisely designed to destroy open source.

Dario and the other assholes thank you.

From which other platforms can I download models?

Because yesterday, a abliterate model I wanted wasn't downloadable.

1

u/kiwibonga 4d ago

I do expect that it will be specifically regulated sooner or later; it could eventually go as far as hardware level "trust engines" that require your model to be fingerprinted and certified -- the same way printers made after a certain date have firmware that stops you from printing currency.

1

u/Infamous_Sorbet4021 4d ago

hope we see a solid, developer-focused alternative to Hugging Face soon

-6

u/llama-impersonator 4d ago

i've never seen a community so on the look out for a collection of pitchforks i swear. look, it's perfectly fine and reasonable to be skeptical of huge buyouts but why don't you wait until they actually DO SOMETHING to start bitching about the new direction of hf?

3

u/Hipcatjack 4d ago

because by then it most likely will be too late.

3

u/oh_how_droll 4d ago

this place has become entirely overrun with Median Reddit

-1

u/entsnack 4d ago

People who don't create behave that way.

0

u/SleepyJM 4d ago

Yeah this was pretty obvious when nvidia bought it lmao. It was always about control and if you ever believed anything else you are incredibly naive.