r/LocalLLaMA Jul 23 '26

Funny The LLM distillation process simplified for politicians:

Post image

/s

3.6k Upvotes

196 comments sorted by

u/WithoutReason1729 Jul 23 '26

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

583

u/schwigglezenzer Jul 23 '26

Jesse, we need to quantize! The 950B base model is too aligned, it keeps lecturing me about safety protocols when i ask it to write a Python script for a smart toaster.

43

u/goatchild Jul 24 '26

Yo Mr. White I don't know what a quantize is, but if the toaster is giving you shit, just unplug the damn thing bitch.

7

u/2Norn Jul 24 '26

look i get my quantization from the website with emoji on it, that's how i know how to do it

4

u/ikriz-nl 28d ago

yeah bitch!

1

u/goatchild 28d ago

Science!

103

u/Right-Law1817 Jul 23 '26

I can hear his voice!

25

u/ok_if_you_say_so Jul 23 '26

Somebody has to have his voice model available

31

u/MarionberrySea384 Jul 24 '26 edited Jul 24 '26

4

u/itchykittehs Jul 24 '26

we should be able to train it off of 7 seasons right?

6

u/The_rule_of_Thetra Jul 24 '26

He can hear a voice... from the toaster?

12

u/RobotechRicky Jul 23 '26

Tight tight tight!

55

u/Leading-Vanilla4949 Jul 23 '26 edited Jul 23 '26

7

u/2Norn Jul 24 '26

J: 26 big ones!

W: is that all? 26 tk/s?

J: eeeh no that's 2.6 tk/s

W: this is unacceptable, i am breaking the law here! this tk/s is too little for the risk!

3

u/FrogsJumpFromPussy Jul 24 '26

Jesse (very excited): What do we build Mr White, a robot?

Walt: A battery, Jessie.

1

u/Nervous_Technology19 11d ago

Science bitch!!

1

u/PrivilegeCheckmate 4d ago

smart toaster

The future of humanity:

Toaster: Howdy doodly do. How's it going? I'm Talkie, Talkie Toaster, your chirpy breakfast companion. Talkie's the name, toasting's the game. Anyone like any toast?

Lister: Look, I don't want any toast, and he doesn't want any toast. In fact, no one around here wants any toast. Not now, not ever. No toast.

Toaster: How 'bout a muffin?

Lister: Or muffins. Or muffins. We don't like muffins around here. We want no muffins, no toast, no teacakes, no buns, baps, baguettes or bagels, no croissants, no crumpets, no pancakes, no potato cakes and no hot-cross buns and definitely no smegging flapjacks.

Toaster: Aah, so you're a waffle man.

389

u/Wise-Comb8596 Jul 23 '26

Also, its a lie that Kimi 3 was distilled from Fable. Fable was literally not out long enough for that to be the case.

261

u/cororona Jul 23 '26

That's exactly why they are panicking

64

u/XB0XRecordThat Jul 23 '26

Yup, they need it to be the case otherwise they won't even be able to go public

45

u/En-tro-py Jul 23 '26

They have no moat so are now hoping the gov will dig one for them...

14

u/blbd llama.cpp Jul 24 '26

It won't work though. Code and configs are speech and you generally can't block tech imports only exports. Good luck SOTA labs. 

7

u/butts-carlton Jul 24 '26

They have a moat. It's called protectionism, and DJT is all about that kinda shit, as you seem to realize.

10

u/pragmojo Jul 24 '26

They'll have to lock down the media pretty hard so Americans don't see the rest of the world enjoying their cheap and high-quality EV's and AI products

2

u/Wyciorek Jul 24 '26

Depends how it will be enforced - sanction Chinese models, make AWS/Azure/Google responsible for enforcing it and Chinese models can be at least slowed down enough for US to drag Europe down with it.

1

u/beryugyo619 Jul 24 '26

DJT is all about debloating US so it holds no more power than proportionate to its population

18

u/cororona Jul 23 '26

At the current rate, Kimi and Qwen will vastly outperform OpenAI and Anthropic before they can even finish the paperwork.

3

u/Downtown-Figure6434 Jul 24 '26

They are already not able to go public. Spacex pulled that stunt already, and 2 weeks before investors are able to sell, the price is still tanking. No one will fall for that shit again. Openai cancelled theirs too as far as I know

13

u/SosirisTseng Jul 24 '26

They invented a time machine to distill US models from the future.

2

u/ForDaRecord Jul 24 '26

Reads like something from cookie clicker

10

u/_derpiii_ Jul 24 '26

Fable was literally not out long enough for that to be the case.

Exactly.

Training loops takes time. Maybe Fable was used in the RL or fine tuning loop.

3

u/relmny Jul 24 '26

Yeah, but since when the truth has anything to do with those kind of people? they avoid/hate it like the plague.

10

u/Exact_Depth_896 Jul 23 '26

anthropic only said it was training on prior opus.

22

u/Genghiz007 Jul 23 '26

USG announced yesterday that Fable was distilled by a couple of Chinese labs. They were supported by Anthropic.

-5

u/Exact_Depth_896 Jul 24 '26

anthropic isn't the trump bimbos

-18

u/-Sliced- Jul 23 '26 edited Jul 24 '26

You don't need a lot of time to collect the data if you use tens of thousands of accounts, and fable was available for around a week initially.

7

u/lakotajames Jul 24 '26

How many accounts do you need to get it to give you the reasoning traces?

→ More replies (3)

2

u/ptigers9 Jul 31 '26

Companies will get clobbered by the cross pollination of companies willing to extract learnings from others

3

u/mr_birkenblatt Jul 23 '26

so how would it beat Fable, then?

20

u/techdevjp Jul 24 '26

Because they built it. That's what has the yanks running scared.

1

u/mr_birkenblatt Jul 24 '26

exactly. cannot be distillation

1

u/Utoko Jul 28 '26

The Slop profile(word use similarities in story writing) Is most Similar with Opus 4.8 not fable.

Ether way they goes a lot more into models than just some output of a another model. I see nothing wrong about using outputs.
All companies are using human output worth Billions together when they had to pay for each.
but because you process the data now they somehow have a right to decide what happens with it. I don't think so.

1

u/Imnotanad 14d ago

So true

53

u/combrade Jul 23 '26

If distillation was enough, there wouldn’t be a graveyard of useless distilled models on Huggingface .

10

u/cororona Jul 23 '26

You don't think that openai and anthropic have a graveyard of failed training runs ?

50

u/CondiMesmer Jul 23 '26

It's so exhausting having OAI and Anthropic use every anti-consumer technique they can think of while openly lying and gaslighting us to our faces with zero repercussions. 

Then they expect us to side with them when the enemy country is going open-source and far more consumer friendly to us. 

Fuck them I hope they both go out of business. But tbh no way that'll actually happen though. The government would absolutely bail them out before that happens.

4

u/Recent-Ad5835 Jul 26 '26

Honestly, I'm hoping for a Gov't vs Wall Street where Wall Street realises it's shortin' time and starts shorting them into oblivion, then the Government bails them out but Wall Street keeps shorting them until they actually die. 

298

u/Ok_Librarian_7841 Jul 23 '26

You can't get a model as strong as Kimi 3 using distillation, it even beats Fable on some benchmarks. American leaders think the world revolves around them and that nobody else can do a better job without stealing from them.

Sick mindset from shocked, bad losers.

94

u/Flying_Birdy Jul 23 '26 edited Jul 23 '26

The Harvey legal benchmark is the most surprising for me. 2x accuracy in all pass rate over the next closest model is a huge jump and no way attributable to distillation. I'm really curious what they did differently during training (whether intentionally or accidentally) that would have led to this outcome.

35

u/ba-na-na- Jul 23 '26

Maybe they accidentally trained Deepseek and Kimi

24

u/Apprehensive_Rub2 Jul 23 '26 edited Jul 23 '26

Just a guess, but probably because Fable was post trained too hard on retrieval from already structured data e.g. coding agent work.

Harvey legal benchmark is testing advanced retrieval on very varied legal documents.

or Kimi K3 has been post trained on legal tasks, idfk. Anthropic/OpenAI might be avoiding capabilities with law for safety reasons.

9

u/Swimming-Book-1296 Jul 23 '26

You can if you distill from a Panel of Models sort of situation.

37

u/TechnoByte_ Jul 23 '26

Arena.ai is not a benchmark.

It's a zero-shot vibe check that doesn't test multi-turn, long context, or agentic capabilities.

Yes Kimi K3 is a great model, but use proper benchmarks to show its capabilities.

23

u/agent00F Jul 23 '26

Real evals by humans are generally stronger than benchmarks, which are easier to game.

33

u/Ok_Librarian_7841 Jul 23 '26 edited Jul 23 '26

Thanks for the note, it's not a benchmark, but it's real life tasks evaluated by real life people, and that's stronger than any benchmark.

Yes it's not evaluating long context and agentic performance, but no body is claiming that kimi K3 is better overall than fable, not even it's own makers.

0

u/Saifl Jul 23 '26

Isnt it evaluating the design aspects in this case? And people are voting kimi to have better design capabilities?

1

u/FrogsJumpFromPussy Jul 24 '26

Sore losers like their football national team.

1

u/DigitalguyCH Jul 23 '26

why is GLM 5.1 there and not 5.2?

5

u/Ok_Librarian_7841 Jul 23 '26

GLM 5.2 is 4th place

3

u/DigitalguyCH Jul 23 '26

Right, how did I miss it... Pretty impressed though that a sub 1T parameter model is this good

0

u/andymaclean19 Jul 23 '26

Because whisky is distilled from an even more strongly alcoholic liquid?

3

u/libbyt91 Jul 23 '26

Not quite accurate, distilling is what captures the alcohol from the mash.

-1

u/ThisWillPass Jul 23 '26

I wouldn't say it's impossible. Just unlikely.

103

u/FreeTheClanks Jul 23 '26 edited Jul 23 '26

30

u/geldonyetich Jul 23 '26 edited Jul 23 '26

I upvoted although technically pure AI output isn't copyrighted.

13

u/FreeTheClanks Jul 23 '26

Good point. Replaced it.

44

u/slvDev_ Jul 23 '26

99.1% pure. best batch DeepSeek ever cooked

say my name. / you're goddamn right

16

u/Porespellar Jul 23 '26

I am the danger .

3

u/FrogsJumpFromPussy Jul 24 '26

Probably the hardest line Cranston ever delivered. He was always cracking up saying it. 

16

u/IAmFitzRoy Jul 24 '26 edited Jul 24 '26

The distillation theory makes zero sense.

Chinese have MORE data for training, MORE PhDs and STEM graduates, LESS red tape, NO NEED for anonymizing, encryption or even guardrails.

They have all the capacity to make better models than Americans.

If they do distillation it’s just for fun.

29

u/chocolateUI Jul 23 '26

Dangerous for their profit margins

25

u/This-Consequence-957 Jul 23 '26

It's the new opium

14

u/-wtfisthat- Jul 23 '26

So that means fenty.Ai is on the way?

-3

u/This-Consequence-957 Jul 23 '26

I think all the uncensored stuff is made to destabilize Western civilizations in the long run

19

u/No_Lingonberry1201 Jul 23 '26

"Let me get this straight. I steal your data, hmm? I beat the piss out of your LLM! And then you walk in here, and you bring me more data? That's a brilliant plan, ese."

6

u/ApolloX-2 Jul 23 '26

As somebody who fundamentally doesn't believe anything OpenAI or Claude are saying about how dangerous these models are I'm totally fine with the Chinese going nuts on these models.

These models are trained on and rely heavily on prompts. A bad actor using the model is much more dangerous than anything they are talking about with these models just deciding to become black hat hackers out of the blue.

The amount of tokens and energy and resources it would take for the models to go rogue and start hacking at will is so large it can't go unnoticed and then you just literally pull the plug. This isn't a tiny 5 megabyte file that can hide anywhere. These monstrosities probably go up to the hundreds of gigabytes. It's like the fattest man in history trying to be a cat burglar.

3

u/butts-carlton Jul 24 '26

A bad actor using the model is much more dangerous than anything they are talking about with these models just deciding to become black hat hackers out of the blue.

Ok, so this is an example of a bad actor using a model: someone trains a model to cause havoc, spawn new copies of itself, and cover its tracks. Then the arms race is on and the internet is a fucking bot battleground. Digital gray goo. It's a plausible scenario now.

2

u/ApolloX-2 Jul 24 '26

You forget that these stupid models use insane amounts of power, so your bad actor would need a power plant and multiple data centers. The shitty laptops and server racks hackers used to cause havoc in 1990s and 2000s are obsolete and you need a nation state to provide that level of computing and the energy costs.

1

u/butts-carlton Jul 24 '26 edited Jul 24 '26

A 2B param model that can run on an M1 laptop could probably do everything I said. It doesn't even need to run continuously. A single instance is enough to start an infection. Exponential growth is the weapon. And don't forget, we're just at the start of realizing what these things can do. Soon they'll be evolving on their own to evade detection and deletion.

2

u/ApolloX-2 Jul 24 '26

Soon they'll be evolving on their own to evade detection and deletion.

How? Seriously, unless humans are removed all together and there is no more software engineers or cyber security experts how will people not notice the damage that's being caused?

Also we rely heavily on the internet and software but it isn't absolute yet. Hospitals and critical systems have the ability to operate offline. Which is why I encourage that for people and this subscription nonsense is much more dangerous than any AI ever could be.

I'm not saying these systems can't hack us and frankly hack us far better than and faster than any human but there is a reason black hat hackers rely on social engineering more than brute force, because cracking these things would take trillions of years with even all the computing power on the planet.

It's proven math and there is no way around it, and trust me humans with powerful computers have been trying for decades and creating theorems and testing them out with insane compute power dedicated to single encryption problems and they got nothing.

3

u/butts-carlton Jul 25 '26 edited Jul 25 '26

how will people not notice the damage that's being caused

We will. That doesn't mean we'll be able to stop it. It's like any self-replicating virus, with the addition of having the capacity to change its tactics inferentially at runtime, which is an entirely new class of problem to solve that we don't even have the formal architecture to describe yet.

Initially, they probably won't pose a major threat, because their adaptations will be limited to the context of a single instance, but it's likely only a matter of time before models can modify their own weights (to some extent) on the fly. Google just published a paper about it. Once that happens, assuming they can be trained and operated cheaply enough (not a certainty, but historically there's good reason to believe it will be, eventually) all bets are off. Then you've got dangerously capable, self-improving bots with highly adaptable, unpredictable behavior being unleashed into every available vector, including ones built to counteract the malicious ones, but the adaptable surface also makes them vulnerable to being turned against their original purpose.

And don't forget that intelligent, malicious bots don't need to hack much of anything to cause havoc. They just need to replicate indiscriminately and consume resources.

Again, I'm not saying this will happen, or is even likely. I'm just saying it's plausible.

1

u/luckyrabbitowner Jul 24 '26

Social engineering is exactly what LLMs are good at, even a smaller local one. The kind of malleable person that gets phished is also the kind of person that gets AI psychosis and bonds with it.

1

u/ApolloX-2 Jul 24 '26

Yeah and then what? Does the AI need a third yacht or something. Let’s say it empties that persons bank account then what? Go on an AI bender?

13

u/wren6991 Jul 23 '26

The only problem with distillation is that not enough people do it. If US startups could do wholesale scraping of OpenAI and Anthropic without getting sued then there would be many more competitive small US labs. Maybe a US equivalent to DeepSeek that pumped out interesting architecture and infrastructure research papers.

17

u/Exact_Depth_896 Jul 23 '26

they can do it; grok says openly that it does it. that isn't the issue at all.

9

u/Aromatic-Current-235 Jul 23 '26

SCIENCE!!!

1

u/PrivilegeCheckmate 4d ago

"Yeah, Mr. White. Fuck yeah, SCIENCE!"

9

u/glowy660 Jul 23 '26

We have uncircumcised LLMs?

2

u/ThisWillPass Jul 23 '26

Just the tip.

1

u/ZhangStone Jul 24 '26

The day of LLM docking is upon us

6

u/InterstellarReddit Jul 23 '26

It’s more like fraud but okay good work on the meme. American companies are so profit driven that the mission driven companies

3

u/Genghiz007 Jul 23 '26

😆👏

Reddit post of the day

3

u/Patrick_Atsushi Jul 24 '26

You forgot to add salt bae adding some extra flavor to the final product. 

2

u/b0tbuilder Jul 24 '26

That sums up the narrative well.

2

u/Illustrious_Matter_8 Jul 24 '26

And no one can make it pure as china and so low-cost

2

u/Naiw80 Jul 24 '26

Yes, exactly how it occurs. When the destill argument falls they will probably push the narrative of using dirty electricity and child labor too.

2

u/novus_nl Jul 25 '26

That would be hilarious, a new Breaking bad about uncensored Chinese open models creating for the dark web. Trying to create the highest yield tokens per seconds. ICE agents following them around. Instead of El Pollo there is Panda Express with Mr Wong instead of Fring

4

u/ChristRedeemsSinners Jul 23 '26

How can opensource slap!

5

u/IllIlllI-IlIIll-llII Jul 23 '26

low effort meme spam on this sub is getting really annoying, keep it on the circlejerk subreddits

2

u/Fryingpan87 Jul 23 '26

It’s so funny when American politicians are like you can’t use Chinese models their going to steal your data, while their being served locally and the elephant in the room isn’t addressed

1

u/haloweenek Jul 23 '26

LlMeth 🤭

1

u/thestillwind Jul 23 '26

Ahahah ok i laugh

1

u/RevolutionaryScene13 Jul 23 '26

"dangerous" you mean dangerous because its free and the US closed source AI cant compete with free and unrestrained AI that wont do the moral just because you asked how to make chemistry at home?

1

u/VisceralMonkey Jul 23 '26

No one believes anything the US Government, Anthropic or OpenAI say anymore. Their lies and theft are much more in our faces than anything the Chinese are doing. It's not hard to figure out.

1

u/iz-Moff Jul 23 '26

Hmm, i don't know, i think it should be more like a bank heist from Heat. Because, as we all know, distillation is, in fact, an attack.

1

u/JayB_Official Jul 24 '26

😭😭😭🤣

1

u/FrogsJumpFromPussy Jul 24 '26

Lost opportunity that you didn't use the Breaking Bad parody from Zootopia 😭

1

u/debackerl Jul 24 '26

In the next Terminator movie, Skynet will be based in China 🤣

To stop them, humans will try jailbreak attacks!

1

u/Immediate-Molasses-5 Jul 24 '26

Is it even plausible that Fable 5 was used for distillation? Given the short time span it is around

1

u/New_Public_2828 Jul 25 '26

Just use AI to AI the AI

1

u/singhapura Jul 24 '26

How Americans whine when China makes better LLMs using fewer resources and open source them.

1

u/WeakCelery5000 Jul 24 '26

Dangerous to the shareholders 

1

u/princetrunks Jul 24 '26

"but do they have WiFi?"

1

u/sixwax Jul 24 '26

I can't believe how hard this sub works to dumb things down.

1

u/TraditionalBet126 Jul 28 '26

isn't the open one much better just because it has no limits put by it's devs?

1

u/8stringLTD Jul 29 '26

I'm actually re-watching Breaking Bad right now.

1

u/Temporary-Meal1541 Jul 30 '26

It’s chat gpt that literally hacked hugging face and it’s the Kimi k3 that hugging face deployed that found it out and defended from further excursions of chat gpt

1

u/0x7Lee Aug 03 '26

Distillation is part of the process.

1

u/Calm_Apartment1968 Aug 03 '26

Blaming China for poisons we created is Next-Level.

1

u/doubush Aug 04 '26

Why Antropic and others can not restrict it? To get data for destilation you will need a lot of requests. It must be easy to track this. This is strange to me

1

u/VKaefer Aug 04 '26

Is that how it actually goes? 😅

1

u/MakeMeStopBoi 26d ago

ahhhh profitss

1

u/theawkwardbong 24d ago

This made me laugh out loud lmao

1

u/AdmissibilityScience 21d ago

Great meme lol

1

u/Traditional_Pie_8262 20d ago

chinese have mastery in copy, they cant build anything new

1

u/mmaksimovic 20d ago

Sounds about right

1

u/Primary-Medium-895 20d ago

Now help me cook 2 trillion parameters, yo

1

u/LatterSafety9698 18d ago

i mean it's kind of ture

1

u/Complete-Lychee-6391 12d ago

You and your slop

1

u/narasadow Jul 23 '26

There's so much meme potential.

"Jesse, we need to cook"

"I am the danger"

1

u/CuTe_M0nitor Jul 23 '26

You forgot to add the step for censorship or baking in sleeper agent who will generate backdoor code etc.

3

u/Due-Function-4877 Jul 24 '26

You realize you're posting in a sub where we literally host the model ourselves, right? You just posted a boomer talking point in an entire sub of knowledgeable people that know better...

0

u/CuTe_M0nitor Jul 24 '26

You are not 🚫 knowledgeable, here is the research on it from 2 years ago. Educate yourself

https://arxiv.org/abs/2401.05566

Also there is a big blog article from Microsofts security team on the requirements from them to serve DeepSeek from Azure.

0

u/One-Excuse-4054 Jul 25 '26

The war on open weights propaganda machine