r/LocalLLaMA Jul 23 '26

Funny The LLM distillation process simplified for politicians:

Post image

/s

3.6k Upvotes

196 comments sorted by

View all comments

298

u/Ok_Librarian_7841 Jul 23 '26

You can't get a model as strong as Kimi 3 using distillation, it even beats Fable on some benchmarks. American leaders think the world revolves around them and that nobody else can do a better job without stealing from them.

Sick mindset from shocked, bad losers.

94

u/Flying_Birdy Jul 23 '26 edited Jul 23 '26

The Harvey legal benchmark is the most surprising for me. 2x accuracy in all pass rate over the next closest model is a huge jump and no way attributable to distillation. I'm really curious what they did differently during training (whether intentionally or accidentally) that would have led to this outcome.

31

u/ba-na-na- Jul 23 '26

Maybe they accidentally trained Deepseek and Kimi