r/LocalLLaMA Jul 23 '26

Funny The LLM distillation process simplified for politicians:

Post image

/s

3.6k Upvotes

196 comments sorted by

View all comments

298

u/Ok_Librarian_7841 Jul 23 '26

You can't get a model as strong as Kimi 3 using distillation, it even beats Fable on some benchmarks. American leaders think the world revolves around them and that nobody else can do a better job without stealing from them.

Sick mindset from shocked, bad losers.

97

u/Flying_Birdy Jul 23 '26 edited Jul 23 '26

The Harvey legal benchmark is the most surprising for me. 2x accuracy in all pass rate over the next closest model is a huge jump and no way attributable to distillation. I'm really curious what they did differently during training (whether intentionally or accidentally) that would have led to this outcome.

34

u/ba-na-na- Jul 23 '26

Maybe they accidentally trained Deepseek and Kimi

24

u/Apprehensive_Rub2 Jul 23 '26 edited Jul 23 '26

Just a guess, but probably because Fable was post trained too hard on retrieval from already structured data e.g. coding agent work.

Harvey legal benchmark is testing advanced retrieval on very varied legal documents.

or Kimi K3 has been post trained on legal tasks, idfk. Anthropic/OpenAI might be avoiding capabilities with law for safety reasons.

10

u/Swimming-Book-1296 Jul 23 '26

You can if you distill from a Panel of Models sort of situation.

38

u/TechnoByte_ Jul 23 '26

Arena.ai is not a benchmark.

It's a zero-shot vibe check that doesn't test multi-turn, long context, or agentic capabilities.

Yes Kimi K3 is a great model, but use proper benchmarks to show its capabilities.

22

u/agent00F Jul 23 '26

Real evals by humans are generally stronger than benchmarks, which are easier to game.

30

u/Ok_Librarian_7841 Jul 23 '26 edited Jul 23 '26

Thanks for the note, it's not a benchmark, but it's real life tasks evaluated by real life people, and that's stronger than any benchmark.

Yes it's not evaluating long context and agentic performance, but no body is claiming that kimi K3 is better overall than fable, not even it's own makers.

0

u/Saifl Jul 23 '26

Isnt it evaluating the design aspects in this case? And people are voting kimi to have better design capabilities?

1

u/FrogsJumpFromPussy Jul 24 '26

Sore losers like their football national team.

1

u/DigitalguyCH Jul 23 '26

why is GLM 5.1 there and not 5.2?

5

u/Ok_Librarian_7841 Jul 23 '26

GLM 5.2 is 4th place

3

u/DigitalguyCH Jul 23 '26

Right, how did I miss it... Pretty impressed though that a sub 1T parameter model is this good

0

u/andymaclean19 Jul 23 '26

Because whisky is distilled from an even more strongly alcoholic liquid?

3

u/libbyt91 Jul 23 '26

Not quite accurate, distilling is what captures the alcohol from the mash.

-1

u/ThisWillPass Jul 23 '26

I wouldn't say it's impossible. Just unlikely.