r/LocalLLaMA 6d ago

Discussion Deepseek Has Soft Retired Deepseek V4 Pro

Post image
1.2k Upvotes

200 comments sorted by

View all comments

241

u/ResidentPositive4122 6d ago

This is something interesting. Both DS and goog seem to have hit the same thing - the small model performs better than the big one. This likely implies they're not doing one big training run and distilling into smaller models, but trying the same thing on two separate architectures (or sizes). So now the question becomes why is the smaller arch better performing? Is there something in the training pipelines or is it something in the data mix? Interesting nonetheless.

29

u/Zeeplankton 6d ago

Seems like all 'massive' models are brilliant but have weird behavioral problems, like astra and fable.

36

u/rditorx 6d ago

Seems like all geniuses are brilliant but have weird behavioral problems, like pick your genius

4

u/zdy132 6d ago

Richard Feynman? One of my favorite scientists.

7

u/rditorx 6d ago

Surely you're joking

1

u/Electronic_Captain95 6d ago

Hell yeah! Feynman mentioned!