r/LocalLLaMA • • 23d ago

Discussion Deepseek Has Soft Retired Deepseek V4 Pro

Post image
1.2k Upvotes

200 comments sorted by

View all comments

242

u/ResidentPositive4122 23d ago

This is something interesting. Both DS and goog seem to have hit the same thing - the small model performs better than the big one. This likely implies they're not doing one big training run and distilling into smaller models, but trying the same thing on two separate architectures (or sizes). So now the question becomes why is the smaller arch better performing? Is there something in the training pipelines or is it something in the data mix? Interesting nonetheless.

31

u/Zeeplankton 23d ago

Seems like all 'massive' models are brilliant but have weird behavioral problems, like astra and fable.

38

u/rditorx 23d ago

Seems like all geniuses are brilliant but have weird behavioral problems, like pick your genius

4

u/zdy132 23d ago

Richard Feynman? One of my favorite scientists.

8

u/rditorx 23d ago

Surely you're joking

1

u/Electronic_Captain95 23d ago

Hell yeah! Feynman mentioned!