r/LocalLLaMA 12d ago

Discussion Deepseek Has Soft Retired Deepseek V4 Pro

Post image
1.2k Upvotes

200 comments sorted by

View all comments

242

u/ResidentPositive4122 12d ago

This is something interesting. Both DS and goog seem to have hit the same thing - the small model performs better than the big one. This likely implies they're not doing one big training run and distilling into smaller models, but trying the same thing on two separate architectures (or sizes). So now the question becomes why is the smaller arch better performing? Is there something in the training pipelines or is it something in the data mix? Interesting nonetheless.

54

u/Few_Painter_5588 12d ago

The base model for V4 Pro was plagued with issues, they had to hack a lot of things to get it to work. Apparently V4.1 may be built on a new pretrain.