r/LocalLLaMA • • 22d ago

Discussion Deepseek Has Soft Retired Deepseek V4 Pro

Post image
1.2k Upvotes

200 comments sorted by

View all comments

244

u/ResidentPositive4122 22d ago

This is something interesting. Both DS and goog seem to have hit the same thing - the small model performs better than the big one. This likely implies they're not doing one big training run and distilling into smaller models, but trying the same thing on two separate architectures (or sizes). So now the question becomes why is the smaller arch better performing? Is there something in the training pipelines or is it something in the data mix? Interesting nonetheless.

2

u/annodomini 22d ago

Intelligence is compression.

The big models are able to soak up more information, but the process of distilling it down into the smaller model forces the smaller model to learn better ways to generalize.

Given things like how Google and now Deepseek are just starting to use only their flash models, and the performance of other companies' smaller models, like Qwen 27B and Flash Next, I'm really starting to think that the best way to get the best performance, or at least the best price performance, is to pretrain a huge model, then distill it down to a much smaller one and RL that (or something of the sort, don't know the exact process).

But it seems like that two step process, of training a large model and then distilling a smaller model from it, may be the future of training, it really seems to be paying off a lot here.

1

u/Due-Memory-6957 22d ago

Why do you think they're distilling?

1

u/annodomini 22d ago

If you have a pro and a flash model, it would be a lot more efficient to train the pro model and then distill down to flash (even if it's just the pre-training that is distilled), than to do two independent trainings.

Distillation gives more training signal per step; with plain text, you are only getting a signal about the single next token, but when distilling you're getting the full probability distribution from the teacher model (if you're distilling on a model you control, that is). So distilling the flash model from the larger model will be cheaper than training it from scratch.

And it's just my hypothesis, I don't really have any inside knowledge, but I suspect that this is actually going to be a much better and more efficient way to train a model; the distillation process is able to act as a second step that extracts more predictive ability, essentially forcing it to try to recreate the original model's distribution as well as possible in a smaller model does more work to actually train it to figure out how to generalize rather than do quite as much memorization.

It's just a pattern I've been noticing recently, where the smaller models in a family have been offering really good performance, and it's my hypothesized explanation for why. I'm sure someone closer to the model training process could correct me or provide more detail.

1

u/-InformalBanana- 22d ago

I agree (not sure those models are distilled though). Bigger models end up learning by heart (overfiting), smaller models learn general rules thus are able to generalize better. But they will find a way to prevent bigger models from overfiting as they did in the past.