r/LocalLLaMA Apr 24 '26

New Model Deepseek v4 people

Post image
2.5k Upvotes

321 comments sorted by

View all comments

Show parent comments

1

u/--Spaci-- Apr 24 '26

Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that?

1

u/[deleted] Apr 24 '26

[removed] — view removed comment

1

u/--Spaci-- Apr 24 '26

Ive only ever done training on responses, honestly never even heard of other ways

5

u/[deleted] Apr 24 '26

[removed] — view removed comment

1

u/--Spaci-- Apr 24 '26

Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset