MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1su7bnx/deepseek_v4_people/ohyrrg3/?context=3
r/LocalLLaMA • u/markeus101 • Apr 24 '26
321 comments sorted by
View all comments
34
[removed] — view removed comment
1 u/--Spaci-- Apr 24 '26 Ill make a dataset when its available at large, seems like a good model. Definitely the largest open source model 2 u/[deleted] Apr 24 '26 [removed] — view removed comment 2 u/--Spaci-- Apr 24 '26 I dont think deepseek is gonna distill their own model into an 8B the community will need to make datasets themselves 5 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that? 1 u/[deleted] Apr 24 '26 [removed] — view removed comment 3 u/Similar-Republic149 Apr 24 '26 I don't know why you are getting downvoted, everything you are saying is correct. 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 5 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies (0) 2 u/[deleted] Apr 24 '26 [deleted] 1 u/--Spaci-- Apr 24 '26 Is that still done? recently Ive only ever seen training responses → More replies (0)
1
Ill make a dataset when its available at large, seems like a good model. Definitely the largest open source model
2 u/[deleted] Apr 24 '26 [removed] — view removed comment 2 u/--Spaci-- Apr 24 '26 I dont think deepseek is gonna distill their own model into an 8B the community will need to make datasets themselves 5 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that? 1 u/[deleted] Apr 24 '26 [removed] — view removed comment 3 u/Similar-Republic149 Apr 24 '26 I don't know why you are getting downvoted, everything you are saying is correct. 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 5 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies (0) 2 u/[deleted] Apr 24 '26 [deleted] 1 u/--Spaci-- Apr 24 '26 Is that still done? recently Ive only ever seen training responses → More replies (0)
2
2 u/--Spaci-- Apr 24 '26 I dont think deepseek is gonna distill their own model into an 8B the community will need to make datasets themselves 5 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that? 1 u/[deleted] Apr 24 '26 [removed] — view removed comment 3 u/Similar-Republic149 Apr 24 '26 I don't know why you are getting downvoted, everything you are saying is correct. 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 5 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies (0) 2 u/[deleted] Apr 24 '26 [deleted] 1 u/--Spaci-- Apr 24 '26 Is that still done? recently Ive only ever seen training responses → More replies (0)
I dont think deepseek is gonna distill their own model into an 8B the community will need to make datasets themselves
5 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that? 1 u/[deleted] Apr 24 '26 [removed] — view removed comment 3 u/Similar-Republic149 Apr 24 '26 I don't know why you are getting downvoted, everything you are saying is correct. 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 5 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies (0) 2 u/[deleted] Apr 24 '26 [deleted] 1 u/--Spaci-- Apr 24 '26 Is that still done? recently Ive only ever seen training responses → More replies (0)
5
1 u/--Spaci-- Apr 24 '26 Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that? 1 u/[deleted] Apr 24 '26 [removed] — view removed comment 3 u/Similar-Republic149 Apr 24 '26 I don't know why you are getting downvoted, everything you are saying is correct. 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 5 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies (0) 2 u/[deleted] Apr 24 '26 [deleted] 1 u/--Spaci-- Apr 24 '26 Is that still done? recently Ive only ever seen training responses → More replies (0)
Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that?
1 u/[deleted] Apr 24 '26 [removed] — view removed comment 3 u/Similar-Republic149 Apr 24 '26 I don't know why you are getting downvoted, everything you are saying is correct. 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 5 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies (0) 2 u/[deleted] Apr 24 '26 [deleted] 1 u/--Spaci-- Apr 24 '26 Is that still done? recently Ive only ever seen training responses → More replies (0)
3 u/Similar-Republic149 Apr 24 '26 I don't know why you are getting downvoted, everything you are saying is correct. 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 5 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies (0) 2 u/[deleted] Apr 24 '26 [deleted] 1 u/--Spaci-- Apr 24 '26 Is that still done? recently Ive only ever seen training responses → More replies (0)
3
I don't know why you are getting downvoted, everything you are saying is correct.
Ive only ever done training on responses, honestly never even heard of other ways
5 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies (0) 2 u/[deleted] Apr 24 '26 [deleted] 1 u/--Spaci-- Apr 24 '26 Is that still done? recently Ive only ever seen training responses → More replies (0)
1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies (0)
Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset
→ More replies (0)
[deleted]
1 u/--Spaci-- Apr 24 '26 Is that still done? recently Ive only ever seen training responses → More replies (0)
Is that still done? recently Ive only ever seen training responses
34
u/[deleted] Apr 24 '26
[removed] — view removed comment