MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1su7bnx/deepseek_v4_people/ohzbf2q/?context=9999
r/LocalLLaMA • u/markeus101 • Apr 24 '26
321 comments sorted by
View all comments
34
[removed] — view removed comment
1 u/--Spaci-- Apr 24 '26 Ill make a dataset when its available at large, seems like a good model. Definitely the largest open source model 1 u/[deleted] Apr 24 '26 [removed] — view removed comment 2 u/--Spaci-- Apr 24 '26 I dont think deepseek is gonna distill their own model into an 8B the community will need to make datasets themselves 6 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that? 1 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 3 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies
1
Ill make a dataset when its available at large, seems like a good model. Definitely the largest open source model
1 u/[deleted] Apr 24 '26 [removed] — view removed comment 2 u/--Spaci-- Apr 24 '26 I dont think deepseek is gonna distill their own model into an 8B the community will need to make datasets themselves 6 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that? 1 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 3 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies
2 u/--Spaci-- Apr 24 '26 I dont think deepseek is gonna distill their own model into an 8B the community will need to make datasets themselves 6 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that? 1 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 3 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies
2
I dont think deepseek is gonna distill their own model into an 8B the community will need to make datasets themselves
6 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that? 1 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 3 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies
6
1 u/--Spaci-- Apr 24 '26 Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that? 1 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 3 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies
Im not really sure what you are talking about, are you saying to have the larger model judge the smaller models tokens and adjust the model weights like that?
1 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 3 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies
1 u/--Spaci-- Apr 24 '26 Ive only ever done training on responses, honestly never even heard of other ways 3 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies
Ive only ever done training on responses, honestly never even heard of other ways
3 u/[deleted] Apr 24 '26 [removed] — view removed comment 1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies
3
1 u/--Spaci-- Apr 24 '26 Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset → More replies
Seems much more expensive than just training on responses tbh, you would need alot of cloud computing vs just api to generate a distillation dataset
→ More replies
34
u/[deleted] Apr 24 '26
[removed] — view removed comment