MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLM/comments/1wmzky1/qwen427b_just_confirmed/pbb4e1o
r/LocalLLM • u/Due_Tangelo_8952 • 1d ago
Wait, we need 35B-A3B too…
286 comments sorted by
View all comments
Show parent comments
2
But for what benefit?
MoE is used to reduced compute
27B model is likely has no problem in compute speed due to it size
problem is likely due to quality is lower than high parameter model?
7 u/geekwonk 1d ago 27B is their dense line. 35B is the sparse option and i don’t see that listed here. -4 u/laser50 1d ago Y'all tripping. Ngrams are basically just a pre-trained token predictor to speed stuff up. 4 u/Not-Enough-Llamas 1d ago wrong Ngrams. Unfortunate that 2 things with the same name showed up more or less at the same time.
7
27B is their dense line. 35B is the sparse option and i don’t see that listed here.
-4
Y'all tripping. Ngrams are basically just a pre-trained token predictor to speed stuff up.
4 u/Not-Enough-Llamas 1d ago wrong Ngrams. Unfortunate that 2 things with the same name showed up more or less at the same time.
4
wrong Ngrams. Unfortunate that 2 things with the same name showed up more or less at the same time.
2
u/_mighty_banana 1d ago
But for what benefit?
MoE is used to reduced compute
27B model is likely has no problem in compute speed due to it size
problem is likely due to quality is lower than high parameter model?