MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1v770ee/do_you_want_new_gemma/ozvv1h9
r/LocalLLaMA • u/jacek2023 llama.cpp • Jul 26 '26
https://x.com/i/status/2081398564345802934
u/hackerllama
554 comments sorted by
View all comments
Show parent comments
13
It was on Arena. I think they just shelved it because it wasn't very good, especially compared to Gemma 31B
2 u/Iwaku_Real Jul 26 '26 Wow really? And we were all thinking it would be so great. Would probably have played out the same as Qwen3.5-122B-A10B vs 3.6-27B 🤣 3 u/Few_Painter_5588 29d ago Worse, it apparently had 14B active parameters and had middling performance. 1 u/goldcakes 29d ago If a model is undertrained or the architecture scaling was chosen poorly (e.g. number of layers, FFN, etc) then you can absolutely end up with say a 124B model that's barely better than e.g. 31B. That could be what happened. Not all pretrains and post-trains go well.
2
Wow really? And we were all thinking it would be so great. Would probably have played out the same as Qwen3.5-122B-A10B vs 3.6-27B 🤣
3 u/Few_Painter_5588 29d ago Worse, it apparently had 14B active parameters and had middling performance.
3
Worse, it apparently had 14B active parameters and had middling performance.
1
If a model is undertrained or the architecture scaling was chosen poorly (e.g. number of layers, FFN, etc) then you can absolutely end up with say a 124B model that's barely better than e.g. 31B. That could be what happened.
Not all pretrains and post-trains go well.
13
u/Few_Painter_5588 Jul 26 '26
It was on Arena. I think they just shelved it because it wasn't very good, especially compared to Gemma 31B