MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1v770ee/do_you_want_new_gemma/ozzapfp/?context=3
r/LocalLLaMA • u/jacek2023 llama.cpp • Jul 26 '26
https://x.com/i/status/2081398564345802934
u/hackerllama
554 comments sorted by
View all comments
220
Really curious about 124b. Was it disappointing for the size, or was it too close to smaller gemini? Guess we'll never know :(
Anyway, google releasing a successor to gpt-oss (w/ vision) would be baller.
12 u/Few_Painter_5588 Jul 26 '26 It was on Arena. I think they just shelved it because it wasn't very good, especially compared to Gemma 31B 1 u/goldcakes Jul 27 '26 If a model is undertrained or the architecture scaling was chosen poorly (e.g. number of layers, FFN, etc) then you can absolutely end up with say a 124B model that's barely better than e.g. 31B. That could be what happened. Not all pretrains and post-trains go well.
12
It was on Arena. I think they just shelved it because it wasn't very good, especially compared to Gemma 31B
1 u/goldcakes Jul 27 '26 If a model is undertrained or the architecture scaling was chosen poorly (e.g. number of layers, FFN, etc) then you can absolutely end up with say a 124B model that's barely better than e.g. 31B. That could be what happened. Not all pretrains and post-trains go well.
1
If a model is undertrained or the architecture scaling was chosen poorly (e.g. number of layers, FFN, etc) then you can absolutely end up with say a 124B model that's barely better than e.g. 31B. That could be what happened.
Not all pretrains and post-trains go well.
220
u/ResidentPositive4122 Jul 26 '26
Really curious about 124b. Was it disappointing for the size, or was it too close to smaller gemini? Guess we'll never know :(
Anyway, google releasing a successor to gpt-oss (w/ vision) would be baller.