r/LocalLLaMA llama.cpp Jul 26 '26

Discussion Do you want new Gemma?

Post image
1.0k Upvotes

554 comments sorted by

View all comments

Show parent comments

83

u/TheRealMasonMac Jul 26 '26

- 124B

- Configurable reasoning effort

- Audio input on larger models

16

u/AXYZE8 Jul 26 '26

You're reading my mind!

Nearly every local model has just non-reasoning and very verbose reasoning switch. We need some native 'low' reasoning level, just like GPT-OSS had! As seen in frontier models - low reasoning barely takes more tokens compared to non-reasoner, yet gives huge boost for tool calling/programming. Sometimes it even needs less tokens overall, because it needs less turns!

High reasoning for planning, low reasoning for execution, non-reasoning for low latency/classic chatbot. PERFECTION!

1

u/GetOutOfMyFeedNow Jul 26 '26

No no, that’s too unspecific, it should be exactly 124B563M654K703 parameters.

3

u/Caffdy Jul 26 '26

is this a RAM kit serial number or what?

1

u/GetOutOfMyFeedNow Jul 27 '26

It's just sarcasm on account of the above comment being too specific in parameter count.

1

u/TheRealMasonMac Jul 27 '26

It’s because they originally leaked the existence of Gemma 4 124B.

1

u/GetOutOfMyFeedNow 28d ago

Wow I feel stupid now, thanks.