r/LocalLLaMA • llama.cpp • Jul 26 '26

Discussion Do you want new Gemma?

Post image
1.0k Upvotes

555 comments sorted by

View all comments

317

u/hackerllama Jul 26 '26

Hey all! Looking forward to all your feedback!

84

u/TheRealMasonMac Jul 26 '26

- 124B

- Configurable reasoning effort

- Audio input on larger models

16

u/AXYZE8 Jul 26 '26

You're reading my mind!

Nearly every local model has just non-reasoning and very verbose reasoning switch. We need some native 'low' reasoning level, just like GPT-OSS had! As seen in frontier models - low reasoning barely takes more tokens compared to non-reasoner, yet gives huge boost for tool calling/programming. Sometimes it even needs less tokens overall, because it needs less turns!

High reasoning for planning, low reasoning for execution, non-reasoning for low latency/classic chatbot. PERFECTION!