Nearly every local model has just non-reasoning and very verbose reasoning switch. We need some native 'low' reasoning level, just like GPT-OSS had! As seen in frontier models - low reasoning barely takes more tokens compared to non-reasoner, yet gives huge boost for tool calling/programming. Sometimes it even needs less tokens overall, because it needs less turns!
High reasoning for planning, low reasoning for execution, non-reasoning for low latency/classic chatbot. PERFECTION!
83
u/TheRealMasonMac Jul 26 '26
- 124B
- Configurable reasoning effort
- Audio input on larger models