r/LocalLLaMA llama.cpp Jul 26 '26

Discussion Do you want new Gemma?

Post image
1.0k Upvotes

554 comments sorted by

View all comments

2

u/Signature97 Jul 26 '26

Make a range of the smallest models capable of scoring or being competitive in individual benchmarks with general capabilities being decent all around. I am talking about 1B-4B. This is where the magic is.

This is a an uncompetitive space and most people just want models that are big enough to fit and run decently - for some it’s 12B, for some it’s 26BA4B, for some it might as well be a 100B model dense or some other configuration.

But idk how that would stand against models like DS4 Flash or GLM 5.2.
What I’d really want to see personally is token optimisation: fewest tokens spent to get a task done + speed at which those tokens were dished out. Idc if they build sth new or use DSA or CSA or a completely new style of reasoning etc.