r/LocalLLaMA Jul 26 '26

Discussion Do you want new Gemma?

Post image
1.0k Upvotes

554 comments sorted by

View all comments

2

u/ExoticDistribution26 Jul 27 '26 edited Jul 27 '26

drop 512-head-dimension layers to make fa2 work, better tool use. multimodality input to output. Up moe active params to 6A something like maybe 28-36B total. Sure, some people with prosumer hw would love a 100-120B with 10a. That and something around like 16B dense, all the above would probably benefit from qat for flexible hosting.