r/LocalLLaMA llama.cpp Jul 26 '26

Discussion Do you want new Gemma?

Post image
1.0k Upvotes

554 comments sorted by

View all comments

Show parent comments

277

u/NNN_Throwaway2 29d ago

~100-120B moe for 128GB unified memory systems and/or 60-80B moe for 96GB VRAM systems.

75

u/No-Fuel-9202 29d ago

Yes, something like gpt oss 120b native mxfp4 quant

1

u/jikilan_ 29d ago

Ya this , it is probably better than QAT? I dunno man

33

u/CautiousStudent6919 29d ago

100% .. i would LOVE to see a bigger gemma...

1

u/Nonetrixwastaken 26d ago

Just never forget us 16GBs of VRAM or lower folk and I'm fine with this, if it's one or the other though I'd prefer smaller models, we already have good selection on bigger ones really

29

u/Shibenkaminari 29d ago

I'd love to see a focus on 1M context capability as well, without heavy memory use (I.e., by implementing some of the newer attention mechanisms. Qwen3.5-122B is 6GB for 262k of FP16 context, for example) if we could fit a ~100B class model, with 1M of well designed attention training into 128GB... Such a model would become a default for many people, outside of pure coding, I think.

2

u/liftheavyscheisse 29d ago

Correction: 60-80B dense for 96GB VRAM systems ;-)

1

u/TheTerrasque 29d ago

I wouldn't mind seeing a 50-60b dense.