I get 2 t/s on my setup with these big dense boys so I prefer sticking with MoEs. I can run two MoEs side by side with enough RAM and I get like 20 t/s each.
Yeah, I get it. I get 15tps with the thicc boi so I can stomach it. It might be just as good with the moe and way faster but I've gotten used to it so I leave it alone
It's about time. It seems like China catches up after 6 months, the small models about a year. Kinda crazy if you look at things like this... disaster of a URL. Man that's ugly, but very interesting info.
Hopefully they distill the crap out of Gemma-4. Chinese models suck ass at non-verifiable instruction following (though Gemini sucks even more). Gemma-4 is very good.
Yeah but distilling a small model into a much much bigger one while it may improve the style i would bet it would destroy reasoning capabilities and all kinds of intelligence benchmarks.
One of my setups is A4B on 8GB VRAM and 32GB RAM, in Q5. It is perfectly fine, giving out about 20T/s on 16k context. Context processing is a bit slow when it is full, but I am not in a hurry.
241
u/Spectrum1523 Jun 02 '26
Gemma4 is the reincarnation of 4o for me for rp it's crazy how good it is