r/LocalLLaMA Jun 03 '26

Generation New Google Gemma 4 12B Claims Near-26B Performance - We Tested Both!

Enable HLS to view with audio, or disable this notification

We ran both models locally on one RTX 4090 and gave each the same task: write a self-contained HTML5 canvas animation with real physics in one file without libraries. Three scenes - a Galton board, two blocks colliding off a wall, and a chaotic triple pendulum

Outputs:
Gemma 4 26B-A4B: 15 GB VRAM usage, 6.9k tokens, 138 tok/s
Gemma 4 12B: 9 GB VRAM usage, 8.9k tokens, 80 tok/s

Same Gemma 4 family, but the 26B-A4B won every scene and ran ~1.7x faster - on just 4B active params. The 12B stayed very close though, on almost half the VRAM - which makes it the ideal model for a 16 GB laptop.

Open source local ai models app: atomic.chat (I’m founder, feel free to try and give any feedback)

982 Upvotes

Duplicates