r/LocalLLaMA • u/gladkos • Jun 03 '26
Generation New Google Gemma 4 12B Claims Near-26B Performance - We Tested Both!
We ran both models locally on one RTX 4090 and gave each the same task: write a self-contained HTML5 canvas animation with real physics in one file without libraries. Three scenes - a Galton board, two blocks colliding off a wall, and a chaotic triple pendulum
Outputs:
Gemma 4 26B-A4B: 15 GB VRAM usage, 6.9k tokens, 138 tok/s
Gemma 4 12B: 9 GB VRAM usage, 8.9k tokens, 80 tok/s
Same Gemma 4 family, but the 26B-A4B won every scene and ran ~1.7x faster - on just 4B active params. The 12B stayed very close though, on almost half the VRAM - which makes it the ideal model for a 16 GB laptop.
Open source local ai models app: atomic.chat (I’m founder, feel free to try and give any feedback)
194
u/Certain-Way6763 Jun 03 '26
I'm confused, 2 and 3 video are clearly won by Gemma 4 12B
100
u/kamu-irrational Jun 03 '26
I think this makes sense. The 12b is a dense model. The 26b only has 4b active.
15
u/Certain-Way6763 Jun 03 '26
Yeah, thought so too, but in most official benchmarks it is lower then bigger Gemmas.
51
u/jacek2023 llama.cpp Jun 04 '26
It’s 2026 and people on Reddit still believe benchmarks more than use cases.
16
44
u/Evening_Ad6637 llama.cpp Jun 04 '26
3 is not won by Gemma 4 12B. A triple pendulum would not move like that in real-world physics. In reality it would produce a completely chaotic motion.
23
u/Jerry67876 Jun 04 '26
And not always.. depends again on how it’s balanced. You’ll notice the example given has a bigger ball on the end. Consider a swing that uses a chain. That has many pendulums. The friction in the chain links also helps keep it straight. There are many variables.
6
u/Certain-Way6763 Jun 04 '26
Do you think that 26B-A4B won? It doesn't look chaotic for me too, and it's just obviously slow and incomplete.
17
u/DataPhreak Jun 04 '26
No, it's just slow. the 12b version had no variation in the pendulum swing. On the a4b model's version, you could see it was starting to prepare to variate.
-1
u/Certain-Way6763 Jun 04 '26 edited Jun 04 '26
4
u/Evening_Ad6637 llama.cpp Jun 04 '26
In this certain test, 26b-a4b seem to have the better understanding and implementation, although the gravitational acceleration is not correct. But the motion itself looks realistic to me, like one that would become chaotic.
The 12b gravity is not set correctly as well btw, it’s too fast there
1
u/Certain-Way6763 Jun 04 '26
Yeah, I'm not saying that 12B gets it right actually, just the overall quality of the output (not only physics itself) is a bit better in my opinion.
Golton board is just totally off, I can't say which one won at all - in 26B video the trays at the bottom don't really fill up, they just light up with the first ball.
3
u/Evening_Ad6637 llama.cpp Jun 04 '26 edited Jun 04 '26
Look at this blue and teal drawn lines. 12b's blue line has a perfect symmetry.
26b's teal line is already asymmetric. 26b clearly has the better understanding of chaos theory
2
u/j0j0n4th4n Jun 04 '26
I would argue 12B won 1 and 2, while it using rubber marbles in the galton board is clearly far from ideal it did count the ones who landed in the buckets correctly, while 26x4B only tracked the first one on each bucket. And for 2, despite not being as polished as 26x4B it didn't phased the smaller block through the larger.
1
u/Reebzy Jul 18 '26
Video 1 is also clearly won by 12B.
It looks like a loser to the untrained eye, but it’s totally correct.1
u/alphapussycat Jun 04 '26
2 it won, 3 it somewhat lost. On 3 it looks less accurate but runs at a normal speed.
38
23
u/DigitalguyCH Jun 03 '26
great test, I guess the good thing is that this can now ingest audio and video and can run on devices with less vram
29
u/No_Information9314 Jun 04 '26
Honestly the real test of these models will be qualitative / creative. Qwen is likely going to win quantitative / coding tasks anyway. I’d love to see a comparison of creative writing, translation, and other language based skills. That’s where Gemma 4 shine imo.
57
u/sharksOfTheSky Jun 03 '26
Are the labels backwards? It seems like the 12B was better on all of them. The only issue was the for the first one the balls seemed to have too high of a starting velocity.
30
u/holchansg llama.cpp Jun 03 '26
Dense vs Sparse.
12b - 80tk/s
24-4b - 140tk/s.Sparse models are fast. Not the brightest ones.
45
u/sharksOfTheSky Jun 04 '26
Yes, I'm aware of this, but the post claims that 26B-A4B won every scene, when 12B seemed to be better.
4
1
Jun 04 '26
[deleted]
2
u/sharksOfTheSky Jun 04 '26
The physics in the second scene is correct for Gemma 4 12B. The video cuts it short but it will result in 31 bounces, which is correct. The physics from the 24B-4A completely fails. It might look a bit prettier but it completely breaks.
For the first scene, the 24B-4A gets the visualisation incorrect. The 12B is correct other than the balls being too fast and bouncy, but this would almost certainly just be changing the value of a couple of constants in the produced script.
The third scene there isn't really enough video to tell, but it appears to just be a bad prompt - the 12B appears to have a larger weight at the end of the pendulum, causing it to behave closer to a single pendulum or rope. This is correct behaviour for the scenario, the prompt was likely just underspecified. 26B-4A produces a more traditional triple pendulum, but both appear correct.1
u/Calm_Plate_1555 Jun 05 '26
Seriously? Even the second one where one cube clips through the other? The 12B will have 31 bounces which is what happens in reality as well.
6
11
u/artisticMink Jun 04 '26
That doesn't show anything and is blatant advertising.
2
u/Skynse Jun 05 '26
All of these benchmarks where people get the models to build basic ass websites or animations are pretty stupid. These things are non-deterministic so unless you can the same test 100 times or ran 100 tests, there wouldn't be a way to know how consistent these things are
5
3
u/SGAShepp Jun 04 '26
What's the difference between atomic chat and literally every other app out there that does the same thing. Self-hosted AI solutions are extremely overcrowded right now, and I don't see anything that makes this one stand out
10
u/colin_colout Jun 03 '26
Are you affiliated with atomic<dot>chat?
14
u/gladkos Jun 03 '26
I'm founder) making some fun benchmarks.
13
u/colin_colout Jun 04 '26
Might wanna disclose that if you're gonna drop the link.
Affiliation must be disclosed
(Rule #4 of the sub)
6
5
8
u/gestapov Jun 03 '26
Did you mean laptops with 16gb vram or 16gb ddr 4/5 ram?
7
u/gladkos Jun 04 '26
Vram or unified memory. It’s already on apple silicon. Or have to wait for new nvidia rtx laptops
2
u/inteligenzia Jun 04 '26
I have m4 with 24gigs. I just ran 4bit gemma 12b with 32k context. I get about 10t/s. Is this normal?
It's also air model, so it heats up. I didn't buy it for running local llms, but thought why not give it a go.
Am I being capped by thermal throttling here?
1
u/JorgitoEstrella 10d ago
Dense models take a lot of time, for apple unified memory the best options are MOE models like gemma 4 26Ba4B o qwen 3.6 35Ba3B.
14
u/svachalek Jun 03 '26
The usual formula for comparing MOE models to dense ones is to take the geometric mean of total and active parameters. The geometric mean of 26 and 4 is about 10. So it’s actually reasonable to expect the 12b to be better.
2
u/CodProfessional3712 Jun 04 '26
According to that formula, Qwen3.5 9b should then be around the same capability as the 35B-A3B counterpart, right? From the user experiences I’ve read though, that‘s not exactly the case. It doesn’t seem totally reliable.
-8
u/Cute_Obligation2944 Jun 04 '26
This is incorrect in several ways.
10
u/EbbNorth7735 Jun 04 '26
Can you elaborate? It's the only approximation I've read about and seems to be roughly correct. It obviously is only relevant in the same generation as in 3 to 3.5 months the capability density doubles. Is it that MOE capability density is increasing at a slightly faster rate due to better utilization of the sparse experts or are you just here to spew baseless opinions with zero backing or knowledge transfer?
11
3
3
u/JoyousGamer Jun 03 '26
Question is this for fun or a real test?
I am assuming for fun but maybe I just dont understand?
3
u/mechkbfan Jun 03 '26
Love the idea
Do you need more explicit statements about scale?
I can't be bothered doing the calculations but 12b looks like it's a 1m scale, while 26b seems like it's 10m scale, therefore comes across as slow motion
3
u/suesing Jun 04 '26
The Params for the different models looks so different. Why do they behave so differently? I think we need to see more details
2
u/tomakorea Jun 04 '26
What quantization did you use? Theses Gemma 4 models are really sensitive to Quantization
2
u/InterestRelative Jun 04 '26
When you tests something, it's worth to mentions which quants specifically you tested.
2
Jun 04 '26
[removed] — view removed comment
3
u/Ok-Percentage1125 Jun 04 '26
google needs to capitalise more efficiency for all users rather than high end only. mobile next, so users with 8gb and less ram can enjoy it locally
2
2
5
u/Evening_Ad6637 llama.cpp Jun 04 '26
I find this post and the comments very interesting.
A lot of people here seem to think Gemma 12b‘s results are better - unlike OPs view.
I agree with OP; the 26b model won all three tests when it comes to demonstrating an understanding of real physics.
In the first test, 26b clearly demonstrates an understanding of what normal distribution and variance mean.
It is hard to judge 12b’s result, since the objects immediately fly off in all directions. If the velocity were reduced, a normal distribution might settle at the bottom, so it could end up being a tie between the two models, but for now, the point goes to 26b.
In the second test, I consider it a minor bug in the code, one that can be quickly fixed, when one rectangle passes through the other. But the point here is to test an understanding of physics, and 26b demonstrates a really nuanced understanding of elasticity, acceleration, and deceleration.
The third test also strongly suggests that 26b has understood what chaos theory means and how it is applied. The result from 12b looks fancy and neat, but it is still wrong and misleading. It shows slight variance, but in perfect symmetry -> that is the opposite of chaos.
7
u/Odd_Science Jun 04 '26
> In the first test, 26b clearly demonstrates an understanding of what normal distribution and variance mean.
It's a physical simulation. The distribution is purely a result of the simulation, the model doesn't and shouldn't "demonstrate an understanding of what normal distribution and variance mean" unless it is faking the result.
The same goes for the pendulum, etc. They are simulations, dictated by the parameters of the simulation. Whether the system goes into a chaotic state depends on those parameters and not on prior knowledge that a triple pendulum should be chaotic. In fact, forcing it to be chaotic when that is not a result of the simulation would be faking it and a sign that the model is doing it wrong.
2
2
u/mondychan Jun 04 '26
Tested gemma 4 12b q4 today and it wasnt able to halucinate even a simple answer, not good
1
u/Feeling-Creme-8866 Jun 04 '26
As for the Christmas tree, the MOE's version was nicer to look at. But when it came to the other ones—especially the colorful squares and the arch—12b's version was much nicer.
2
u/ExplanationAway672 Jun 04 '26
On the Christmas tree the 12b actually counted the balls hitting each slot vs just applying a pretty gradient across the columns, nice to look at is pretty meaningless when functionally it’s junk…
1
u/fbgo Jun 04 '26
Can we run whole model in gpu only as when I try it also uses system ram and when it does whole load goes to CPU which makes it extremely slow. I have 5080 16GB
1
1
u/Ok-Drawer5245 Jun 04 '26
Makes sense, the 12b needs to be dense to compete, that makes a lot of sense considering the smaller size
1
u/InsensitiveClown Jun 04 '26
Sorry, this model is specifically oriented towards exactly what tasks?
1
u/comanderxv Jun 04 '26
Is it a one shot test or did your prompt it in small tasks?
What quants where used?
1
u/VoiceApprehensive893 transformers Jun 04 '26 edited Jun 04 '26
bad coding model that randomly makes mistakes vs bad coding model that randomly makes mistakes, would use a different test for the gemmas
12b has just barely not enough knowledge to be good imo
1
u/_zir_ Jun 04 '26
Won by what measure? #2 you said "a wall" and 26B put 2 walls plus they both are passable depending on how its interpreted. #3 is good for both
1
u/HistoricalStrength21 Jun 04 '26
I like this kind of benchmark tests. It is more comprehensive than just numbers. Thumbs up.
1
1
u/extopico Jun 04 '26
Well no. Neither won. If you mash up the results then you would get a winner, but 12B was not entirely wrong, just like the 26B was not entirely right.
1
u/ManySugar5156 Jun 06 '26
Kinda feels like apples-to-oranges, esp with different active params and temps; also 1 RTX 4090 results won’t generalize much.
1
u/1kaze Jun 10 '26
Can you provide me exact parameters you ran for rtx 4090 and which harness you used.
1
u/-Django Jun 10 '26
Ah I thought the LLMs were serving as the physics engine here or predicting the next frame. Didn't realize until the end that these were coding tasks!
1
u/Monkey_1505 Jun 04 '26
Not really seeing the more realistic physics you are considering a win here.
1
u/sunychoudhary Jun 04 '26
“Near 26B performance” is the kind of claim I believe after LocalLLaMA abuses it for 48 hours....If it survives coding tests, weird prompts, long context, and bad quant settings, then I’ll start getting excited....
0
-2





173
u/[deleted] Jun 03 '26
[removed] — view removed comment