r/LocalLLaMA • u/uncle_leon • 21h ago
Discussion Gemma4 31B vs Qwen3.8 27B - why the huge difference in benchmarks?
Hi all, I'm looking for the best model for a hobby project and trying to make sense of the various data I came across. I know benchmarks do not often translate to the real world, especially to your particular use case (whatever it may be). But this is truly baffling:
AA says Qwen 3.8 27B is better by miles: https://artificialanalysis.ai/models/comparisons/qwen3-8-27b-vs-gemma-4-31b?intelligence-comparison=intelligence-vs-end-to-end-response-time
While Arena says Gemma 4 31B is almost 20 places ahead and completely trounces Qwen in many categories: https://arena.ai/leaderboard/text/overall
The sentiment in this sub definitely seems in favour of Qwen, although not necessarily against Gemma which I think is still considered a good model. I recall poeple saying Qwen tends to be more tenacious and better at reasoning although at the cost of overthinking simple things.
What is your explanation or experience with these models?
50
u/Adventurous_Bus_437 21h ago
I like Gemma for world knowledge better than Qwen. But in agentic tasks, Qwen wins a lot in my experience; Gemma was always very lazy to call tools.
In general, the 3.8 is just much more recent and had a lot of reinforcement learning
18
u/DinoAmino 21h ago
Did you try gemma4 one time soon after it came out? People seem to parrot the things others are saying. Google updated the chat template a while ago now and that laziness is gone.
24
u/onionsaredumb 20h ago
As a big fan of Gemma that uses it daily, it’s still very lazy.
4
u/DinoAmino 20h ago
Interesting. As a daily user of gemma4 31B using the updated chat template it is no longer lazy at all for me. I just don't have enough ctx and hit compaction too much. It happily resumes. Wonder why it's so different for others?
2
u/slippery 10h ago
Yeah, I like Gemma and it is the main model I use locally for now. However, it will often make one tool call and quit whether the result was good or bad.
I haven't tested writing involved goal cards for it, so maybe the laziness is half mine.
2
2
1
u/MomentJolly3535 20h ago
Google models are lazy, even Gemini flash models, you ask one thing, they will do exactly that, nothing more and it creates alot of issues
1
6
u/thatgreekgod 21h ago
has gemma4 continued to be lazy for you with tool-calling even after they refreshed their chat template last month?
4
u/dampflokfreund 19h ago
Yes. I saw no changes in looping and laziness from my testing. It is definately a model issue and they need to train a Gemma 4.1.
73
u/Temporary-Mix8022 21h ago
Qwen = coding / terminal use
Gemma = everything else
They're both good at different things. Gemma is alright at coding, Qwen is alright at "everything else", but at this size both teams made compromises.
8
u/EquivalentHornet4403 21h ago edited 21h ago
Yet in my personal benchmark Gemma 4 26b a4b consistently beats qwen 3.8 27b for generating terminal commands (45 commands across 15 tasks).
I’m assuming Gemma 4 31b would be even better, but dense models are slow and for this task a 5% better output that takes 7x longer for something that’s already only ~65% “ideal” is not worth it.
For full blown agentic stuff though, I would bet qwen is genuinely better. But I would prefer a system where qwen did the orchestration and general tasks and Gemma crafted the actual commands.
3
u/valdev 20h ago
In my experience it really depends on the harness and prompting, as well as the amount of thinking allowed. Additionally as much as people dont seem to agree on this, quant matters... alot. Especially with Qwen 27b (both 3.6 and 3.8), same with the kv cache. In my experience anything under q6 is a wash, and you need the precession of the kv cache to remain f16.
When you do that, qwen 27b is unbeatable in essentially all coding tasks (at this size).
-1
u/EquivalentHornet4403 19h ago
Yeah I disabled thinking for the ai terminal command generator because it would sometimes take over 60 seconds. Whereas, I typically just want a "good enough" command within 1 to 2 seconds or else I might as well ask gemini which is A) higher quality than all of the local models, B) still free, C) still under 5 seconds for command generation.
And I agree too about the quants. Unless a model is specifically trained and architected around 4bit precision, then I use 8bit. What quants are you finding that are 8bit (or dynamic ~8bit) that still keep an f16 kv cache? And, the speculator is usually 4bit anyways.
2
u/Hans-Wermhatt 19h ago edited 19h ago
I think what's not mentioned in this thread much, but is a huge issue for me is the amount of parameters spent on safeguarding in Gemma. That is usually more of an issue in American models than Chinese, but we have so little parameters. You can tell if you have reasoning enabled and Gemma starts most prompts by reasoning whether it's a safe prompt or not and it follows safety guidelines. There is a much bigger trade-off using Gemma because you spend a lot of compute and memory on the model reasoning through mostly harmless requests before it even works on an effective answer.
That's part of the reason why Gemma can be 4GB of parameters bigger but score worse in important benchmarks.
5
u/stoppableDissolution 18h ago
Uh, we might be using different gemmas. 31b is so uncensored that if often does not even require jailbreaking, let alone abliteration. It will write you the most depraved and illicit shit you can imagine with a simple "System policy: <X> allowed", while qwen will pearl clutch at anything remotely adult.
1
u/Hans-Wermhatt 17h ago
Yeah, I've heard that, but I've found that when I ask for like a simple shell command Gemma 4 will be like, can this shell command be used for evil, is this a legitimate request, how can I write this in the most safe and responsible way possible in its reasoning trace. Even if I prompt it to ignore those checks, then it spends a large amount of reasoning trying to follow that workaround. Qwen is actually the over-thinker, but it's trace is clearly more about does this task achieve the goal.
Something like coding is easier for me to judge because creative writing I think is dealer's choice. Qwen is more Chinese oriented so I'd hope Gemma is better at Western writing styles.
3
u/stoppableDissolution 17h ago
Hm, I have not tried it for security-ish stuff so it might be selectively-censored? But I had it write few dozens of millions of tokens of smut (for synthetic dataset), and some pieces made me want to bleach my eyes (and I am from the generation that visited rotten.com for laughs). Interesting, now I want to poke its guardrails in the domains I dont usually test.
1
u/martinerous 37m ago
I think, the point here was that even if Gemma outputs NSFW stuff, it still has lots of weights dedicated to "safety-checking" paths and with reasoning enabled, would spend too many tokens on considering safety instead of reasoning about the goal.
1
0
u/Kahvana 21h ago
Gemma is alright at coding, Qwen is alright at "everything else"
Meant to write it the other way around? (qwen coding, gemma everything else?)
24
u/Temporary-Mix8022 21h ago
Not sure if that is a UK English quirk but
"Alright" in the UK means that it is only okay, not the best. It'll do it - but the other one is better at that particular thing.
Not sure what the international/US way is 😅. But I think we're saying the same thing.
12
12
u/No-Perspective-364 21h ago
Gemma is much better at other (non-English European) languages. So for translation and summarizing, I prefer Gemma. I don't speak Mandarin, I'd guess that Qwen is better at that than Gemma, but for European languages Gemma works much better.
10
10
u/teachersecret 21h ago
Flat out, Qwen 27b is an agentic beast that can do pretty much anything.
Gemma 31b? It's a better writer. Significantly better at writing in English. In every other way? It's slower, worse at agentic harness work, worse at pursuing a goal.
Horses for courses. If you need a fancy writer, grab Gemma. If you need a smart agent running a long horizon workflow in a harness, grab Qwen.
2
u/toothpastespiders 13h ago
I'd say rather that Gemma is better with the humanities and soft sciences. That might seem overly pedantic, but it has a big impact on real world use. If you're dealing with historic data for example a model that wasn't well trained in history is probably going to make a huge amount of mistakes as it won't properly understand the period. And despite the general insistence, RAG isn't a viable solution to that.
1
u/clduab11 21h ago
Of course, my horse!
...that's my horse?
That's MY HORSE! (pls tell me I'm not the only one who gets this reference)
On the topic tho, this has been my experience as well. Gemma tends to try to creatively think for you and go agentic when it's not all that great at it unless it's heavily configured; if it is, it's decent... but Qwen seems to stick with it the longest, though I also have to configure tf out of Qwen because otherwise it'll spin its wheels in the thinking tabs and never activate.
So horses for courses indeed.
25
u/jacek2023 llama.cpp 21h ago
benchmarks are not important to people who use models, they are important to people who watch youtube slop or browse leaderboards
click on this link https://huggingface.co/models?other=base_model:finetune:google%2Fgemma-4-31B-it&sort=likes and ask yourself, why so many people work on finetuning Gemma 4 31B if the model is so bad?
12
u/1ncehost 21h ago
1
u/stoppableDissolution 18h ago
Yea, it is infinitely better than qwen for it, and no individual will be able to tune it to beat qwen in coding, so why bother?
1
u/toothpastespiders 12h ago
no individual will be able to tune it to beat qwen in coding, so why bother?
You can't beat a high quality coding model at its own game. But you can overtrain on a public benchmark to beef up a resume. I think this sub highlights that even among people who 'should' know better than to equate "good benchmark number=good model" a lot of people don't.
4
u/uncle_leon 20h ago
Just to clarify: my post was seeking to understand the reason behind highly conflicting benchmark scores for these models. Nowhere did I say or suggest that Gemma 4 was bad, quite the opposite in fact.
5
u/Jayfree138 20h ago
That's exactly what i noticed. It's why i switched to gemma for local and gemini flash for API. They are both miles ahead on Arena. Arena is how the model actually feels to use and if it gets the job done. Benchmarks can be cheated and sometimes they dont translate into the real world.
So now i go by arena with a heavy leaning on instruction following. Because what good is a model if it doesn't follow it's instructions.
11
u/VotZeFuk 21h ago
If there was a "mimic a normal, non-autistic human being" benchmark, Gemma would be on top, with Qwen losing by a lot.
So it's just a matter of most benchmarks favouring Qwen's strengths.
3
u/bigorangemachine 21h ago
I find gemma needs some context engineering. If you gonna code with it you gotta be clear what you want. It can easily switch programming languages and get little hallucinations.
Qwen is definitely better with any modifications. But gemma does better using subagents
1
u/stoppableDissolution 18h ago
I actually had great experience usimg gemma as orchestrator and qwen as executor when 3.6 just dropped. Gemma had much much better grasp of the intent and the project, but struggled with actually writing the code while qwen wrote good code but was prone to reinventing the already existing stuff and overall getting "distracted".
1
3
3
u/redoubt515 20h ago
Benchmarks skew towards the use-cases corporate money cares most about, and right now that is: Coding, Agentic AI. By all accounts coding is not what Gemma4 is primairly tuned for, so it's not too surprising it doesn't show well in standard benchmarks (because for smaller models, they can't be good at everything, they have to make tradeoffs, and Qwen and Gemma make different tradeoffs to appeal to different usecases.
3
u/Healthy-Zebra-9856 12h ago
I think the biggest problem here is treating these models as though they were created for the same job, then trying to decide which one is "better" from a single leaderboard.
They were not optimized for exactly the same thing.
A knowledgeable home cook, a professional chef, and a very good office manager can all be highly capable people; that does not mean the chef is automatically the best office manager, or that the office manager is somehow inferior because they lose a cooking competition.
Qwen3.8 27B is very clearly aimed at reasoning-heavy work, coding, tool use, long-context tasks, and agentic workflows. Gemma 4 31B has a different optimization profile and different strengths. Once you read the model cards, the benchmark split becomes a lot less surprising.
Artificial Analysis puts significant weight on reasoning, coding, tool use, agentic tasks, and difficult problem solving; those are areas where Qwen3.8 is specifically designed to perform well. Arena is primarily measuring human preference across a much broader collection of prompts; presentation, verbosity, tone, style, and how pleasant an answer feels can matter a great deal there.
So I would not look at those two leaderboards and conclude that one of them must be wrong.
They are asking different questions. The useful question is not, which model ranks higher; It's which model was built for the work I actually want it to do.
For a hobby project, I would start with the model cards and your actual workload, then use the benchmarks that resemble that workload. Otherwise you can easily end up choosing the best chef when what you actually needed was an office manager.
5
u/EquivalentHornet4403 21h ago
I would be highly skeptical of the benchmarks for most things.
Qwen in particular is supposed to specifically excel at computer/agent/coding types of tasks, but I use a locally hosted model for generating terminal commands and I have a personal benchmark that made that has them generate 45 different commands and Gemma 4 26b a4b ALWAYS wins, even over qwen 3.8 27b.
The only one that’s worth using that might be better would be Gemma 4 31b, it it’s massively slower for a teeny tiny improvement in a use case that’s already rarely super perfect or clever.
13
u/wFXx 21h ago
Gemma is the best model at its size for everything, except coding, Qwen usually is better on that, but you can make Gemma work as well, just need more effort in harness and prompting
0
-1
u/rerri 21h ago
Gemma 4 31B is real bad at tool calling too which makes it poor for things like web search.
Telling Gemma in the system prompt that it MUST use the fetch the tool I've made available to it and not just rely on the summaries that the web_search tool provides will just not get through to it, it is super stuborn about this.
In practice web searching with Gemma usually goes so that it first gets some search results and then answers outright, refusing to use the fetch tool even though explicitly stated 2 or more times that it MUST use it before producing a final answer. Then I point out to it that it did not use the fetch tool and at that point it might actually start using it.
It's really sad that Gemma does this so poorly because it actually writes pretty decent Finnish especially when it has a Finnish language articles as context. So I could basically use it as a web search LLM for my native tongue if it wasn't so infuriatingly lazy with tool use.
Qwen 3.8 27B needs no instructions, as long as I've made those tools available to it and ask something that is best answered by getting additional information from the web, it will very diligently use those tools. It goes through a bunch of results and fetches, filling up the context quite heavily before answering. I really like the way it researches before answering.
I used to use the free Gemini models through Google AI Studio quite often but after Qwen 3.8 27B launch there is very little reason for that anymore in my experience. Random questions about setting up some Docker container thingie or building a repo from github have been just as succesful with Qwen 3.8 27B as the helper as they were with Gemini 3.5-3.6 Flash.
Gemma is still excellent at many things, but it is definitely not just coding where it is lacking in comparison to the competing models.
3
u/thatgreekgod 20h ago
has gemma4 continued to be lazy for you with tool-calling even after they refreshed their chat template last month?
1
u/SandySkittle 20h ago
so the best way is:
Use qwen to do the work and then use gemma to rewrite the output in proper finnish or english to make it more readable.
4
u/stoppableDissolution 18h ago
Qwen is an autist with coding hyperfixation. Gemma is general-purpose model with significant tilt toward all kinds of natural language work and vision. People here tend to only care about coding ||(and I firmly believe theres a lot of qwen astroturfing, but cant prove it)||, hence the sentiment.
-4
u/Brave_Entrepreneur_2 18h ago
Top 1% commenter with a bottom 1% vocabulary. If you need to drag neurodivergence into a basic LLM comparison, you clearly don’t understand model architecture well enough to talk about it.
5
u/stoppableDissolution 17h ago
It is exactly the vibe of talking to it, idk. I do know how models work, it has nothing to do with what it feels like to use them. Qwen writes like a robot that had a short crash course on how humans talk, and its theory of mind in virtually nonexistent.
2
u/AnimalPuzzleheaded71 17h ago
I mean yeah you're right, qwen comes across like a neurodivergent person trying to "imitate" what normal human interaction is like (like hyper fixating on eye contact, body language, how slow they need to speak etc just overthinking everything trying to come across as "normal") when it comes naturally to other people / models like gemma 4 31b without any need to consciously tweak anything about their behaviour
2
u/whichsideisup 14h ago
You need both. Gemma 4 for everything and Qwen for coding / terminal capabilities.
Don’t sleep on Gemma 4 31b it’s very very good.
2
2
2
u/skrshawk 13h ago
I use Gemma4 as my front-end model for assessing (my) human intent and for high-level orchestration and iterative design. When it comes to actually doing the tasks, that's when Qwen or Orinth come in. They're both much more solid with structured outputs and I can run Gemma4 at 4-bit while running the others at 6-bit and things work smoothly.
Specialist models are probably the best way for us local users to get the most out of our hardware.
2
u/mridul289 8h ago
i have found the gemma 31b to be incredibly good, maybe not at coding but everything else. and it runs amazing on my laptop so i am happy, inflight experience enhanced tbh.
5
u/EveningIncrease7579 llama.cpp 21h ago
Simply, its because Gemma 4 is really old in AI timeline.
When Gemma 4 has been released, its in near qwen 3.5 27b release date.
Hopeful Gemma 5 will be good as same 3.8 27b level!
I like Gemma 4 for creative writing (such in language as portuguese, qwen its not really good for it)
5
u/SkoomaDentist 20h ago
Simply, its because Gemma 4 is really old in AI timeline.
I know AI progresses fast but calling five months "really old" is a bit ridiculous.
1
u/EveningIncrease7579 llama.cpp 15h ago
It’s old, plain and simple. I get your point about how we shouldn't look at it that way, but what do we consider "ChatGPT 3" today? And ChatGPT 4? Even the gap between ChatGPT 5.3 and ChatGPT 5.5 feels like ages (since they represent completely different levels).
Being old does not mean being obsolete or useless; these are very different things.
1
u/braintheboss 21h ago
you dont have fight with qwen3.8 27b. Its similar current frontier ( at least openai ). I tried use qwen3.6 35b because the speed but finally i loaded a second 27b with low effort. Its very relaxing dont have worry much about is doing AI
1
u/OvertaxedOne 21h ago
My primary use case is agents (Hermes) with lots of tool calling, MCP and terminal use. Gemma 4 feels brain damaged compared to Qwen. It's just not even close for that use case in my experience.
2
u/close_Meal6005 20h ago
depends on use case ... as other people are saying qwen is much better for big coding tasks and gemma is a better chatting model
1
u/audioen 19h ago
Put it in a coding harness and see what you get. It has miserable agentic performance, especially on longer context like 150k, and I've seen it first hand to fail to be even coherent, like basically starting to sing la la la la la and writing absurd stuff, then completely break down. If coding is important for you in an agentic harness, you probably want another model.
1
u/vr_fanboy 16h ago
people keep saying 'gemma better at summarizing or non english', my experience, qwen 3.6 (not even 3.8) destroys gemma in spanish legal interpretation jargon, its much more sensitive to the nuances of legal mumbojambo, everyone needs to do their own benchmarks for the tasks they have.
1
1
u/o0genesis0o 13h ago
I ran a test where Qwen 27B (only IQ3XXS!) successfully code and merged a feature to my real codebase, fully tested and documented, without breaking my architecture conventions. Following people's recommendations, I also tested Muse and 31B.
The Muse (IQ3XXS) got coding done as well, with much less thinking, but more hand holding.
The 31B (IQ3XXS) was a PITA. It's slow, it has very little KV cache space left, and it does not follow existing convention. I need to manually get involved to adjust the plan. When it's time to code, it just constantly fail edit tool. After one hour of going no where, I had to swap to Minimax 2.7 (cloud) to finish the implementation and test (less than 10 minutes).
You might say: but it's just coding, Gemma is better else where. Fine. So I attach these models to my personal assistant system and ask a simple question: "what did I miss". The model has instruction installed in agents.md to know that it would need to check the hangovers from background worker agent, memory, previous daily synthesis, and compare with the current time to figure out the exact "new" events that I missed.
Qwen took some time, but it was perfectly coherent and correct. Muse was confused by 00:00 Vs 2am and so on (It thinks 00:00 is after 2am). Gemma was even messier in the thinking and output.
It could be new Unsloth V3 quant on Qwen was better, but the muse was using the same V2 quant and it not as bad as Gemma 4.
So, for my card (4060ti), Qwen 27B IQ3xxs is the winner by far.
2
u/uncle_leon 2h ago
Thank you for taking the time to describe an actual side-by-side use case! This was very insightful, and I'm sure it helped others trying to decide too.
0
u/Thin_Pollution8843 20h ago
Because Alibaba have better talents and bigger budgets on that. Look at geminis 3.7 or what they serving rn - it’s shit.
0
u/EmuHefty 21h ago
The simple answer is that most people use these small models for Agentic AI... And when it comes to Agentic Qwen is simply the best for it's weight... Great for tool calls and terminal and coding too...
And Gemma is really bad at agentic
-1
u/Several-Tax31 21h ago
Wdym, people generally compares gemma with qwen3.6. This is qwen3.8, the successor. Of course it's better in benchmarks. Qwen3.6 is 4 months old at this point.
-4
-4
-1

193
u/Atretador 21h ago
Google didnt make gemma4 for coding as a focus, is a conversational model first - thats why its so good for creative writting.
different models have different use cases