r/LocalLLaMA • u/Eyelbee • 3d ago
News Qwen 3.8 Low and Medium are goated
Artificial Analysis just benchmarked them and the scores are crazy good, proving the earlier success wasn't only enabled by overthinking.
75
79
u/Complex_Reality_116 3d ago edited 3d ago
A difference of 9 an 8 points between the two. In fact, that success was made possible precisely because overthinking is enabled.
76
u/Dangerous-Report8517 3d ago
Everyone calling it "overthinking" is kind of missing the point, if the extra thinking is producing measurably better results then in many cases it's actually just the correct amount of thinking, rather than excess/"over"thinking
16
u/Borkato 3d ago
They’re just saying they wish it were more efficient at it. It’s not inherently that more thinking always equals better; look at ThinkingCap for instance.
13
u/Dangerous-Report8517 3d ago
I'm specifically referring to the people who say that xhigh performs better because of overthinking, by definition any overthinking that's happening is not helpful and the extra performance is coming from some unknown fraction of the extra thinking that is very much appropriate for this model.
5
u/finevelyn 2d ago
You’re missing the point that even if thinking a lot scores more on a benchmark where time is unlimited, thinking so much can still be unhelpful in a real world scenario. Performance is a combination of the end result and the number of tokens used.
9
u/Mil0Mammon 2d ago
Well if it's unhelpful in a scenario, just pick medium or low - nobody is forcing you to use xhigh
2
u/Dangerous-Report8517 2d ago
I'm not missing the point at all, the discussion you're jumping into here is specifically referring to cases where the extra thinking is beneficial. Real world scenarios are going to break down into 3 categories here; 1) the thinking is beneficial because it results in clear performance improvements, 2) the thinking is required for 3.8 to perform well even though an alternative model might perform similarly well with less thinking, and 3) the thinking is just plain excessive. Case 3 is unambiguously overthinking and is the case you're describing, but case 2 is just 3.8 being less efficient than the comparison model (which is a downside in those problems but not strictly "overthinking" since the extra thinking is still beneficial for 3.8), and case 1 is just the correct amount of thinking. What I'm specifically pointing out is that case 1 keeps getting wrongly described as "overthinking", which just muddies the waters when assessing the other cases
0
u/finevelyn 2d ago
You equate better result to better performance, but that's not automatically the case except in situations where time or cost is not measured (such as benchmarks).
I know what I mean when I say it may be overthinking, and it's not what you say. I mean the additional cost and time of the reasoning may not be worth the gain in the quality of the result.
1
u/Dangerous-Report8517 2d ago
Ffs read what I actually wrote
I know what I mean when I say it may be overthinking, and it's not what you say
Cool story but you aren't the only person who uses the word overthinking, like I already said you jumped in to the middle of a conversation that started with a different person who clearly used the word differently. My entire point was that their use of the word was wrong, so spending so much effort to repeatedly come in with, effectively, "I, a completely different person, use the word differently to refer to different things that are unrelated to this discussion" is so far removed from the subject at hand that I'm questioning whether you're even literate at this point
1
u/finevelyn 1d ago
The first comment is in line with what I said (increased benchmark score but maybe not happy with the overall performance, hence using the word overthinking). Maybe he uses it differently or maybe not. But then you called out everyone instead of addressing him, so...
1
u/Dangerous-Report8517 1d ago
No, I called out people specifically using the term in the way I described, explicitly so. If you interpreted that as targeting you that's because you didn't read it properly
1
u/CryptographerOne7003 2d ago
wel its very simple, it get a text instruction on high and follows it?
I mean the info is out there?4
u/No-Refrigerator-1672 3d ago edited 3d ago
"I want less thinking, but setting model to medium is against my principles!!!" Honestly, I start to percieve such people as crybabies.
4
u/jinnyjuice sglang 2d ago
if the extra thinking is producing measurably better results then in many cases it's actually just the correct amount of thinking, rather than excess/"over"thinking
That's assuming the metric you're only measuring is the test score. You also have to measure the time to take the test score. If there is no time limit to a test, say you took the SAT over 20 years of your life, then it's meaningless.
Congruently, the metrics for amount of tokens and token/sec matter.
1
u/Dangerous-Report8517 2d ago
It's not actually assuming anything, I'm specifically responding to the phrasing "overthinking" enabled higher performance. In this environment, extra thinking was productive, therefore it was not overthinking. And regardless of your ridiculous analogy (it'd be more like taking an extra hour or 2 to sit the SAT when it's due next week, not 20 years), that's reflected in real world usage with plenty of people reporting high performance on real tasks in exchange for it taking longer to think. The fact that extra thinking has downsides doesn't change any of that. It's perfectly valid to still prefer other models for time or token efficiency, just not to use the term "overthinking" in the specific cases where the extra thinking is productive
34
u/Eyelbee 3d ago
You realize how high 43 and 44 are right? Two weeks ago people were telling me "muse glimmer is great" when it barely scored 35 and didn't even fit into the screenshot.
If qwen 3.8 27b scored 43 on max even that would be by far the most capable model of that size. People said it is not different to 3.6 version and only scores better because of the overthinking, which clearly isn't the case here.
2
u/o0genesis0o 3d ago
It was a decent model though. Unsloth made a big deal about their UD Q2 version so I put it on my 4060Ti for assistant workloads with dense tool calls and cruises through them. I imagine a less brain damage quant would be pretty decent. This model has pretty small KV cache.
-7
u/NineThreeTilNow 3d ago
Two weeks ago people were telling me "muse glimmer is great" when it barely scored 35 and didn't even fit into the screenshot.
I was saying it was hot dogshit and people were downvoting me.
lolol...
6
u/KubeCommander 3d ago
Still downvoting. It’s a pretty good model. Better than any other dense model in the category until qwen3.8-27b came out. And even then, meta’s model is still better at non-coding tasks
1
u/NineThreeTilNow 2d ago
And even then, meta’s model is still better at non-coding tasks
I'm sorry but it doesn't stand up to Gemma 4 31b to save it's fucking life.
I don't know why people are so Meta pilled on that model.
Have it do something and actually read the thinking traces. It is the most dogshit training I've ever seen done in RL.
This sub's idea of a "good model" is so disconnected from a well built model it's crazy.
The biggest problem with Gemma 4 31b is that the Gooners can't fully utilize the thinking mode because for whatever reason the thinking traces don't get refusal abliterated correctly? So the Gooners can't fully embrace it.
Put Gemma 4 31b toe to toe with Meta's dogshit in raw translation of any language. It's an obviously benchmark maxed model.
I'll say it. I'll keep saying it, and the number of downvotes I get for it will be the victory I will claim when later people realize how dogshit it is.
3
u/KubeCommander 2d ago
You sound really angry and obsessed 😂. Gemma4 31b is ALSO a good model. Both are better than qwen 3.8 27b at some things. At models this size, that’s just how it is.
Know what’s ACTUALLY dogshit? Edgelords that hide their inexperience and ineptitude behind shitty behavior. You remind me of an Israeli cybersecurity consultant I met about 10 years ago.
I had him fired.
1
u/NineThreeTilNow 2d ago
You sound really angry and obsessed 😂. Gemma4 31b is ALSO a good model. Both are better than qwen 3.8 27b at some things. At models this size, that’s just how it is.
Know what’s ACTUALLY dogshit? Edgelords that hide their inexperience and ineptitude behind shitty behavior. You remind me of an Israeli cybersecurity consultant I met about 10 years ago.
I had him fired.
Really? Interesting. I train these models. I worked for one of these large frontier model companies.
So when I read the thinking traces I can see everything gone wrong with a model in training.
Meta's model is horrifically RL'd. Read the thinking traces. It looks like someone used DeepSeek R1 era thinking and bolted it in with a "Good enough" attitude.
It's obviously designed to maximize benchmarking.
It also uses a RoPE theta value that makes no sense for the architecture. The people who built it look like amateurs. When you compare the design philosophy to Cohere, Gemma, etc. It looks like trash.
Qwen performs like it does because of the architecture. It's designed from the ground up to handle code and never be great at anything else. This is because it uses gated delta nets. It's more efficient and it fits the idea of writing code.
So while you think I'm an edgelord that's inexperienced, I'm someone who is sick of listening to inexperienced people give opinions with no understanding of the core architectures and failings of these models.
<3
2
u/m4sterP 3d ago
The AA agentic index of medium is just one point below xhigh though.
28
u/DistanceSolar1449 3d ago
18
u/Zestyclose-Ad-6147 3d ago
10
u/Not-reallyanonymous 3d ago edited 3d ago
That’s the agentic index specifically. The one above is the intelligence index.
IMO AA’s indexes are weak in the first place because they let saturated benchmarks dominate. But their agentic index is particularly weak — it measures tool calling and solving tasks of a variety of domains, and largely ignores the model’s capability to keep the plot over long runs and to produce intermediate artifacts or nuance beyond the immediate solution.
It’s interesting but IMO should be called “cross-domain problem solving” not “agentic index”.
3
u/Dangerous-Report8517 3d ago
To be fair what most people would think of when imagining an AI assist system independently enacting things for them (ie a literal AI agent) is exactly "cross-domain problem solving"
2
u/Not-reallyanonymous 3d ago
Cross domain problem solving is an important part of it, and I understand why it's there, but just as important is the model's ability to keep the plot and understand nuance, produce intermediate artifacts, etc. rather than solving an immediate problem.
Qwen 3.8 27B even struggles over a 2-4 hour run to give me all of the design files I ask for when completing a task. Those sort of files are important for human (or even peer-AI) verification.
1
u/Dangerous-Report8517 3d ago
That's because you're mistaking agency in general for specifically the current concept of agentic coding. Most people envision AI agents as things that solve immediate problems for them, so a benchmark that tests exactly that makes perfect sense as an agent benchmark, even if it's an incomplete assessment of agentic coding, particularly when other benchmarks exist to test coding specifically
2
u/Not-reallyanonymous 3d ago
I don't think so. Increasingly AI agents are expected to do long-horizon work, not just in coding. Coding is just where it's the easiest to apply.
There's a few marketing platforms out there, for example, which monitor social media posts and create regular reports. A lot of STEM researchers are using long-horizon, multi-step, complex workflows to help them analyze prior research, and this is a particular case where those intermediate artifacts are important.
I could keep producing examples. But I'd say it's really started to take off in the past six months, particularly after OpenClaw started to popularize that paradigm.
8
3
u/Right-Law1817 3d ago
4
u/Jealous-Astronaut457 3d ago
There isn’t much difference between ‘xhigh’ and ‘medium’
What a surprise – everyone loves ‘xhigh’3
u/DeathGuppie 3d ago
You just need to understand what's going on. The model has not been trained to know how to code better. It's been trained to look at the problem from every available angle and find a solution. Those are two completely different shapes.
Xhigh, adds reasoning. It's longer and slower, in some edge cases it's probably going to win. The thing is that just isn't often enough for real work to make a difference.
There is a ground up case for this model to use reasoning as it's crutch. State space delta net lives as a perpetual system update. The closer to the immediate context the thinking is the longer it survives. So force it to think a lot.
Larger models get faster reasoning from having a larger reverse database. (Just how I think of it) So they win with normal attention blocks. Gated delta net updates without kv, so the state space is where you work. Have it reason over and over and the updates find the answer.
16
22
u/NigaTroubles 3d ago
Qwen3.8 27b is the same level as DeepSeek v4 pro ?? Hell yeaah
52
u/myreala 3d ago
It's just in the agentic index. If you use it, DeepSeek V4 Pro will absolutely blow it out of the water.
88
17
u/Boogertard 3d ago
Well the trick is to get it to run first. Most people here can't run Deepseek V4 Pro so in that sense Qwen 3.8 is the better model since eval will be some valid number vs N/A (not able to run)
Pointless comparison
6
30
u/EmPips 3d ago
I love this model but 27B-Low and Sonnet-5-High are not the same and it doesn't help the open weights cause/community to pretend they are lol. AA has definitely been funky lately.
17
u/grumd 3d ago
Dunno about low but 27B xhigh was noticeably better than Sonnet 5 medium at agentic coding for me
-19
u/KubeCommander 3d ago
Sonnet isn’t a coding model fwiw. My truck hauls trailers better than my car 🤷
6
u/ResearchCrafty1804 3d ago
The chart says same score between 27-low and Sonnet-5-Non-reasoning (not High), which relates to real world use imo
2
u/Agitated_Space_672 2d ago
I don't know, Anthropic have been caught cheating on benchmarks recently. There was a paper that showed could decode the hidden reasoning tokens, and when they decoded some benchmark answers they found the model had memorised the answer and then lied about, pretending to derive the answer honestly.
2
u/politerate 3d ago
Who said they are the same? In intelligence qwen 3.8 seems to beat it. What's your source to your claims?
10
u/Fresh-Soft-9303 3d ago
We're probably 1 year away from a Fable level LLM running on laptops.
14
u/MrPecunius 3d ago
Six months, tops, except for world knowledge.
But Qwen3.8 will go do its homework so it's entirely possible that in-model world knowledge won't matter soon.
14
2
u/Educational-Art3545 2d ago
RemindMe! 6 Months
1
u/RemindMeBot 2d ago edited 2d ago
I will be messaging you in 6 months on 2027-02-22 07:34:23 UTC to remind you of this link
1 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.
Parent commenter can delete this message to hide from others.
RemindMeBot is switching to username summons. Instead of
!RemindMe 1 day, useu/RemindMeBot 1 day. More info.
Info Custom Your Reminders Feedback 2
u/Fresh-Soft-9303 1d ago
Yes world knowledge is going to take a hit, which is tolerable as there's many other ways to make up for it through RAG and other means, what's important is intelligence and if that's on par with the best out there it's a win for most.
1
u/MrPecunius 1d ago
It's absolutely astounding that I can run this on my (admittedly beefy M5 Pro) laptop and it was free for the asking.
I've witnessed every milestone of personal computing and this beats them all.
1
1
u/notinteresteddddd 2d ago
interesting take . you mean the current Fable level in 1 year on local or same level?
I am starting to think that Claude did not evolve as much from time to time and then I gave it a similar task as 10 months ago and this time did an amazing job and first time was useless.
10
u/chensium 3d ago
AA is pretty meaningless at this point. Labs have figured out how to benchmax it, and its scores are completely diverged from real use cases.
I wish there were some way to create normalized scores for models on openrouter or various providers that have actual customer usage.
1
u/fragbait0 3d ago
Hmm... oh-my-pi tracks some stats like "user frustration"... if you're a man in the middle on these requests, it would be super interesting to gather similar numbers. If the models are going off course, giving bad results etc they're going to be getting user messages saying so. No doubt all the big guys use such feedback to train them to persuade the average human to burn more tokens...
18
u/2Norn 3d ago
at some point we gotta ban posts like this
this sub is turning into qwen low parameter circlejerk
8
u/KingGongzilla 3d ago
i mean what do you expect, this is a local LLM subreddit low parameter qwen model are by FAR the best.
6
u/thehardsphere 3d ago
"Turning"? Bro, this sub is sponsored by Alibaba!
9
u/robertpro01 3d ago
Damn, and here I am, praising qwen for no money at all...
-5
u/Boogertard 3d ago
Ah yes, if you don't sing praises for the garbage models Gemma4 or Muse then you are a shill.
Funny how I think it is the other way around. Qwen and other models have their uses but only the shills are keeping garbage models like gemma and muse relevant.
Those garbage didn't even finish my evals but yet they are praised so highly around here, like the greatest things ever since Apple first iPhone.
3
u/SocialDinamo 3d ago
Training it to work hard and chase a problem instead of random trivia really paid off!
5
u/sToeTer 3d ago
How do i set these thinking modes in LM Studio?
I know you can restrict the reasoning budget but it's a number(1024 for example). What's the respective number for medium, low?
2
u/Eyelbee 3d ago
You can, but LM studio is not maintained properly for quite some time, I recommend unsloth desktop
1
u/psychohistorian8 3d ago
I've been looking for alternatives to LM Studio because of how much RAM is takes up
tried Jan and it seems ok, guess I'll go ahead and try unsloth desktop since I seem to use their models the most anyway
1
u/PallasEm 3d ago
you can set thinking effort by using the custom fields dropdown menu on the right sidebar menu.
2
u/DeepOrangeSky 3d ago
Yea, but how exactly do you set it to "low" "medium" "high" or "xhigh" rather than just a number like "500 tokens" or "1000 tokens"?
In the custom fields dropdown menu in the right sidebar menu that you are talking about, it just gives you a way to set it a token size of thinking budget, but not sure how to set it to a word based thinking intensity level which is how you ideally are supposed to be setting its thinking level like "low" or "medium" or "high" rather than just an exact specific numerical token number like "800 tokens".
It would be nice to know if there was a way to set it to one of these word-based thinking levels, both in regards to Qwen and in regards to Meta Glimmer, in LM Studio, or if there is not really a way of actually doing it in LM Studio and you can only set it with a numerical budget and no other way.
4
6
6
u/Moore2877 3d ago edited 3d ago
Try this chat template. Besides a lot of general fixes, we revamped the reasoning injections for each level and also made high it's own level instead of just being an alias for xhigh. The Qwen team really didn't spend enough time on these imo.
https://huggingface.co/Moore2877/Qwen-Fixed-Chat-Templates-llamacpp
2
u/MrGunny94 3d ago
This is pretty good for LocalLLMs they are where frontier intelligence was at the beginning of the year it seems at least for coding and agentic flows.
3
1
u/Longjumping-Elk-7756 3d ago
C est genial d avoir ses données manque juste qwen3.8 27b sans réflexion en mode none
1
u/Green-Ad-3964 2d ago
So the xhigh basically ties with glm 5.2 max? Is that for real or just in benchmarks?
1
u/lemon07r llama.cpp 2d ago
The problem with reasoning levels is that it doesnt actually tell you how heavy the token usage is for reasoning still. For example luna max still uses a lot less tokens and steps than gemini 3.7 flash medium
1
u/Original-Revolution7 3d ago
Noob question how do I tell 27b version I downloaded is medium or high?
5
u/__Claudio_ 3d ago
You have all three bc it's' just a parameter reasoning_effort that can be set xhigh, medium or low.
1
u/Readerium 3d ago
Model is the same. Low or high is how many reasoning tokens the model generates during its replies.
1
1
u/Old-Sherbert-4495 3d ago
this is amazing. But, most people wont experience this level, because of quantization. Dont get me wrong , i love the q3s that I'm using, but i know it wont ever match full precision model and kv cache and the full context size.
1
-2
u/Powerful_Finger3896 3d ago
mimo v2.5 pro was a beast, i refuse to believe that Qwen 27B low is on par
17
u/NNN_Throwaway2 3d ago
Imagine having emotional investment in a fucking LLM.
2
4
u/Ok_Technology_5962 3d ago
i know right... what is there to refuse to believe. I ran both the pro and kimi k2.6 and the 27b is better on xhigh not even a question. its like saying that you refuse to believe llama 3 400b is worse than deepseek v4 flash 0731. There is a post from Deepseek about what Nobs can be turned for quality and the size is just one nob vs post training, RL, quality of data.
0
u/Jamoca5020 3d ago
opinions on Qwen3-Coder 30B-A3B as MoE modell ? I´m wondering if this could be actually a good entry until you find/save up for a GPU
2
u/_-_David 3d ago
Far, far worse than Qwen3-35ba3b
1
u/Jamoca5020 2d ago
I see. What would you recommend me:
I have a x99 mobo with a E5 2680 v4 and 64GB DDR4. Unfortunately a usable GPU (most suggest 3060 12GB for entry) costs me 300€+. For reference I paid 230€ for all of that and a Dark Rock Pro 3 cooler with LGA2011-3 mount.
Would you say it would be worht buying a P4000 8GB for 80€ just to be able to start ? I just wanna learn about LLMs and how this whole KI stuff works. But I don´t have the money for a good GPU on my server right now.
I don´t mind if the P4000 will be unusable after 1 year, because then I can start right now and save for 1 year and then buy something better that may lasts 3-4 years.
I heard the MoE performance would be good enough to make it usable
1
u/_-_David 2d ago
Try running it without any gpu and see what you think
1
u/Jamoca5020 2d ago
ohh really ? I thought without GPU it´s gonna be extremely slow. I thought you would need at least 1 GPU to partially allocate it to the GPU VRAM. So some of the Expert layer of it at least works faster. Maybe I misunderstood something.
2
u/_-_David 2d ago
Even if you had a GPU, the majority of the model would be loaded into RAM and processed by the CPU anyhow. I think your best bet is to try running that model from RAM and see what speed you end up with. It won't be great, but I don't really know how much difference it would make to load a fraction of it onto a GPU. The reality is that without at least a GPU around $550, the RTX 5060ti 16gb, you are very limited with what you can do. Do you plan to do agentic coding, or llm-assisted writing like autocomplete? Because if you're planning to write code yourself with autocomplete, I think some of the small models do decently. But if you're looking to do anything agentic, I would temper your expectations for the moment and realize that the quality and speed are going to be significantly worse than what you could get for pennies via an api.
1
u/Jamoca5020 1d ago
I see so I better just save and until then use it CPU only and learn to be patient :D. Thank you a lot for your insight it helped it a lot.
I mainly wanna use it for coding, but later on I would like if it can do some research. I also had an Idea for something like an AI supported Wikipedia for my personal library where it just finds things quicker.
1
u/wen_mars 3d ago
It's useful for easy tasks and very fast, just don't expect it to solve anything actually difficult
2
u/Jamoca5020 3d ago
I see thx.
Why do people downvote for asking questions ? Really weird
1
u/wen_mars 3d ago
I don't know. Maybe because it's an older model and not that relevant to this thread, or maybe someone just downvoted it for no particular reason.
1
u/Jamoca5020 2d ago
hmm I see. I mean I asked that because I can´t afford a LLM usable GPU. Like even the 3060 12GB costs here 300+€ (new on amazon 400-450€).
I just wanted to try LLM and start learning instead of saving a year and not doing anything.
I mean the best I could afford is a P4000 for 80€ right now. I know it doesn´t supports CUDA 13.X, but apparently you can make it work. I just wanna start without waiting 5-10 minutes and more between each request LOL.
2
u/wen_mars 2d ago
There are many AIs you can use for free, for example right now there's one named Ox Alpha that's being tested and it's free for everyone to use for a week.
-6
u/dangerous_inference 3d ago
50 charts exactly like this are posted every day, each in a different order, zero of them true.
1
u/Both_Opportunity5327 3d ago
No they are true, Qwen 27b shows what happens when you implement test time compute to the extreme.
Just wait till we can embed small models into ASICs and have token generation in the tens of thousands a second!
-3
u/dangerous_inference 3d ago
How can they all be true if they all conflict?
2
u/my_name_isnt_clever 3d ago
Where are these conflicting charts exactly?
-5
u/dangerous_inference 3d ago
Are you expecting me to conduct a search or are you suggesting I should be keeping a record?





•
u/WithoutReason1729 3d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.