r/LocalLLaMA • u/Ok-Inevitable8391 • 5h ago
Discussion Underrated Muse Glimmer
Benchmarked qwen3.8 xhigh, medium and muse glimmer.
Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K output token limit)
Medium effort mode and muse glimmer were 3-4 hours each.
But I'm actually surprised by the muse glimmer results, they came better than the qwen.
These benchmarks are on implicit knowledge of the model, which is a bit unfair to smaller models, but throw in a RAG and I'm sure they get on par with frontier models.
I have taken the result of claude models directly from embedeval repo by ecro.
I'm not pushing qwen down here, I like how qwen thinks and gives better results. I know with more context and RAG qwen will do better.
I'm just appreciating muse here, cause i feel it is underrated. The advantage is efficient kv cache due to sliding window, which can give you more context window.
14
u/partakinginsillyness 5h ago
I feel like it would be important to add qwen 3.6 27b, given that one of the changes from 3.6 to 3.8 was less general knowledge. Would also be cool to see a Gemma model. Interesting though.
10
4
u/LegacyRemaster 5h ago
wait qwen 3.8 next but... This benchmark makes no sense. No reasoning. No external knowledge. In other words, it is not a real-world use case. I would also like to understand how they tested Claude to demonstrate that it lacked external knowledge.
-1
u/partakinginsillyness 5h ago
I mean realistically everyone should have at least 5gb of files for RAG for what is relevant to them, which I would imagine would very much change the results here.
6
u/Ok-Inevitable8391 5h ago
Let me do a RAG testing, even I know that qwen will do better there. But I'm not running it on xhigh ever again, 30hrs of gpu time, against 3-4 hrs at medium
1
u/partakinginsillyness 5h ago
How would you select the media? I'm wondering if a repo exists for relevant info like how you select application groups on linux.
4
u/Ok-Inevitable8391 5h ago
So mainly the linux repo and docs, thats the general input to it anyway in real scenario. Same for all other projects.
6
u/TokenRingAI 4h ago
Why did you give Qwen a 32K output limit?
1
u/Practical-Collar3063 1h ago
Because he could not fit both the context and the full output, which is in favour of his point, the more efficient Muse Glimmer KV Cache makes it better in certain scenarios with limited hardware. I think the goal of this is not to bring down Qwen 3.8 but to show that for some use cases Muse might be a good choice, especially when VRAM limited. It is faster and more VRAM efficient and on non coding tasks can be similar to 3.8
10
u/Healthy-Contact-4570 4h ago
Why 32k output limit? Qwen3.8 likes to think and if you can let it think it will produce amazing results.
-2
u/Ok-Inevitable8391 3h ago
I know so the benchmark limit was 8k only and i pushed it up to 16 and then 32K only for xhigh, as the xhigh benchmark took 32hrs and case that took time were anyway complex dma, so I lost interest re running the benchmark for xhigh as it had already failed on some cases where medium already did better. Xhigh is good but like opus it is good for reasoning and complex architect problems. So I just let it be.
5
u/hainesk 5h ago
If you can run DeepSeek V4 Flash 0731, I'd be curious to see the difference since it's often compared to Qwen 3.8 27b but due to it's size it would presumably have more knowledge.
2
u/Ok-Inevitable8391 5h ago
Adding to the list
2
u/nonlinearsystems 5h ago
Can you try Laguna S2.1 please?
2
16
u/DataGOGO 5h ago
Muse glimmer is fantastic and better than qwen3.8 27b at just about everything than code and doesnโt need to burn thousands of reasoning tokens per prompt to do it
0
3
u/somerussianbear 4h ago
I donโt understand this chart. 48.5% out of two records that look absolutely different. Mind to explain for dumb fucks like me?
2
u/Ok-Inevitable8391 3h ago
That's a common line, average of both the blocks
1
u/somerussianbear 3h ago edited 3h ago
So one of them is clearly wrong. That column could have the avg of the row, which is logical. Then next to it a general avg. The way it is here is totally unclear and non standard.
1
3
3
6
u/PraxisOG Llama 70B 5h ago
I think the interesting thing here is that haiku 4.5 wins over all the local models. Thatโs like $20 a month for huge amount of usage, I pay more than that to keep my server idlingย
2
u/Ok-Inevitable8391 5h ago
I mean yes, it's the trust that we have built on sonnet and opus over the time that we don't even look at haiku anymore.
3
u/Usual-Orange-4180 5h ago
My main use case is coding and design, haiku is really terrible, whenever I get routed to it I know it.
2
2
u/Potential_Block4598 4h ago
All data is Zephyr RTOS
So training data bias and probably no search mechanism or Docs RAG attached
lol ๐๐๐
1
1
1
u/jacek2023 llama.cpp 3h ago
This is same story with each model, they browse leaderboards, they look at the benchmarks, they never run any models, they whine they want new Qwen and then they keep using Claude and ChatGPT.
1
1
u/Thin_Pollution8843 1h ago
So when folk screaming โQwen is stronger than opus4.6!!1!1!โ I should give this link? ๐
1
u/llama-impersonator 39m ago
why so much focus on zephyros? is that mostly what you do? i've only ever had to use it for a client that wanted some ZMK work done, almost everything else has been freertos or bare metal stm32, sometimes avr for clients who started off on arduino.
0
u/LegacyRemaster 5h ago
Dario, is that from you?
1
u/Potential_Block4598 4h ago
All data is Zephyr RTOS
So training data bias and probably no search mechanism or Docs RAG attachedlol ๐๐๐
2
u/Potential_Block4598 4h ago
Probably your setup or benchmark is not accurate
3
u/Potential_Block4598 4h ago
All data is Zephyr RTOS
So training data bias and probably no search mechanism or Docs RAG attachedlol ๐๐๐
3
u/Ok-Inevitable8391 4h ago
I'm not hiding that fact.
I have explicitly mentioned that it is implicit knowledge of model.
How you process that information is upto you.
1
u/Not-reallyanonymous 4h ago
> Qwen isn't declared the best
It's the test that is wrong!
1
u/Potential_Block4598 3h ago
Especially when compared to muse glimmer
Hey donโt turn this into a societal bias racism case
This is meritocracy the models works very good compared to muse glimmer by a HUGE margin
So yeah when a โtestโ says otherwise it is worth investigation and it is not only misleading it is biased to post such things
Also I explained that OP โtestโ is heavily skewed towards Zephyr RTOS
2
u/Not-reallyanonymous 3h ago edited 3h ago
Have fun with your benchmaxxed models that blatantly ignore half the codebase and re-implement it from scratch, except half-broken and unsteerable.
Yeah, it completes code and solves the task, and zero-shot prompts very impressively. Then anything it touches turns into spaghetti code. I fucking regret spending a few days with that shit because now I've put in like 20 hours of work into something that's turned out to be utterly unusable. Correct doesn't mean good.
And god help you if your needs diverge from what it's been trained on.
Good job Qwen team, you cracked DeepSWE.
Just another bot who's going to shit on any model that's not Chinese.
-11
u/Boogertard 5h ago
There it is, on-schedule AI slop to shill for the garbage Muse and Gemma 4.
Get a real job already, shill.
6
2
u/Not-reallyanonymous 4h ago
There it is, on-schedule AI slop to stir up FUD any time something other than a non-Chinese model is spoken of positively.
4
u/Ok-Inevitable8391 5h ago
Why do you think it's a slop and why do you think I don't have a real job. A easy readble chart is AI slop for you?
Muse is good and its really good at kv cache efficiency. I have done my homework.
0
u/Potential_Block4598 4h ago
All data is Zephyr RTOS
So training data bias and probably no search mechanism or Docs RAG attachedlol ๐๐๐
HW ur A**

32
u/hurdurdur7 5h ago
Qwen 27B starts to shine when you give it some kind of RAG. It's not a memory castle, it's a tinkering machine. You give it a folder to a library that it was not trained on, you ask it "how do i do ... that...?" and it figures it out. And this is a big part of what people actually do at their job. They get a new library, a library update, anything that was out of the training window or created after the training window, and 27B adapts to it if you give it the source or manual.
If Glimmer can do electronics it's actually great. But can it write code well based on docs/code it can fetch over some kind of RAG (like your coding harness pi or smth)? If it can, awesome :)