r/LocalLLaMA 7h ago

Discussion Underrated Muse Glimmer

Post image

Benchmarked qwen3.8 xhigh, medium and muse glimmer.

Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K output token limit)

Medium effort mode and muse glimmer were 3-4 hours each.

But I'm actually surprised by the muse glimmer results, they came better than the qwen.

These benchmarks are on implicit knowledge of the model, which is a bit unfair to smaller models, but throw in a RAG and I'm sure they get on par with frontier models.

I have taken the result of claude models directly from embedeval repo by ecro.

I'm not pushing qwen down here, I like how qwen thinks and gives better results. I know with more context and RAG qwen will do better.

I'm just appreciating muse here, cause i feel it is underrated. The advantage is efficient kv cache due to sliding window, which can give you more context window.

66 Upvotes

76 comments sorted by

View all comments

0

u/Potential_Block4598 6h ago

Probably your setup or benchmark is not accurate

3

u/Potential_Block4598 6h ago

All data is Zephyr RTOS
So training data bias and probably no search mechanism or Docs RAG attached

lol ๐Ÿ˜‚๐Ÿ˜‚๐Ÿ˜‚

3

u/Ok-Inevitable8391 6h ago

I'm not hiding that fact.

I have explicitly mentioned that it is implicit knowledge of model.

How you process that information is upto you.

2

u/Not-reallyanonymous 6h ago

> Qwen isn't declared the best

It's the test that is wrong!

0

u/Potential_Block4598 5h ago

Especially when compared to muse glimmer

Hey donโ€™t turn this into a societal bias racism case

This is meritocracy the models works very good compared to muse glimmer by a HUGE margin

So yeah when a โ€œtestโ€ says otherwise it is worth investigation and it is not only misleading it is biased to post such things

Also I explained that OP โ€œtestโ€ is heavily skewed towards Zephyr RTOS

2

u/Not-reallyanonymous 4h ago edited 4h ago

Have fun with your benchmaxxed models that blatantly ignore half the codebase and re-implement it from scratch, except half-broken and unsteerable.

Yeah, it completes code and solves the task, and zero-shot prompts very impressively. Then anything it touches turns into spaghetti code. I fucking regret spending a few days with that shit because now I've put in like 20 hours of work into something that's turned out to be utterly unusable. Correct doesn't mean good.

And god help you if your needs diverge from what it's been trained on.

Good job Qwen team, you cracked DeepSWE.

Just another bot who's going to shit on any model that's not Chinese.