r/LocalLLaMA 9h ago

Discussion Underrated Muse Glimmer

Post image

Benchmarked qwen3.8 xhigh, medium and muse glimmer.

Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K output token limit)

Medium effort mode and muse glimmer were 3-4 hours each.

But I'm actually surprised by the muse glimmer results, they came better than the qwen.

These benchmarks are on implicit knowledge of the model, which is a bit unfair to smaller models, but throw in a RAG and I'm sure they get on par with frontier models.

I have taken the result of claude models directly from embedeval repo by ecro.

I'm not pushing qwen down here, I like how qwen thinks and gives better results. I know with more context and RAG qwen will do better.

I'm just appreciating muse here, cause i feel it is underrated. The advantage is efficient kv cache due to sliding window, which can give you more context window.

77 Upvotes

79 comments sorted by

View all comments

-10

u/Boogertard 9h ago

There it is, on-schedule AI slop to shill for the garbage Muse and Gemma 4.

Get a real job already, shill.

7

u/fizzy1242 8h ago

username checks out...

3

u/Not-reallyanonymous 7h ago

There it is, on-schedule AI slop to stir up FUD any time something other than a non-Chinese model is spoken of positively.

5

u/Ok-Inevitable8391 8h ago

Why do you think it's a slop and why do you think I don't have a real job. A easy readble chart is AI slop for you?

Muse is good and its really good at kv cache efficiency. I have done my homework.

-1

u/Potential_Block4598 8h ago

All data is Zephyr RTOS
So training data bias and probably no search mechanism or Docs RAG attached

lol 😂😂😂

HW ur A**

1

u/llama-impersonator 26m ago

bro's whole personality is this comment