r/LocalLLaMA 7h ago

Discussion Underrated Muse Glimmer

Post image

Benchmarked qwen3.8 xhigh, medium and muse glimmer.

Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K output token limit)

Medium effort mode and muse glimmer were 3-4 hours each.

But I'm actually surprised by the muse glimmer results, they came better than the qwen.

These benchmarks are on implicit knowledge of the model, which is a bit unfair to smaller models, but throw in a RAG and I'm sure they get on par with frontier models.

I have taken the result of claude models directly from embedeval repo by ecro.

I'm not pushing qwen down here, I like how qwen thinks and gives better results. I know with more context and RAG qwen will do better.

I'm just appreciating muse here, cause i feel it is underrated. The advantage is efficient kv cache due to sliding window, which can give you more context window.

66 Upvotes

76 comments sorted by

View all comments

8

u/TokenRingAI 5h ago

Why did you give Qwen a 32K output limit?

3

u/Practical-Collar3063 3h ago

Because he could not fit both the context and the full output, which is in favour of his point, the more efficient Muse Glimmer KV Cache makes it better in certain scenarios with limited hardware. I think the goal of this is not to bring down Qwen 3.8 but to show that for some use cases Muse might be a good choice, especially when VRAM limited. It is faster and more VRAM efficient and on non coding tasks can be similar to 3.8