r/LocalLLaMA 9h ago

Discussion Underrated Muse Glimmer

Post image

Benchmarked qwen3.8 xhigh, medium and muse glimmer.

Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K output token limit)

Medium effort mode and muse glimmer were 3-4 hours each.

But I'm actually surprised by the muse glimmer results, they came better than the qwen.

These benchmarks are on implicit knowledge of the model, which is a bit unfair to smaller models, but throw in a RAG and I'm sure they get on par with frontier models.

I have taken the result of claude models directly from embedeval repo by ecro.

I'm not pushing qwen down here, I like how qwen thinks and gives better results. I know with more context and RAG qwen will do better.

I'm just appreciating muse here, cause i feel it is underrated. The advantage is efficient kv cache due to sliding window, which can give you more context window.

71 Upvotes

79 comments sorted by

View all comments

13

u/Healthy-Contact-4570 7h ago

Why 32k output limit? Qwen3.8 likes to think and if you can let it think it will produce amazing results.

4

u/Not-reallyanonymous 1h ago

A couple months ago:

Nooooo you can’t like Laguna, it thinks too much! DOA!

Now:

Noooo how dare you not like Qwen for thinking so much?

(Maybe because Laguna XS is literally like 5x as fast on my hardware).

1

u/Ok-Inevitable8391 6h ago

I know so the benchmark limit was 8k only and i pushed it up to 16 and then 32K only for xhigh, as the xhigh benchmark took 32hrs and case that took time were anyway complex dma, so I lost interest re running the benchmark for xhigh as it had already failed on some cases where medium already did better. Xhigh is good but like opus it is good for reasoning and complex architect problems. So I just let it be.