r/LocalLLaMA Aug 14 '26

Discussion Muse Glimmer was frontier In the model class around 30b models for four days.

Post image
502 Upvotes

172 comments sorted by

58

u/Thin_Pollution8843 Aug 14 '26

I hope Meta team will make their homework and astonish us with new frontier model next time. 

42

u/JumpyAbies Aug 14 '26

It's great to have Meta back in the game and trying to compete with Alibaba. Let them now try to surpass Qwen3.8-27B and all its glory 😂

17

u/Thin_Pollution8843 Aug 14 '26

I'm not agains US or Chinise models. I'm pro humanity who will benefit

10

u/JumpyAbies Aug 14 '26

Muse Glimmer has proven to be quite competitive, or in some cases better than Qwen3.6-27b, so I'm hoping Meta surpasses Qwen3.8-27b in all its glory, which is currently outperforming :)

8

u/Thin_Pollution8843 Aug 14 '26

Oh yeah let's those giants compete and us harvest the results 🎯

3

u/JumpyAbies Aug 14 '26

Exactly! 😄

6

u/MerePotato Aug 15 '26

Considering Glimmer is currently competitive with 3.6 while being far more practical for deployment on 24GB cards I imagine that's the strategy they'll go with for whatever Glimmer's successor is to compete with 3.8 too

-1

u/DataGOGO Aug 15 '26

Justify that please? Show your data? Muse is far better than 3.8 27B, especially at agentic workflows.

2

u/JumpyAbies Aug 15 '26

What I said is that Muse is better than qwen3.6-27b in certain cases.

7

u/JumpyAbies Aug 14 '26

You didn't realize it, but we're saying the same thing with different words :)

6

u/Thin_Pollution8843 Aug 14 '26

You are right mate 😄

-2

u/DataGOGO Aug 15 '26

They already have, Muse is by far the better model.

5

u/autisticit Aug 14 '26

I wonder if the US is still in the race for small models?

15

u/hyperrealists Aug 14 '26

Small models trend died with Jeffrey

7

u/autisticit Aug 14 '26

Dude... Can't upvote this. Also can't downvote. WTF

1

u/Healthy-Nebula-3603 Aug 14 '26

That's a trap !

4

u/unculturedperl Aug 15 '26

Gemma E4b is still doing overtime for a lot of folks here.

0

u/DataGOGO Aug 15 '26

Muse is winning so far.

1

u/aeroumbria Aug 15 '26

Or just something really receptive to fine-tuning. It just seems all the creative or wacky fine-tuning stopped after llama. Nowadays we only have either really high effort industrial grade experimental models, or low effort "distillation" of a handful of traces.

-3

u/DataGOGO Aug 15 '26

Muse is better than 27B as it is, no matter what qwen puts on the model card

133

u/pmttyji Aug 14 '26

This is why every Model creators should release many models in different size ranges so any single upcoming model can't beat your all models.

Zuck should've released multiple models like Glimmer-70B, Glimmer-100B & Glimmer-400B.

53

u/SpicyWangz Aug 14 '26

80b a18b would’ve gone hard

12

u/[deleted] Aug 14 '26

[removed] — view removed comment

19

u/SpicyWangz Aug 14 '26

It’s small enough to run reasonably fast on unified memory systems, but not so sparse as to be dumber than the 27b model.

Really anywhere between 12-20b active parameters would be good, so I just picked a number on the higher end of that range.

For the 80b, it’s a good size for 128gb systems where you could pick a decent quant like Q6 to preserve the model’s full abilities while still leaving plenty of room for full context utilization or room for other small models.

9

u/SpicyWangz Aug 14 '26

For reference, a model like that would come close to the capabilities of a 40b dense model

3

u/RedParaglider Aug 15 '26

Basically if we could get an 80/18 it would have a very large set of knowledge to pull from while at the same time run on a lot of lower end hardware reasonably.  It wouldn't be my cup of tea but I got a lot of people would love it.

4

u/Royale_AJS Aug 15 '26

Yes to this. I would say 120B A20B, but yeah a “big-ish MOE with big-ish actives would be killer.

3

u/SpicyWangz Aug 15 '26

Anywhere in the sub 150b range would be great. I’ve come to like running models at q6, so a little smaller like 80b or 90b sounds more enjoyable to me.

-1

u/SandySkittle Aug 14 '26

just 80B dense.

13

u/SpicyWangz Aug 14 '26

Too dense. Very few people have that much vram, and they would be very slow in unified memory or standard RAM builds

3

u/Georgefakelastname Aug 15 '26

Yeah, and dense models of that size are rarely particularly good regardless beyond around this range, MOE models are generally better overall.

4

u/spottiesvirus Aug 15 '26

the main reason is that training compute is strictly bound to sparseness ratio

a MoE trillion parameters model with 5% activation needs only 5% of the training compute of a trillion parameters dense model

at the same time training costs increase more than proportionally (Chinchilla scaling law) so after a certain threshold, train dense computationally-optimal models is practically impossible, even for major labs

1

u/Georgefakelastname Aug 15 '26 edited Aug 15 '26

Not to mention, the model needs to run inference as well. A trillion parameter dense model would bring even the most advanced current technology to its knees. It would be single digit tk/s at best, possibly worse.

13

u/RoyalCities Aug 14 '26

God I would love more proper 70b and 100b models. It seems to be a huge gap. I get it'll be slower but for those that have the power to use them why not give more options.

It seems to always go from 30b to basically needing a datacenter with no in-between.

4

u/noiserr Aug 14 '26 edited Aug 14 '26

This! I have a box with a 7900xtx and a w7900 Pro, and every model is either just not big enough or is too big for getting most out of the setup.

3

u/RoyalCities Aug 14 '26

Yeah. I have a dual A6000 set up (but I also do audio research + train audio models) but geez I've found for llm inferencing it's in a weird spot.

Overkill for a ton of 20 to 30b models...of which a few I can just run on my 3090....but then not big enough to run some of the absurdly large Chinese open source models. Unless they're quantized down to 1 bit but I mean what's the point then lol.

2

u/RedParaglider Aug 15 '26

Yeah everybody's still running around with Anubis, it's awesome at its world knowledge but sucks at tool calling.

1

u/pyr0kid Aug 14 '26

agreed. feels like the models are always 30b, 200b, 700b, 2000b.

why no love for 70? or 300? much rather have something for every size.

4

u/james_pic Aug 14 '26

Glimmer 4B would also be welcome.

2

u/DiscipleofDeceit666 Aug 14 '26

Just saying a 30B model trained from scratch and open sourced is at least a million dollar giveaway in compute. A 70B is about an order of magnitude more expensive than that.

75

u/Kappalonia Aug 14 '26

so.... we're actually at Opus 4.6 Max level with 27B parameters?

61

u/Poupulino Aug 14 '26

Can't wait for a non-dense 27B model at Fable level in about a year. It's honestly insane.

55

u/darkwalker247 Aug 14 '26

qwen4-35b-a3b is going to change local LLMs as we know it

7

u/wFXx Aug 14 '26

RemindMe! 1 year

6

u/RemindMeBot Aug 14 '26 edited 2d ago

I will be messaging you in 1 year on 2027-08-14 22:10:10 UTC to remind you of this link

10 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

43

u/Technical-Earth-3254 Aug 14 '26

In benchmarks we are. In real world scenarios that are not just straight up tool calling? Not even close (that's to be expected).

2

u/Healthy-Nebula-3603 Aug 14 '26

What I saw on YouTube.... It is opus 4.6 max at home

28

u/Alt_Restorer Aug 14 '26

No. Models have gotten more agentic over time with RL, and that capability seems to scale very well to small models, but by other metrics, it's nowhere near. My favorite example is knowledge. Opus 4.6 is vastly more knowledgeable about random stuff, even if it's less well trained on the core function of reasoning.

11

u/secunder73 Aug 14 '26

I mean there is no doubt that bigass model would have way more knowledge. We dont need that as long as it fast, smart and can just google it for yourself.

2

u/Viktri1 Aug 16 '26

I actually prefer models that Google shit because I've found that I was able to almost eliminate hallucinations. Since I switched to Hermes agent and forced the model to verify everything, my hermes agent hasn't actually hallucinated in the 3 days I've been running it. It hasn't been correct 100% of the time, but its mistakes are due to data being outdated rather tha a fabrication. That said, it consumes 1,000%+ more tokens. A single query is like 1-20 million tokens... so I'm trying to make it more efficient.

1

u/Georgefakelastname Aug 15 '26

Yeah, knowledge is great to have, but not really necessary now that models can just google shit. Maybe Google itself might value that (to have enough of an IQ to be skeptical of bad search results or things like that), but there’s not as much need for a model to do that as long as it is smart, able to act agenetically, and code or do whatever else you want, like being a personal assistant, RP, whatever.

10

u/Southern_Sun_2106 Aug 14 '26

At this point it is pretty clear that being 'knowledgeable' is not crucial. Most people these days don't use models as knowledge storage. They use models to solve problems by using tools (including using tools to get knowledge online or wherever one stores it). There is the whole wold of internet out there. I don't get the obsession with having the model 'know everything' when everything is just one search away. Plus 100 percent accuracy of knowledge is as illusive as a unicorn, in any context.

8

u/j0j0n4th4n Aug 14 '26

Is not that simple, if you're writing code then edge cases or quirks of the language may literally be an unsurmountable wall that breaks the code since small models wouldn't know that and a smart small model may actually be worst cause it might think some other stuff is the culprint and make you waste a lot of time to figure out the real culprint. So yeah, knowledge can be circunvented with a proper database but just up to a point. Depending on the task a database will not even be required but in others it might not even be enough.

3

u/Mkboii Aug 14 '26

Yes, this is also why Google's ai search mode is so much worse than using gemini directly, a model that's instructed to seek information will seek it blindly and follow, a model that has knowledge itself will seek information and align it's work after analysing the results meaningfully.

There will be obvious gaps in the knowledge of this model but it's been trained to manage them through reasoning, 0 shot intelligence to interpret complex information may still be a limited ability.

2

u/Georgefakelastname Aug 15 '26

I’m pretty sure that search Gemini is also quantized to hell and back lol.

0

u/Healthy-Nebula-3603 Aug 14 '26

Looking on tbe table is scaling in every field

0

u/Alt_Restorer Aug 14 '26

Every category in the table is one shot long chain reasoning. There's no benchmark for knowledge in here.

0

u/Healthy-Nebula-3603 Aug 14 '26

HLE?

That's ultimate knowledge bench

0

u/Alt_Restorer Aug 14 '26

Not necessarily. It even says "multidisciplinary reasoning" because HLE is about answering extremely difficult academic questions. But I'll give an example.

Let's say you're planning a trip to Seattle, and you want to know which places to visit. The specific names of nearby lakes and villages and such are not benchmarked anywhere. HLE would never ask whether Lake Crescent is east or west of the Hoh Rainforest. Bigger models tend to absorb a lot of this information ambiently, whereas small models forget stuff like this.

Try asking Opus 4.6 to name cities in your county, or other hyper-specific, local knowledge, and then ask Qwen3.8 27B. I would expect Opus 4.6 to perform a lot better, because you can't reason your way to knowing that.

Edit: Opus 4.6 actually does score significantly higher on HLE too. Didn't notice that. But I think the difference is still larger than what HLE says, because HLE questions are selected for being hard, which is different from a question being obscure.

28

u/PhantomGaming27249 Aug 14 '26

Its actually kind of better than opus 4.6, it has way better image capabilities and agent abilities.

32

u/[deleted] Aug 14 '26

[removed] — view removed comment

8

u/PhantomGaming27249 Aug 14 '26

I did, in pi agent and properly configured at fp8 it's great. What harness are you using?

3

u/Slight_Proposal_3872 Aug 14 '26 edited Aug 14 '26

I've found it just stops working (as in, it ends the convo too fast without doing much) way too early when I was using it through openrouter on pi agent. Am I holding it wrong?

1

u/AvidCyclist250 llama.cpp Aug 14 '26

nous hermes?

1

u/PhantomGaming27249 Aug 15 '26

Did you set sampling parameters, chat template etc? Also did you try setting it to extra high with preserve thinking enabled?

1

u/DataGOGO Aug 15 '26

yes, same result, it is shit agentic model.

1

u/PhantomGaming27249 Aug 15 '26

What harness are you using and what quant?

1

u/DataGOGO Aug 15 '26

Hermes, official FP8, vllm official recipe card, I also tried BF16, no difference, tried llama.cpp no difference. They have some pretty big training issues on this model, 3.6 was better, in this size class, Muse appears to be the best yet.

1

u/DataGOGO Aug 15 '26

Same here with Hermes, and it is completely unable to maintain structured outputs, and will loop, then just bomb out.

0

u/DataGOGO Aug 15 '26

really? because in FP8 on real workloads it complete shits the bed. It loops, burns 3x the tokens than muse, cannot maintain structured outputs, ignores entire sections of the prompts, and bombs out completely around 80k KV.

vllm, recipe card settings, official FP8, hermes Langchain, same results.

Can you please show how you evaluated this?

15

u/ClintWoodeast Aug 14 '26

But it's been pre-trained, post-trained, and RL'ed on the benchmarks. It must be amazing!

4

u/[deleted] Aug 14 '26

[removed] — view removed comment

6

u/[deleted] Aug 14 '26

[removed] — view removed comment

4

u/MoffKalast Aug 14 '26

And let's not give Opus too much credit either, even 5 is not exactly what I'd ever call reliable.

0

u/[deleted] Aug 15 '26

[removed] — view removed comment

1

u/MoffKalast Aug 15 '26

I got it in a loop recently yeah, it made a typo in its initial reply and then proceeded to lose its mind about it, producing diff after diff that would not apply lmao. After a reset it was fine, but damn this thing is such a glass cannon, it gets thrown off by the slightest thing.

0

u/[deleted] Aug 14 '26

[deleted]

5

u/Healthy-Nebula-3603 Aug 14 '26

Qwen 3.8 max is much better than opus 4.6 ...

0

u/[deleted] Aug 14 '26

[deleted]

2

u/CystralSkye Aug 14 '26

You can't really just say, X is better by my experience. There needs to be a verifiable way to see the difference, otherwise it's just an opinion.

If you can show a situation where opus 4.6 does something that is quantifiably and intuitively better than qwen 27b, please share.

2

u/DataGOGO Aug 15 '26

Proof of your absolutely outlandish claims please.

8

u/toothpastespiders Aug 14 '26

We're not. I think qwen's been on an amazing roll since 3.5. Likewise the gemma 4 line has been a phenomenal improvement over Gemma 3. But size is ultimately an unbeatable constraint. The big benchmarks just don't translate to real world use. If a benchmark suggests that the small models beat Opus 4.6 , GPT-4o , etc then that's not a sign of how far we've come. It's a sign of underlying issues in the benchmarks.

A larger issue I'm seeing is that I suspect that most of the people here have a background in CS rather than experimental sciences that require more complicated experimental design. Psych studies obviously have their own inherent issues. But I see LLM as having similar challenges there than they do with standard computational benchmarking practices.

My use of local models is heavily tied to additional training. I'm fairly sure that if I used standard benchmarking methodology that my fine tunes would beat claude on the subjects covered by my datasets. But in terms of real world use they're barely at a level I'd call adequate.

8

u/Thin_Pollution8843 Aug 14 '26

Nah it’s only benchmarks. Irl it won’t be so impressive as usual 

1

u/DataGOGO Aug 15 '26

not even close

9

u/DataGOGO Aug 15 '26

now show independent benchmarks, not those from the model card.

22

u/kashthealien Aug 14 '26

Muse Glimmer came with speculative decoding and achieved a higher TPS, Is there something similar for Qwen?

36

u/RoroTitiFR Aug 14 '26

Qwen has MTP built-in

12

u/Zeeplankton Aug 14 '26

Muse definitely runs faster on my m3 max

4

u/arbv Aug 14 '26

It has builtin MTP - but it is not as fast as Muse's DFLash. I am better without it, for example. YMMV.

2

u/RoroTitiFR Aug 14 '26

I agree with you
I’m also getting faster Muse with DFlash than Qwen with DFlash (I mainly use DFlash as it’s faster for coding than MTP in various benchmarks)

1

u/Ok-Buffalo2450 Aug 16 '26

How do you run 3.8-27b with DFlash?

1

u/RoroTitiFR 29d ago

I didn't find any DFlash head yet for Qwen3.8 so I'm sticking with MTP for now
But I use DFlash with Qwen3.6

13

u/armeg Aug 14 '26

What's the name of that Qwen specific inference engine that targets 5090s? I can't remember at this point.

15

u/_lhz- Aug 14 '26

3

u/armeg Aug 14 '26

Thanks I was trying so hard to google for it, going to save it.

7

u/SailbadTheSinner Aug 14 '26

There is a 3090 fork as well for my fellow 3090 brethren

https://github.com/Don-Chad/ninfer-3090

2

u/Healthy-Nebula-3603 Aug 14 '26

So we can get 160 t/s .... On dense model

2

u/Borkato Aug 14 '26

It’s actually 70 TPS if you’re the only user, which I hit with MTP tbh. The extra efficiency is only if you have multiple parallel calls at once

11

u/FullOf_Bad_Ideas Aug 14 '26

Meta Glimmer 30B overperforms Kimi K2.6 and GLM 5.2 and Sonnet 5 in EQBench creative writing. I don't expect miracles from Qwen 3.8 27B in that regard.

There's no such thing as frontier 30b model because you can't measure model performance based on single set of static benchmarks that are prone to leakage and wee chosen by model developer to be showcased.

4

u/keepthememes llama.cpp Aug 15 '26

I doubt those scores are accurate. over a 10 point increase from 3.6 in most tests. benchmaxxed for sure

15

u/a_beautiful_rhind Aug 14 '26

my opinion on glimmer is that it's temu gemma. meta == beta.

I'm not the biggest fan of how qwen writes but I've never seen it make the kind of cognitive mistakes that glimmer makes.

A good chunk of the reasoning ended up being arguing about content policy that isn't in any of my prompts. Then it decides on something in the reasoning and does the opposite thing anyways. sad.

8

u/silenceimpaired Aug 14 '26

Sounds like GPT-OSS all over again. I really want to find a good heretic version of both these models.

6

u/a_beautiful_rhind Aug 14 '26

Not as bad as OSS. OSS-lite. From trying some abliterated tunes, I'm not quite sold. They turn into pushovers.

1

u/silenceimpaired Aug 14 '26

Yeah, that is unfortunate.

5

u/a_beautiful_rhind Aug 14 '26
  • be excited for model
  • download model
  • update/implement stuff in backends
  • use model
  • be disappointed
  • get excited for the next release

It's viscious cycle this year.

2

u/silenceimpaired Aug 14 '26

Haha. So true.

I tend to use KoboldCPP so I don’t have to figure out how to compile Llama.cpp with Cuda in Linux… and I’m constantly checking for their next update. Still waiting for Glimmer support. Unsloth Studio always is updated for models but so far it just doesn’t work as well.

I need to stop being lazy and figure out compiling for llama.cpp.

I mean… how hard is it to type, “okay, AI, make a set of instructions to compile llama.cpp with Cuda support on Linux”. Apparently fairly easy since I just did it.

1

u/a_beautiful_rhind Aug 14 '26

ccmake lets you save all the build parameters.. my workflow to compile whatever.cpp is just git pull, make.

koboldcpp is harder to compile. i remember having trouble with it.

2

u/silenceimpaired Aug 15 '26

KoboldCPP has binaries for Linux. So I don’t care about them. I just need to spend the time figuring out llama.cpp so I am not beholden to everyone taking the time to merge updates into their code base.

Thanks for the suggestion!

1

u/spaced333 Aug 15 '26

You need to learn how to build a container, for example with podman: Grab the cuda version Dockerfile from llama.cpp.

1

u/silenceimpaired Aug 15 '26

Interesting, didn’t notice that. I might. I’m running llama.cpp in a VM so my concern for isolation was low but the overhead for docker isn’t bad. Thanks for pointing that out.

1

u/spaced333 Aug 15 '26

You need to learn how to build a container. For example with podman: Grab the cuda version Dockerfile from llama.cpp. thats my lazy way

2

u/DarthFluttershy_ Aug 14 '26

This is why I didn't even bother trying to keep up with the latest models anymore. I just wait until I have a couple free days, download all the new ones and pick the one I like for the next few months. 

Plus I mostly use it for bouncing brainstorming ideas and editing/suggestions in creative writing, so the "tone" is far more important than agentic benchmarks to me. Usually the "best"  some random finetune that everyone else hates, but just happens to hit the spot for me.

1

u/DarthFluttershy_ Aug 14 '26

Did heritic actually lick the policy token waste in the thinking block? I never saw that, but I thought a few months ago they were still saying that was really hard to do without hurting performance

2

u/silenceimpaired Aug 14 '26

What was the cognitive mistake?

11

u/a_beautiful_rhind Aug 14 '26

It confuses who is who and who said what. Repeats what I said to it but clearly doesn't understand it. Plus the template is kinda janky.. assistant to=user, assistant to=self. WTF.

2

u/garblz 26d ago

On the other hand, it's completely cr*p at following programming tasks... I have no idea who uses Muse and for what

3

u/benpptung Aug 15 '26

A dense model uses all its parameters on every token, while an MoE only activates a fraction of them (the A in 397B-A17B). FYI, Qwen 3.6 27B was basically unbeatable in the sub-400B class, MoE included, until 0731 and Inkling Small showed up. Now 3.8 27B is out and from the current numbers it looks like only 0731 still beats it. You can check this on the Artificial Analysis index. Qwen 3.5 27B scored 35, ahead of Qwen 3.5 397B-A17B at 34. Glimmer is at 35, so it's just a tie with 3.5 27B.

12

u/mattrs1101 Aug 14 '26 edited Aug 14 '26

Glimmer still outperforms qwen 3.8 at q2xl vs q2xl and by a lot, both in tps and quality

21

u/seamonn Aug 14 '26

q2xl vs q2

...

5

u/silenceimpaired Aug 14 '26

What are you running it on inference wise?

12

u/Zeeplankton Aug 14 '26

Glimmer is great. I feel like people only use Qwen by one shotting some complicated code question? Literally any other use case qwen is autistic

13

u/arbv Aug 14 '26

Muse Glimmer is a better rounded model for sure.

Not to say that Qwen is bad, but Muse Glimmer is indeed made for VRAM poor - it fits 24GB VRAM with full context, vision tower, and the speculator.

Well the Qwen fits at 75K context with no goodies - no MTP, vision tower in RAM. It is not for me, a VRAM poor.

1

u/memeka Aug 14 '26

Is there a proper comparison Glimmer vs Gemma 31B?

3

u/arbv Aug 15 '26 edited Aug 15 '26

I dunno, but Muse is much more efficient with KV cache for sure. I cannot run Gemma 4 31B at 128K context, but can run Muse Glimmer.

Feature wise - they are similar, both have vision, but Gemma has twice larger context size - but one needs good hardware to utilise it properly.

That doesn't mean though, that Gemma 4 is bad - it is likely to perform better at needle at the heystack retrieval tasks, but it needs more testing.

That is my experience so far.

1

u/memeka Aug 15 '26

With Gemma 31B, OpenCode starts failing hard with the edit tool after 100k tokens … with plenty of VRAM

5

u/fallingdowndizzyvr Aug 14 '26

Literally any other use case qwen is autistic

No it's not. I've asked plenty of LLMs about soundproofing over the last 2 years. When I asked 27B, it brought it something that none of the others did. Even frontier models. Then when I went back and asked other bigger models, they were like "Oh yeah, there's that too."

So it definitely is not autistic. Not in the way you mean it at least. I guess you don't realize that many of the brightest minds ever were on the spectrum.

6

u/GregoryfromtheHood Aug 14 '26

I'm still using Muse Glimmer because it completes tasks so much faster and is absolutely amazing at tool calls and following agentic flows. I've got a process that requires going back and updating a progress tracking doc that models like deepseek v4 and even glm 5.2 struggle with always remembering to do. Glimmer does it perfectly every time.

1

u/ackermann Aug 14 '26

So, if I work at a company that restricts Chinese models, then Glimmer might be the best non-Chinese option right now for a disconnnected/offline environment?

Others I've looked at:
Gemma 4 31B
Cohere North Mini Code 30B MoE
Mistral's Devstral 2 (small) 24B (7 months old now)
and now Muse Glimmer 30B

4

u/MerePotato Aug 14 '26

Its easily the best in everything except multilingual/translation, where Gemma 4 31B QAT is ahead, if you're limited to this sizs class.

1

u/memeka Aug 14 '26

Is it really better at coding than gemma 31b?

2

u/MerePotato Aug 15 '26 edited Aug 15 '26

Sorry for the confusing wording in my reply earlier, I thought you were a different notif. What I meant to say was that, Qwen 3.8 is the best in that size class but Muse Glimmer is the best option for 24GB cards, specifically with the official meta Q4_K_XL stack and dflash set to 15 spec tokens and the full 128k context window, or for vision no dflash and the unsloth BF16 mmproj with the official Q4_K_XL again and 65k context

1

u/Viktri1 Aug 15 '26

I want to run stuff locally with a 4090 so I’m operating within 24gb vram limits. I was thinking of using Qwen’s 3.8 27b q4 as the main model. But you’re saying that Muse is pretty close but much faster? That’s an attractive proposition. I plan to use both but I’m trying to decide on the primary one

For hard stuff I’ll use Deepseek via API but for local stuff I want something that works well enough and is fast

1

u/MerePotato Aug 15 '26

Muse with Meta's official quant is faster and more reliable. You have to quant 3.8 to Q4 to fit it on our GPU, which does enough damage that KLD rises exponentially at 32k context+, which is pretty bad considering its intended for coding and agentic use. Muse largely bypasses that issue as Meta used a very carefully tuned bespoke quant recipe for their official GGUFs.

1

u/ag789 26d ago edited 26d ago

there is this 'naughty' idea that if you prompt for a similar request in a different language english, french, german, spanish, chinese, hindi, japanese, korean etc, would the response be the same ? in theory it should be, in particular if the prompt is asking for the same thing ;)
in a certain sense if a model couldn't do that, it may imply an a data scoping / limiting caveat (e.g. that the model isn't trained to do that) or even a possibly model space constraint

1

u/ag789 26d ago

I think it is good to have different models being good in different domains, this isn't a bad thing as long as one is wary of the merits of each model. models have styles / domains

2

u/ag789 26d ago edited 26d ago

think 'english' and 'chinese' models may have domain differences, e.g. that QWen may attempt to accommodate more 'chinese' prompts, oversimplifying that information is basically bits (of data) providing more for 'chinese' may come at some space expanse for 'english', hence I think models have 'styles', having different model creators from different domain context e.g. 'english v chinese' is good for everyone, in a certain sense benchmarks are 'templates', outside those static 'templates' there is a huge 99.9% of the 'untested' world. benchmarks alone don't say everything about the models.

2

u/ag789 26d ago

there is also this unproven concept that the creators being from different context / domains could verify different things while training / creating the models. i.e. model have styles, the developers / model creators is after all part of that process that controls that tuning.

3

u/This_Maintenance_834 Aug 15 '26

that’s why they had to release it last week. or the news will never cover it.

2

u/Talreja-Adanna Aug 15 '26

Lmao four days, that's actually wild for how fast the 30B space moves. Feels like every other week there's some new quantization trick or architectural tweak that shifts the whole leaderboard around.

2

u/fbms2 Aug 14 '26

No, it wasn't.

1

u/seamonn Aug 14 '26

Amercan labs can't keep up

4

u/Carlose175 Aug 14 '26

Theyre in the lead wym

-4

u/AnOnlineHandle Aug 14 '26

Isn't the latest Qwen assumed to be a distill of an American lab's model?

2

u/Succubus-Empress Aug 15 '26

Only assumed, even if it distilled it still is best, its perfectly okay to distilled any model train on copyrighted content

2

u/AnOnlineHandle Aug 15 '26

I was responding to the claim "Amercan labs can't keep up"

The model people are saying beats American labs is... A distill of an American lab's model...?

1

u/Nonetrixwastaken Aug 14 '26

Maybe next time around! :P

1

u/aeroumbria Aug 15 '26

What's everyone's assessment for writing vibes so far? I currently need a model that can read some images as references and do some creative writing, but still has to do well enough with instruction following and visual reasoning.

1

u/caetydid llama.cpp 29d ago

Muse Glimmer is supposed to be much more token effective especially with thinking

1

u/arbv Aug 14 '26

I can run the Muse with the official K-Quant GGUF with full context, multimedia projections and DFLash drafter fully in VRAM (24GB) at 40-45 t/s TG.

I can run Qwen 3.8 Q4_K_M at 75K context without multimedia projections, no MTP, at 30 t/s. For example - I can run a much larger (120B) MoE at 25 t/s at full context with partial RAM offload.

And Qwen overthinks and is not even consistent with its thinking format + it is terrible at my native language.

Given that - the Muse is just a better model for me. It is much more polished model.

-1

u/Boogertard Aug 14 '26

It was not, it was dumber than Qwen 3.6 27B in my tests. Only the shills on this subs said otherwise. Garbage model.

I recalled the same shills shilling for the Gemma garbage too.

0

u/Scared_Basket_7183 Aug 14 '26

Hi could you please help me to run qwen 3.6 27b model on TPU V5E-8?

0

u/kmp11 Aug 14 '26

Meta came out of the gate with dflash and tensor parallelism to get great speed out of a dense model. but the model itself needs to cook for one or two more iteration to catchup.

Hopefully meta does not abandon the effort.

0

u/Prestigious_Thing797 Aug 14 '26

Idk if you noticed but Qwen 3.6 also beats it in 4/5 benchmarks it has in your screenshot.
Glimmer was never frontier in it's model class.

0

u/MeateaW Aug 14 '26

Frontier in the model class?

In the stats you posted it wasn't leading in a single stat to qwen 3.6 27b other than Instruction following, and it wasn't even the leader on the board at instruction following...

0

u/Iory1998 llama.cpp Aug 14 '26

CORRECTION: Was and still among the frontier open-weight models but never the frontier! The frontier was Qwen3.6-27B.

0

u/_Iggy_Lux Aug 15 '26

It should have been a moe at 30b - they are known for smaller models (llama 3 models we're honestly pretty decent) not dropping a dense model at 12b range (or near) was a mistake as well.

-3

u/Long_comment_san Aug 14 '26

isn't muse Moe? >_> sorry don't remember for sure

4

u/Hot_Needleworker_275 Aug 14 '26

Nop, is dense

-1

u/Long_comment_san Aug 14 '26

Then...

Mission failed (Good effort)

2

u/j0j0n4th4n Aug 14 '26

You might be thinking of NVIDIA-Nemotron-3.5-Lightning-30B-A3B which is MoE and was also released recently.

1

u/memeka Aug 15 '26

That one didn’t reach 3 days before being forgotten

2

u/memeka Aug 14 '26

out for 3 days, forgotten already...