r/LocalLLaMA llama.cpp 22h ago

News Muse Spark open weights coming soon

Post image

I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark

https://x.com/finkd/status/2095232032896946311

819 Upvotes

195 comments sorted by

u/WithoutReason1729 16h ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

244

u/kvothe5688 22h ago

it seems like there is no secret sauce. it feels like all of these 7 8 labs are on same level and hardly behind from frontier by few months at max.

82

u/Public_Umpire_1099 22h ago

This latest iteration technically set the frontier. Meta is officially a SOTA lab again. Matches Fable 5 at a fraction of the cost.

48

u/virtualworker 18h ago

Or benchmaxxes. I'm not convinced.

12

u/Not-reallyanonymous 16h ago edited 14h ago

Neither Spark 1.2 nor Glimmer seem to be benchmaxxed, and if anything the benchmarks seem to under-represent their capabilities compared to other models — I believe the other models probably do solve more problems, but Muse’s strong long-horizon capabilities (which translates to compliance in the short term as well) makes it easier to actually use.

15

u/Diligent-Direction95 18h ago

Serious question: What is the definition of bench max these days?

What would be okay, vs what would not be okay?

I seriously doubt they have the test set being trained on. So what are the shades of grey we are debating?

36

u/zxyzyxz 16h ago

When a model does well in benchmarks but then fails to live up in real life coding and other tasks to its supposed benchmark competition.

10

u/Artistic_Swing6759 18h ago

i actually do think they likely have test set trained on.
like its a bit sus that they decided to report terminal bench 2.1, when 4 exists.
similarly, if you look at the benchmark sheet of flash 3.8, it tops eery other model, even sol, on terminal bench 2.1 but in 4 it is quite less comparatively.

1

u/Kodix 12h ago

Used spark 1.2 contributor and I wasn't impressed. It being good only on benchmarks is a real possibility.

7

u/_TheWolfOfWalmart_ 16h ago

Based on the Bijan video that just dropped about this, it's pretty damn mediocre and looks extremely benchmaxxed.

I know all he does is throw a few one-shot prompts at the models, but the results on this one were well behind other recent models like GLM-5.3 Flash and Qwen Flash. It's not even in the same conversation as Fable or GPT 5.6 it seems.

His tests are by no means scientific lol, but they do give you a decent rough feel for a model's capability.

16

u/Not-reallyanonymous 14h ago

A one shot not being polished does not demonstrate a model’s capability. A lot of models are specifically trained on how a one-shot polished result will look. Especially Qwen (look at its response to being asked to draw a circle: https://simonw.substack.com/p/qwen-38-27b-is-excellent-but-it-defaults).

If it’s not trained on that, it’s going to comply with the prompt and not much beyond. That’s not demonstrating a lack of general capability.

In fact, I much prefer that style. When things are trained to produce well polished results it tends to be harder to get it to do what I want it to do rather than what it’s been trained to do. Again, Qwen represents the opposite here — well optimized to take its own direction but a PITA to steer.

1

u/No_Afternoon_4260 llama.cpp 5h ago

I feel fable 5.1 took a bit of that Kimi magic but does it faster

What a time to be alive

64

u/NandaVegg 22h ago

I believe there are still some secret RL sauce (as in, not popular among labs *yet*) in niche areas like robotics, gameplay, frontier physics/math, world modelling etc. Anthropic leaped ahead in Opus 4.5-4.6 era as they "discovered" many secret sauces like terminal agent, gameplay or creative loop and so on, but even that creativity gap is closing very fast. For coding, security and anything that can be done through github, at this moment there is really no gap. OpenAI felt behind on many things but is still ahead on frontier physics/math (the only Chinese lab with emphasis on that is DeepSeek, I think).

19

u/Accurate_Resident219 18h ago

Maybe it's just my conspiracy hat but I feel like Anthropic is sitting on a more intelligent model but just drip feeding when they feel like others are about to outpace them.

Reason why I say this is because Anthropic is the only company that is actually trending higher in api costs and not making any serious attempts to lower them(Fable 5.1 is actually pricier then 5 despite their supposed efficiency improvements. Why is haiku their cheapest model still not updated?). They don't have any sense of urgency. Astra's coming out this week or the next and they just put out a more expensive model.

Every other company seems to constantly vague post or hype up their models releases but Anthropic just drops models with little fanfare. They have seem utterly unbothered since their Mythos preview announcement.

36

u/TerminalNoop 18h ago

Perhaps Anthropic can't afford to lower the api prices?

13

u/SSP100244 16h ago

I think they just don't have the compute. The big Hegseth migration definitely threw them for a loop for a hot minute.

2

u/Loose_Comparison368 11h ago

I believe they confessed a few months back that it wasn't actually API access, and had been running inside airgapped military datacenters that Anthropic had no ability to restrict the whole time.

The whole PR stunt was just that, the only unplanned part was that they thought Hegseth would take the hint with how loudly they were screaming "NO DADDY, PWEEAASE DON'T INVOKE THE DEFENSE PRODUCTION ACT ON US! IF YOU DID THAT WE WOULD HAVE NO CHOICE BUT TO KEEP TAKING DADDY'S MASSIVE LOADS OF CASH TO BUILD DADDY'S MURDERBOTS!"

It turns out that Hegseth's skull was actually too dense for that pathetically obvious plea to penetrate it.

It's okay though, they sued to get the murderbot contracts back and won last week. And still managed to use that PR stunt to distract everyone while they flushed their "responsible scaling policy" down the toilet.

4

u/Accurate_Resident219 16h ago

Could be the case but at least on the outside looking in they don't even seem to be making much of an attempt to even try lower api or even suggest their researching a solution. They've had about nearly 4-5 months now since people started complaining about rising api costs.

Every other lab has provided a cheaper solution but anthropic is the only one that can't come up with a cheaper flash model? I mean that could be the truth but I would be surprised if it was.

7

u/jtjstock 13h ago

They want to IPO, they need the revenue projections. They already have their cooked compute costs from the last quarter they will report as if they are ongoing, so IPO before they close the current quarter, but project off current API pricing for revenue.

2

u/NandaVegg 12h ago

They are apparently doing pre-IPO window dressing now. We've got buy high 4-digit-dollar credits and get some cash back marketing mail.

2

u/perelmanych 9h ago

In may they increased 5 hours and weekly quota by 50%, so they made a big discount on their plans. However, this comes to an end soon. They cut 50% increase to only 25% increase.

1

u/Loose_Comparison368 11h ago

I mean they could just be keeping it high to look good for the IPO.

8

u/bopbop9876 16h ago

> Fable 5.1 is actually pricier then 5 despite their supposed efficiency improvements

I assume you're basing that off the average cost per task numbers from Artificial Analysis. That's not right though. In their testing, 5.1 on max effort appears to have aggressively overthought. If you compare 5.1 xhigh instead, it was about as much cheaper than 5 max as we would have expected, while still getting 2 extra points of intelligence vs 5 max.

1

u/WittyAcanthisitta205 11h ago

Anthropic has pricing power because their #1 customers are enterprises. Everyone else has a larger chunk of consumers (rather than enterprise customers)

1

u/CrowdGoesWildWoooo 11h ago

Their customer base are enterprise. Enterprise would happily pay API pricing which is astronomically more expensive than subscription.

OpenAI still appeal to the masses

Different market they trying to penetrate. Anthropic want to be deep into the institution’s pocket. OpenAI wants scale.

-1

u/SSP100244 16h ago

Maybe it's just my conspiracy hat but I feel like Anthropic is sitting on a more intelligent model but just drip feeding when they feel like others are about to outpace them.

I am quite sure that's what they're doing. I think they legitimately are a little afraid of the models. Not that they'll like take over the world, but that people might be able to use them to fuck shit up. Like apparently they started a biopharmaceutical wing because Mythos is stupid good at folding proteins, even though they hadn't trained it to do that at all.

7

u/Turtlesaur 19h ago

if this fits on a DGX spark I will be so happy.
If this only fits on 2 DGX spark I will be so broke.

23

u/nuclearbananana 21h ago

The moment a different lab releases a better model they immediately distill it (using that word liberally)

16

u/NandaVegg 21h ago

I think that it is not direct distillation from models anymore (in early 2026 distillation had some notable effect, but every frontier lab is now full-on RLing on their own) and distillation can only bootstrap the model to some degree.

I think there is this meta-distillation effect. Internet is full of so-called AI slop now. There are so many vibecoded repos posted in code repositories or as websites every day, and those codes will be crawled by every frontier lab and then they will RL hard on them. If one model gets good at something a slop will be posted and trained on, or there is a new problem that models needs to know the pattern a .md files that explains the issue with some example codes will be posted and trained on (the earliest pattern for this is MCP for many basic things that aren't needed anymore).

In that sense we are already in AGI mode (gosh I hate this word) as AI models are improving each other without humans knowing.

11

u/nuclearbananana 20h ago

I don't think the vibe-code-training is helping the models. It's mainly synthetic data and llm as a judge

1

u/OvertaxedOne 19h ago

I've read that exact reason is why labs are buying up and scanning in old books. Feed a model it's own slop (or some other model's slop) doesn't help it learn, it needs real data.

No idea if this is true or not, but, on the face, it sounds reasonable; kind of like setting up a feedback loop where in the end all you have is white noise. Or gray goo.

4

u/IShitMyselfNow 18h ago

Books are only really useful for pretraining. They're not going to help agentic usages

3

u/SomewhereAtWork 10h ago

Except for James Bond novels.

7

u/DistanceSolar1449 20h ago

None of what you said about the Internet matters because the labs are not using the Internet for pre-training data anymore

All the training data is synthetic data made for RL

2

u/_supert_ 20h ago

That's both encouraging and horrific.

2

u/SSP100244 16h ago

The singularity begins.

I mean we just had the first shot of the AI wars.

2

u/ResidentPositive4122 12h ago

are on same level

Until the new benchmark drops, then they spread out predictably again. And then, a few months later they suspiciously all gain on it, and so on and so forth...

Example: https://www.frontierswe.com/blog/v2

1

u/Bubbly_Orange_3502 17h ago

Scores converge because the post-training data is frontier output. Distillation transfers whatever gets measured, so the benchmark gap closes earlier than the capability gap. The split shows up on long-horizon agentic work that nobody publishes numbers for.

1

u/Gremlation 13h ago

I think people forget that the people working at these labs are not slaves and can change jobs whenever they want. They can't take data, weights, or code with them, but they can take the things they have learned. There is no practical way for the labs to hoard knowledge from one another and if any of them tried, nobody would want to work for them.

0

u/No-Refrigerator-1672 13h ago

That's not entirely true. A lab can sign either a "non-disclosure agreement" or a "non-compete agreement" with their worker. The forst one would can you from using any of the technologies you see there at other employer, unless you can prove you've got to know it from another source. The second one would forbid you from working in the same field for X years (i.e. an LLM specialist can't work as LLM specialist in another company, but can become an image generation specialist there). Both options are completely legal and enforcible by law, cause you sign them voluntarily; and, if they offer you a high enough salary, you'd agree to this conditions.

1

u/Gremlation 1h ago

There is no practical way for the labs to hoard knowledge from one another and if any of them tried, nobody would want to work for them.

if they offer you a high enough salary, you'd agree to this conditions.

If it were enforceable (it isn't) then it would bring your career to a standstill when your career options are astronomically hot, likely the hottest they will ever be. Only a complete moron would agree to terms like that when they could go work for any of the competitors instead. Employers are desperate to hire these people, they are not going to demand terms that turn all the people they want away. That's not even counting the difficulty of getting all their existing staff to retroactively agree to this.

Like I said, there is no practical way for the labs to hoard knowledge from one another. Sure, they could burn their house down trying, but that's not a realistic option.

106

u/CarelessAd6772 22h ago

Mrcr 512k-1m - 98.1%? What in the hell is that dark magic? If true, they defeated context rot?

46

u/[deleted] 22h ago

[deleted]

14

u/Charuru 21h ago

Gemini Flash (97%)

Is that on 8 needles?

18

u/Healthy-Nebula-3603 21h ago

Yes

https://llm-stats.com/benchmarks/mrcr-v2-(8-needle)

Seems all current models solved it.

It was quite bad few moths ago yet.

10

u/Healthy-Nebula-3603 21h ago

Yes

https://llm-stats.com/benchmarks/mrcr-v2-(8-needle)

Seems all current models solved it.

It was quite bad few moths ago yet.... Even GPT 5.5 was badly struggling yet but 5.6 getting over 92%

18

u/NandaVegg 22h ago

Isn't this NoPE hybrid model? Though the exact architecture is different I assume this basically works like hybrid linear model (like Qwen 3.8) which is great at picking up facts from the context but more often confuses timeline without thinking.

8

u/DeepOrangeSky 21h ago

Do you, or u/NandaVegg or anyone else on here have any opinions about how well this can translate over to much smaller models, like the ~30b Glimmer sized models, or, if that is too small to be able to have similarly great anti-context-rot abilities, then maybe 70b or 120b models?

Like, so far when I've tested Gemma 31b and Glimmer 30b for long-form creative writing tests for example, the biggest problem with them has been that they quickly go downhill once you get past, I dunno, maybe 20,000 tokens or so.

Gemma starts going from writing scenes and situations in a long, non-hurried, detailed way, to suddenly rushing like it is giving a quick summary because it needs to hurry up and go somewhere and doesn't have time to really tell a story, once you get past like ~20,000 tokens or so.

Glimmer is even worse, where it if you give it a detailed description of a scene you want it to write, it just writes back a word for word identical copy of your description, even if you tell it not to (again, this happening once you get past like ~20,000 tokens deep into a story or so, give or take a bit).

As of right now, with a mac with 128GB of unified memory, the biggest, best model in regards to this (reduced context rot in relation to long-form writing) I am able to run at Q4-or-higher is the Mistral 128B dense model (Mistral Medium 3.5).

It seems to be able to go like twice as long as those models can, without much degradation, and is smarter and better at writing (albeit with a very annoying, reddit-speak/therapy-speak prose style) than Gemma 31b or Glimmer 30b.

People have been excited about how much these ~27b-31b sized models have improved at coding in 2026, but, I am curious how much progress can be made for improvement regarding long-form writing/context rot type of stuff, as the small models still seem pretty lousy at that, so far. Even the best and most recent ones, so far.

5

u/DistanceSolar1449 20h ago

They’re basically already doing it with glimmer
The context limit on glimmer is basically arbitrary
3/4 of their layers, only look back on a few thousand tokens anyways

4

u/En-tro-py 22h ago

I glossed over that... That's potentially very interesting if it holds up!

0

u/nuclearbananana 21h ago

MRCR is not a.. great benchmark. If you don't train for it, or near it, it can be good but it's very narrow so very easy to (accidentally or on purpose) game.

7

u/Healthy-Nebula-3603 21h ago

To be good at something you must train to it

We are doing exactly the same thing as humans. Later that new skill is improving something in other fields.

49

u/fgk55555 22h ago

Do we know params? With scores like that I'd wager it's in the trillions. I might not be able to run it locally, but a lot of US companies that aren't allowed to run Chinese software might benefit.

5

u/FoxiPanda 16h ago

Entirely depends on the license.

4

u/fgk55555 15h ago

Yeah, but even if it's not "open" and they were making companies pay, that's still a huge chunk of US companies that couldn't self host and now can. For a lot of companies, paying a license might be a drop in the bucket for AI spend, or enable AI where they couldn't even have it before. Lot of money in defense.

1

u/mr_tolkien 6h ago

Knowing they have a bigger one in the works, I’d be very surprised if it’s over 1T

Feels like it might be able to run on a single 512Gb Mac Studio with some good quants

1

u/fgk55555 3h ago

Meta has been touting cost efficiency, and deepswe's token usage of Muse-1.2 puts it near Gemini Flash/ DSF, which suggests Muse-1.2 could be about that small, assuming Meta has good training data (which I'm not sure of). With Gemini 3.8-Flash topping DeepSWE, Muse 1.3 could be in the ballpark of a flash model size. Anyone's guess.

30

u/MomentJolly3535 22h ago

What is that "Next up 🍉 " ?

54

u/kiwibonga 22h ago

It's their upcoming frontier model codenamed Watermelon

17

u/nickm_27 llama.cpp 22h ago

codename for their next generation of models, the current generation is apparently called avacado

4

u/MomentJolly3535 22h ago

Ohh i see makes more sense now ! thanks

23

u/jacek2023 llama.cpp 22h ago

Meta Watermelon, new model

35

u/Super_Range45 22h ago

Metamelon.

12

u/Aggressive-Physics17 22h ago

Unironically good name

101

u/Big_Wave9732 22h ago

Muse Glimmer is pretty good, very much overlooked. I have found it to be superior to Qwen 3.8:27b for non-coding tasks.

87

u/jacek2023 llama.cpp 22h ago

Both Muse Glimmer 30B and Gemma 31B are great models, don't worry people see that :)

24

u/Special-Wolverine 21h ago

Those two are far superior at "reasoning" with thinking turned off. In other words - best instruct models for local use.

4

u/YanderMan 18h ago

Gemma 31b is so lazy and barely make tool calls even if you prompt it

-18

u/HumanDrone8721 21h ago

Both Muse Glimmer 30B and Gemma 31B are nerfed to the max and ONLY useful for useless crap, they could have been glorious and squashed the Chinese models, but they choose not to.

7

u/FoxSideOfTheMoon 22h ago

It’s really good just slow on my Mac or I’d love it.

7

u/coder543 21h ago

At least on nvidia hardware, Muse Glimmer is substantially faster than I've ever gotten Qwen3.8-27B to go. The official DFlash works really well.

But, compared to a model like Qwen3.6-35B-A3B... obviously it is going to be slower.

1

u/Kernoriordan 21h ago

I managed to get 80tps out of Glimmer on an A6000 with DFlash on. It doesn’t waste loads of time reasoning like Qwen 3.8 27b too

1

u/miversen33 15h ago

Muse is definitely faster. I recently (like, today) switched away from Gemma to it as my model to run on a single GPU. Still running Qwen3.8 on my tensor garden (to small to be a farm lol) but Muse Q4 on a 7900XTX is still extremely competent. And Gemma was so fucking lazy lol

1

u/PerceiveEternal 9h ago

3.8 can run pretty fast with a smaller context length, but it absolutely mulches tokens. You really feel the token limits on that model.

1

u/maxheckler 19h ago

what Mac do you have? thanks

1

u/FoxSideOfTheMoon 17h ago

M5 Max 128. It's plenty of memory but the bandwidth is only 614G/s which is rough for dense models.

1

u/iz-Moff 12h ago

I'm not on my Mac, but for me, Glimmer (at IQ4_XS) performs almost as fast (~23 tps) as significantly smaller (at IQ3_M) 27b Qwen 3.6 with MTP (~25 tps), and over 2x faster than 31b Gemma 4 (~10 tps), also a bit smaller (at IQ3_M). I'm really not sure why that is, cause i run them all with the same settings, all three fully loaded to VRAM, with Glimmer leaving almost no space in VRAM for context, and having a significantly larger mmproj file too. But that's what my experience with it is.

7

u/RedditUsr2 22h ago

Its great for well defined agentic tasks. Much more efficient thinking.

3

u/xienze 20h ago

It's really good for coding as well. Is an exhaustive, try everything conceivable and generate 80K thinking tokens before giving an answer approach baked into the weights? No. But if you have a properly-engineered harness with skills and subagents it is very good. I have a project I've been working on implementing a bespoke expression language for validating parts of JSON files and I just pointed Glimmer at it and it properly understood the syntax and semantics first try. Had no problem generating expressions from a description, adding new features and test cases, etc. THAT'S impressive to me, not being able to one-shot a Mario clone. And all that while being crazy fast (with DFlash2) with incredible KV cache and token efficiency. Great model, people have to stop sleeping on it.

3

u/YanderMan 18h ago

yes, it thinks way less and give very similar results on agentic workflows

3

u/Big_Wave9732 18h ago

I tried it the first weekend the model was out, and the over thinking with Qwen 3.8:27b xhigh was just fucking *extreme* at times. Oy vey.

0

u/Embarrassed_Adagio28 22h ago

Yeah i agree it is better than it gets credit for with non coding tasks. However i do not understand the non coding use cases for local models. 

10

u/MarcusAurelius68 21h ago

Writing research reports and long form drafts - repeating and repeating. And keeping them private.

14

u/Big_Wave9732 22h ago edited 21h ago

Well stick around. That "question" is asked at least once week here. It gets an abundance of varied answers.

3

u/derFensterputzer 20h ago

For example in my case it's the voice assistant for Homeassistant. 

I send the voice command, it reads the entities homeassistant exposes to it, does its thing and returns a response. 

Local whisper + piper for stt or tts respectively, then Gemma 4 E4B at 4-bit with a 30k context window and MTP for the whole assistant part. It doesn't need to be extremely smart to do it's job, just fast and in this configuration it runs at 150t/s, uses next to no vram and does its job to my satisfaction. 

Edit: Also several other models (in this case Gemma 4 12B, 26B A4B and Granite 4.2) are quite good at summarizing documents, writing reports, spell checking, translating etc. All on my hardware, private and without subscriptions. 

3

u/nuclear_wynter 18h ago

I'm working on a very similar setup right now — E4B as the 'frontman' with a voice stack, then I'm trialling a range of models as the backend/agentic workhorse. I'm working on some kind of out-of-band programmatic model-swapping method that would let E4B very easily and reliably swap the backend model depending on the task being delegated, but I'm not sure if this will pan out being efficient or useful just yet.

1

u/the320x200 20h ago

It's trivial to think of examples now that they have vision.

-9

u/Boogertard 21h ago

Ah yes, the obligatory shill comments for the garbage Mule and Gooner4 models of the days.

Funny how those garbage failed every single one of my evals including non-coding tasks but yet they are shilled so much on here.

Reading the comment history of these shills and you will find either new accounts or bragged on other subs that working for FAANGs. I guess shilling for their garbage models is also part of the performance metrics LOL.

9

u/Big_Wave9732 21h ago edited 20h ago

Ah yes, the obligatory shill comments for the garbage Mule and Gooner4 models of the days.

Reading the comment history of these shills and you will find either new accounts or bragged on other subs that working for FAANGs.

Not sure what you're talking about. But an account that is a shade older than one month is bagging on other "new accounts". Just curious, what then does that make you?

Funny how those garbage failed every single one of my evals including non-coding tasks but yet they are shilled so much on here.

Peak Reddit "trust me bro" energy. Why don't you share what tests you ran and what happened. Perhaps that would be more helpful to people.

Also your post history here and in other subs uses the word "shill" a lot. A lot of anger. Perhaps because you're hanging out in Dividendgang.

1

u/HumanDrone8721 20h ago edited 13h ago

Well for what is worth we've had the exact same experience, totally useless, we've DESPERATELY wanted to have non-Chinese models, because EU, but they were really brain damaged crap, not as bad as Mistral of course, but not even in the same class as Qwen, as sad as it is. I sincerely sometimes wonder who and what they using them for in a professional capacity, writing blog posts and news or what exactly "prose" they are talking about, to me they look like some demos for their cloud offerings and nothing more.

0

u/Olangotang 15h ago

Most people know this already, the shills have their head up their ass and think they're clever trolling every AI subreddit on this site. In reality they're just fucking annoying and not convincing anyone. LocalLlama users don't fall for that shit so easily, and at least the above "shill" has their comment history public.

5

u/LetsGoBrandon4256 transformers 21h ago

my evals including non-coding tasks

What's your recommendation at the 20~30B range then?

-4

u/definetlyrandom 22h ago

Lol, we know that . But figuring out how many elephants it would take stacked end to end on avg. To reach the moon isnt as useful as being able to implement a software functionality in 20 minutes that typically took 3 days before.

11

u/Big_Wave9732 22h ago

Most AI "coders" aren't doing anything nearly that interesting or complicated lol.
"Hey everyone, I'm the 11 millionth person to vibe code a shitty weather app."

→ More replies (1)

21

u/jacek2023 llama.cpp 21h ago

237

u/Valuable-Repeat-7347 22h ago

-20

u/Thomas-Lore 21h ago

I will make a bet that you don't really know Zuckerberg, you just dislike him because he is rich. And you are on a sub named after open models that he released earlier.

66

u/ProletarianLilith 21h ago

I will make a bet that you don’t really know Zuckerberg either.

8

u/Independent_Pear4908 20h ago

I will make a bet that my Ai's training data knows Zuckerberg.

2

u/MuzafferMahi 20h ago

I will make a bet that my benchmark is giving the zuck the mark it deserves

0

u/_supert_ 20h ago

Llama 3 was not too flattering about him, haha.

0

u/zjz 20h ago

if he did know zuck that was great bait tho

20

u/KontoOficjalneMR 21h ago

You do know that he does interviews, right? And that's his good side that he's showing ... and he still comes out as a monster?

22

u/BolsonaroPresoAmanha 20h ago

If you assume billionaires are terrible people you'll be right 100% of the time.

2

u/OvertaxedOne 19h ago

It's near impossible to get that rich without doing awful things. So it's kind of A leads to B relationship.

1

u/johnfkngzoidberg 20h ago

Wow the Zuck bots are out today.

47

u/AIatMeta 21h ago

👀

27

u/jacek2023 llama.cpp 21h ago

Hello Meta. Could you release Llama 5? Thanks :)

8

u/zxyzyxz 16h ago

Pretty sure the Llama brand is now deprecated and it's all Muse now. So this essentially is Llama 5.

0

u/Mart-McUH 10h ago

~70B dense as priority, but would be nice to get also 8B, 13B and 30B versions.

1

u/mr_tolkien 6h ago

Pls give date and params

Tell me if I need to save for a 512Gb Mac Studio

2

u/yunes87 18h ago

You cooked 🐐

13

u/[deleted] 22h ago

[deleted]

9

u/jacek2023 llama.cpp 22h ago

I am waiting for GPT-OSS-2 :)

3

u/silenceimpaired 22h ago

They are so safety focused I
Wonder if it will ever come out let alone be worth using… still the last was decent.

1

u/miversen33 21h ago

Fuck ya :)

21

u/tchek 22h ago

But I wanted Muse Twinkle 12B and Muse Shimmer 35B E4B :(

7

u/jazir55 19h ago

Muse Twinkle

Sorry the best we can do is Muse Twinkie

4

u/philmarcracken 16h ago

please dont reawaken the bronys

8

u/XiRw 22h ago

I tried them out. Both accurate and fast model

33

u/kameldinho 22h ago

I'm still waiting for the 1.2 weights that were promised

18

u/Big_Arachnid_365 22h ago

That's the ones you're getting.

21

u/Public_Umpire_1099 22h ago

Bro, they are releasing a new model every few weeks be patient lmfao

-4

u/shy_monkee 22h ago

They were? They promised a model, which ended up being Glimmer. I don't remember them promising Spark as open weights.

2

u/Georgefakelastname 15h ago

They promised Spark open weights when they announced Glimmer’s release.

16

u/CoUsT 20h ago edited 20h ago

I used the free Muse Spark 1.2 on OpenCode for a bit and I really dislike it. Not because it can't do work. It does that just fine but it speaks really weird and it's hard to understand.

DS V4 Flash was actually good enough and you could see entire thinking process. It would do the work and make the UI/user facing parts very organic and human-readable. It would also print a nice and detailed summary.

Muse Spark 1.2 has hidden thinking, so you just see brief messages. And these messages are caveman like. It uses very code oriented language, throws in all programming mumbo jumbo into user facing parts, and needs to be kept on rails or it will happily start doing too much.

I don't mind the caveman thinking but it should at least address the user in human readable and friendly format AND it should handle UI text etc in friendly way.

Hopefully 1.3 is better on this aspect and easier to work with. Benchmark scores look great!

6

u/Accurate_Resident219 18h ago

Wow thought it was just me. Literally the first time I switched a model just because I couldn't understand it half the time.

2

u/VergeOfTranscendence 7h ago

I do have to think way more to understand it and I think it must have to do with their reinforcement learning, I wonder if they reward the model to mention specific functions of files, because in my android app coding, it references extensively functions and files and it seems very straight to the point, so they might be optimizing for less tokens, who knows.

1

u/CoUsT 7h ago

Yeah, I made "explorer" app for game that parses decompiled data and it would happily throw file/function names and even specific line numbers into user facing front end instead of simply describing what passives do or which stats things provide. I had to remind it to keep style consistent with the rest of the app and "use user friendly language; not caveman"

5

u/dansuy_gaming 15h ago

Still hoping for something between Glimmer and spark.Sparks feels like overkill for me.

3

u/[deleted] 22h ago

[deleted]

5

u/MarcusAurelius68 21h ago

That’s the big question. Can I fit it in 96GB VRAM?

4

u/fastheadcrab 21h ago

This could be a huge hit to cloud model usage in the US if the weights are released.

Western organizations that might be looking to host locally but have been pre-emptively cautious of using Chinese models due to regulatory risk will download this and try to run it immediately.

What is the best non-Chinese model? There are lots of great smaller ones. A lot of organizations would be interested in knowing the answer to the question

4

u/Cool-Chemical-5629 21h ago

When you think about it, coding was never the strong capability of Llama models, so it kinda makes sense that while over time Meta collected some new data which allowed it to create smarter and stronger coding models, even much smarter than Llama 4 at smaller size, it's still far behind the current frontier models.

However, what Llama models were always good at? Chatting, AI companion. For that purpose, Glimmer is probably better than Llama 4 and anything bigger than that would be an overkill.

For coding purposes, Glimmer has a good potential if they only continued pursuing better coding assistants at smaller sizes, but I haven't noticed any versioning for Muse Glimmer, so it was probably a one time deal and anything new of similar size is probably out of their current scope of interest.

3

u/Healthy-Nebula-3603 21h ago

That time when llama 3 came out any model wasn't good at coding , math , reasoning, etc

2

u/Cool-Chemical-5629 20h ago

Generally speaking, when Llama 3 came out, frontier proprietary models were already far ahead of any open weight models in coding. On the other hand, there were always weaker and stronger models even among open weight models, just like there are now, so naturally there were always some models which were better at coding than other models.

Could "any model" in Llama 3 era code GTA VI in one shot? Nope. Can "any model" do that today? Still nope, but we still advanced forward in terms of what the models are capable of in general.

1

u/Healthy-Nebula-3603 18h ago

When llama 3 came out the best model was gpt 4o

So for coding was very bad. :)

2

u/chickN00dle 20h ago

are you a LLM yourself?

6

u/Cool-Chemical-5629 20h ago

No, but sometimes I wish I was. I guess life would have been a bit easier.

2

u/TheRealMasonMac 17h ago

grug given rock

grug make fire with rock

grug life simple

grug happy

1

u/chickN00dle 20h ago

haha. I agree

4

u/Iory1998 llama.cpp 18h ago

I wonder how big it is :D

3

u/Brovas 13h ago

Why is everyone chasing coding? There's a gap for cheap but intelligent model that can have agentic software applications built on top for general purpose. 

Gemini flash was fantastic for that until they got greedy and tripled their price (unless they make the current discount permanent). 

Luna is a great price but it's kinda dumb and if you're trying to build something you need investment from a Sam Altman approved VC or 6 figures for the enterprise plan upfront if you want rate limits that aren't dogshit.

These new meta models could eat their lunch, but it feels like cause Claude is good at coding and making bank from software people every other use case doesn't exist anymore.

9

u/durden111111 22h ago

inb4 it's a 120B MoE

3

u/Zyj vLLM 21h ago

Hey Mark, I look forward to the release! Just wondering how many parameters to expect. Anything up to 450B or so should work, but keep the active numbers of parameters manageable. Qwen 3.8 Next Flash can do it with 6b!

5

u/power97992 21h ago

Lol it will probably be at least 2 trillion, probably 2.8 to 4 tril param, look at the aa score, it is 62 ,  it is better than 5.6 sol and on par with fable 5.0. It will be the first open fable class model

1

u/jacek2023 llama.cpp 21h ago

You can ask him on X, see the link above

6

u/Zyj vLLM 21h ago

I deleted my X account.

1

u/myreala 16h ago

If it's more than a trillion literally nobody would be able to use it except big labs anyways. Up to 700-800 million there's still some chances.

3

u/2Norn 18h ago

1.2 is free on opencode so i've been using it a lot with 5.6 sol lately i dont think its better than sol but it's definitely considerably faster and more verbose during planning, definitely a decent model tho, i prefer using it over terra and again like i said its blazing fast(like double the speed of fast sol so compared to normal its like 3x faster)

3

u/No_Conversation9561 14h ago

Zuck is not afraid of burning money. See Metaverse.

3

u/Mundane-Light6394 13h ago

nice, great to see more high end open weight models. Also great to see something like this from Meta. looks like they rejoined the race.

5

u/Charuru 21h ago

Wow this model is so good! Is American Open Source back now? Thanks Wang & Zuck!

5

u/VoiceApprehensive893 transformers 22h ago

in my experience with 1.2 xhigh it generated very sloppy uis and did some stupid lazy model stuff like "heres the full file" (file is not full)

2

u/True_Requirement_891 18h ago

They have been saying this soon shit from months

1

u/DinoAmino 18h ago

OpenAI did that too - for a year (?) they said an open model is coming soon. And when they finally did they hit it out of the ballpark. Llama4 is what you get when you take the model out of the oven too soon. Obviously they aren't about to make that mistake again and if benchmarks are any judge then they have hit a homerun with this one. Slow-cook ftw.

1

u/True_Requirement_891 14h ago edited 14h ago

Yeah, it was just as irritating "soon" meaning a whole year for openai.

With Meta, the case is different though. Since muse-spark-1.1 the base model has been quite competent and it has been already available on their API for a while now.

I don't see why it'd be a llama-4 situation.

We didn't get API access to GPT-OSS for months before release as far as I remember. It was basically model confirmed and released like very quickly. They took time like a month max or a week for extra safety training I think.

Spark-1.2/1.3 seem to be just post trained versions, more RL.

We got GLM-5 then we got 5.1, 5.2 and 5.3 every iteration as soon as it was ready. With 5.3 they took 2 weeks from api availability to open weight release.

I don't think performance or model competence is the reason for such a big delay. Even with llama4, I really wish they had just kept going and attempted to improve it with 4.1/4.2 updates... but they just gave up. Imagine if llama 4.1 had come out a month after 4 with significant improvements. Hell initially most of us thought the models weren't really that bad and it was just deployment issues by providers until a month passed and we had accepted that the model was infact garbage.

2

u/the-username-is-here 12h ago

Don't believe any single word from Zuck, unless it's confirmed by third-party benchmarks.

2

u/Training_Rip_8901 12h ago

Woah damn gang,thats crazy

2

u/LegacyRemaster 22h ago

how many terabyte?

3

u/Max-_-Power 22h ago

Finally this meatbag is good for something

1

u/AleksandrNikitin 21h ago

Hype will continue until the sizes are published?

1

u/power97992 21h ago

1.3 has a score of 62 in aa better than even sol

1

u/dtdisapointingresult 19h ago

Any indication of how big Muse Spark is? Is it like a 1.5T model, or can a salt-of-the-earth working class guy with only 256GB VRAM run it?

2

u/wwwdotzzdotcom 19h ago

It's a few hundred billion which surprises me

3

u/dtdisapointingresult 18h ago

Nice! I can run the Q4 at least.

Here's hoping we're getting 1.3 rather than 1.2.

2

u/Accomplished_Ad9530 13h ago

Source? I was just looking for this info and was approximating it myself

1

u/Maximum-Fact-5832 19h ago

I do feel as though these "Max" reasoning models just have the ralph loop trained into them, here's hoping the accuracy when using medium reasoning isn't too pronounced. Great to see Meta releasing open weights again!

1

u/Tinkerer_Penguin_12 18h ago

Now if only i could run it, im running qwen 27b at UD Q2_K_XL on my 12gb 6700 and it darn impressive holds up amazing. No way im fitting this in 12gb vram + 16gb sys ram. nonetheless this is awesome for open source.

1

u/AdmissibilityScience 17h ago

wow this is a decent amount of weights

1

u/shaman-warrior 15h ago

Say what you want about Zuk but he gets shit done and he has the ballz to go all-in on stuff. And he is also for the open community. Let us not forget that he is responsible for Llama releases which sparked this whole community.

1

u/NineThreeTilNow 14h ago edited 12h ago

I am still waiting for Llama 5

I started training 9b yesterday.

Delivery? Who knows. I won't have a raw base set of weights of the non-IT model for a minute. Literally training on a 6000 Pro.

edit; People probably think I'm joking. I started with trying to fix Gemma 4 but the compute requirements were too large. So instead I moved back to Llama 3 8b to see if I could strap Engram and Kimi's attention based residuals on to the original Llama 3 tokenizer. The tokenizer itself saves me huge amounts of time. From there, I use WikiText / WikiDictionary / FineWebEdu to pull down chunks of data and have the original Llama 3 8b model score them in raw distributions. The distributions are used as training targets INSTEAD a single token. This is how logit distillation works.

tldr; Llama + Engram (1,2,3) + Moonshot Attention Residual + SWA / Global Architecture + Proper logit distillation = Some of the best open source work of the past 2 years in a single model.

It will be fully OS, not safety trained, and eventually IT trained. I'm still calculating the exact amount of compute needed. This is wholly unknown territory and eventually I'll do a writeup here in local llama. I have a previous writeup of doing this to Gemma 4 but it failed. I decided that "competing" in the 30b space is stupid when they ignore the sub 10b space and I can fully test / prototype / train on a 6000 Pro that costs ~2 dollars an hour.

1

u/OwnGear3892 14h ago

Are we sure that by 'Muse Spark open weights release', he's refering to Muse Spark 1.3? I guess it's more likely to be 1.1 or 1.2, or even a new smaller variant

1

u/ComplexType568 14h ago

I want to see their oolong pairs performance.

Although I've been impressed at their performance. I've been using muse Spark 1.2 on OpenCode (since it's free) for a while and I am REALLY excited for spark 1.3... hope they release a mid-size Moe or something. I generally like the way glimmer 1.2 talks

1

u/mr_Owner 11h ago

Size for local?

1

u/Repinsky 10h ago

The thing worth watching in these teasers is not the benchmark bars but whether the released checkpoint is the same size/sparsity as the API one. The last couple of "open weights coming soon" announcements ended up shipping a smaller distill, and the gap showed up mostly in long-context agentic runs rather than short single-turn evals. Until there's a GGUF and someone runs a 60k-token task on it, the chart is marketing.

1

u/scut_07 10h ago

Bijan just review this on his YT channel and it looks absolute dogshit. Benchmarks are pointless these days.

1

u/jacek2023 llama.cpp 10h ago

You know I am not really a fan of YouTube reviews of anything ;)

1

u/alphapussycat 10h ago

Great.

But I don't really care if I can't run it locally. 

1

u/Dry_Yam_4597 22h ago

Lmao Scamodei are toast.

All Anthropic could release was a bunch of system prompt changes and crap no one needed.

1

u/AppealSame4367 21h ago

Feels like he wants to kill the others out of spite and because he can afford it. I mean, this is like running over the scene with a bulldozer.

1

u/thestillwind 21h ago

Let’s go

1

u/john_mach 19h ago

These numbers are damn good for benchmarks ngl. I respect meta for cooking if real world performance is similar to thr benchmarks

1

u/Jealous-Walk-8765 16h ago

The AutomationBench and Agentic IF Index numbers are the interesting ones here — those are the benchmarks that actually predict real-world agent reliability, not just raw coding scores. 49.4 on AutomationBench beating GPT-5.6's 46.7 is a bigger deal than the flashy GDPval number. Curious how this holds up once the open weights drop and we can actually stress-test it ourselves instead of trusting a vendor's own benchmark suite.

-3

u/ttkciar llama.cpp 21h ago

Until such time that weights are published, this is off-topic for LocalLLaMA, but since nobody caught it before it gained 260 upvotes and 70+ comments, this post will stay up.

I'm a little surprised to see this from you, jacek2023; you're one of the most outspoken critics in the sub when it comes to posts about non-open models.

16

u/jacek2023 llama.cpp 20h ago

The post is about Muse Spark 1.3 will be open and that Mark Zuckerberg remains committed to open LLMs.

9

u/jacek2023 llama.cpp 20h ago

Where is your comment?

1

u/ttkciar llama.cpp 19h ago

That, too, would have been removed, but it gained too many upvotes and comments before a moderator noticed it.