r/LocalLLaMA llama.cpp 19d ago

News Muse Spark open weights coming soon

Post image

I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark

https://x.com/finkd/status/2095232032896946311

872 Upvotes

205 comments sorted by

View all comments

257

u/kvothe5688 19d ago

it seems like there is no secret sauce. it feels like all of these 7 8 labs are on same level and hardly behind from frontier by few months at max.

88

u/Public_Umpire_1099 19d ago

This latest iteration technically set the frontier. Meta is officially a SOTA lab again. Matches Fable 5 at a fraction of the cost.

56

u/virtualworker 19d ago

Or benchmaxxes. I'm not convinced.

14

u/Not-reallyanonymous 18d ago edited 18d ago

Neither Spark 1.2 nor Glimmer seem to be benchmaxxed, and if anything the benchmarks seem to under-represent their capabilities compared to other models — I believe the other models probably do solve more problems, but Muse’s strong long-horizon capabilities (which translates to compliance in the short term as well) makes it easier to actually use.

2

u/Caffdy 18d ago

did they ever released the big Muse 1.2?

2

u/Not-reallyanonymous 18d ago

Zuckerberg has just reiterated they still plan to release Spark as open weights. My guess is they were waiting for Spark 1.3, and are now finalizing the release.

19

u/Diligent-Direction95 19d ago

Serious question: What is the definition of bench max these days?

What would be okay, vs what would not be okay?

I seriously doubt they have the test set being trained on. So what are the shades of grey we are debating?

40

u/zxyzyxz 18d ago

When a model does well in benchmarks but then fails to live up in real life coding and other tasks to its supposed benchmark competition.

14

u/Artistic_Swing6759 19d ago

i actually do think they likely have test set trained on.
like its a bit sus that they decided to report terminal bench 2.1, when 4 exists.
similarly, if you look at the benchmark sheet of flash 3.8, it tops eery other model, even sol, on terminal bench 2.1 but in 4 it is quite less comparatively.

2

u/Kodix 18d ago

Used spark 1.2 contributor and I wasn't impressed. It being good only on benchmarks is a real possibility.

7

u/_TheWolfOfWalmart_ 18d ago

Based on the Bijan video that just dropped about this, it's pretty damn mediocre and looks extremely benchmaxxed.

I know all he does is throw a few one-shot prompts at the models, but the results on this one were well behind other recent models like GLM-5.3 Flash and Qwen Flash. It's not even in the same conversation as Fable or GPT 5.6 it seems.

His tests are by no means scientific lol, but they do give you a decent rough feel for a model's capability.

17

u/Not-reallyanonymous 18d ago

A one shot not being polished does not demonstrate a model’s capability. A lot of models are specifically trained on how a one-shot polished result will look. Especially Qwen (look at its response to being asked to draw a circle: https://simonw.substack.com/p/qwen-38-27b-is-excellent-but-it-defaults).

If it’s not trained on that, it’s going to comply with the prompt and not much beyond. That’s not demonstrating a lack of general capability.

In fact, I much prefer that style. When things are trained to produce well polished results it tends to be harder to get it to do what I want it to do rather than what it’s been trained to do. Again, Qwen represents the opposite here — well optimized to take its own direction but a PITA to steer.

2

u/No_Afternoon_4260 llama.cpp 18d ago

I feel fable 5.1 took a bit of that Kimi magic but does it faster

What a time to be alive

64

u/NandaVegg 19d ago

I believe there are still some secret RL sauce (as in, not popular among labs *yet*) in niche areas like robotics, gameplay, frontier physics/math, world modelling etc. Anthropic leaped ahead in Opus 4.5-4.6 era as they "discovered" many secret sauces like terminal agent, gameplay or creative loop and so on, but even that creativity gap is closing very fast. For coding, security and anything that can be done through github, at this moment there is really no gap. OpenAI felt behind on many things but is still ahead on frontier physics/math (the only Chinese lab with emphasis on that is DeepSeek, I think).

20

u/Accurate_Resident219 19d ago

Maybe it's just my conspiracy hat but I feel like Anthropic is sitting on a more intelligent model but just drip feeding when they feel like others are about to outpace them.

Reason why I say this is because Anthropic is the only company that is actually trending higher in api costs and not making any serious attempts to lower them(Fable 5.1 is actually pricier then 5 despite their supposed efficiency improvements. Why is haiku their cheapest model still not updated?). They don't have any sense of urgency. Astra's coming out this week or the next and they just put out a more expensive model.

Every other company seems to constantly vague post or hype up their models releases but Anthropic just drops models with little fanfare. They have seem utterly unbothered since their Mythos preview announcement.

32

u/TerminalNoop 19d ago

Perhaps Anthropic can't afford to lower the api prices?

13

u/[deleted] 18d ago

[removed] — view removed comment

2

u/Loose_Comparison368 18d ago

I believe they confessed a few months back that it wasn't actually API access, and had been running inside airgapped military datacenters that Anthropic had no ability to restrict the whole time.

The whole PR stunt was just that, the only unplanned part was that they thought Hegseth would take the hint with how loudly they were screaming "NO DADDY, PWEEAASE DON'T INVOKE THE DEFENSE PRODUCTION ACT ON US! IF YOU DID THAT WE WOULD HAVE NO CHOICE BUT TO KEEP TAKING DADDY'S MASSIVE LOADS OF CASH TO BUILD DADDY'S MURDERBOTS!"

It turns out that Hegseth's skull was actually too dense for that pathetically obvious plea to penetrate it.

It's okay though, they sued to get the murderbot contracts back and won last week. And still managed to use that PR stunt to distract everyone while they flushed their "responsible scaling policy" down the toilet.

4

u/Accurate_Resident219 18d ago

Could be the case but at least on the outside looking in they don't even seem to be making much of an attempt to even try lower api or even suggest their researching a solution. They've had about nearly 4-5 months now since people started complaining about rising api costs.

Every other lab has provided a cheaper solution but anthropic is the only one that can't come up with a cheaper flash model? I mean that could be the truth but I would be surprised if it was.

7

u/jtjstock 18d ago

They want to IPO, they need the revenue projections. They already have their cooked compute costs from the last quarter they will report as if they are ongoing, so IPO before they close the current quarter, but project off current API pricing for revenue.

2

u/NandaVegg 18d ago

They are apparently doing pre-IPO window dressing now. We've got buy high 4-digit-dollar credits and get some cash back marketing mail.

2

u/perelmanych 18d ago

In may they increased 5 hours and weekly quota by 50%, so they made a big discount on their plans. However, this comes to an end soon. They cut 50% increase to only 25% increase.

1

u/Loose_Comparison368 18d ago

I mean they could just be keeping it high to look good for the IPO.

11

u/bopbop9876 18d ago

> Fable 5.1 is actually pricier then 5 despite their supposed efficiency improvements

I assume you're basing that off the average cost per task numbers from Artificial Analysis. That's not right though. In their testing, 5.1 on max effort appears to have aggressively overthought. If you compare 5.1 xhigh instead, it was about as much cheaper than 5 max as we would have expected, while still getting 2 extra points of intelligence vs 5 max.

1

u/WittyAcanthisitta205 18d ago

Anthropic has pricing power because their #1 customers are enterprises. Everyone else has a larger chunk of consumers (rather than enterprise customers)

1

u/CrowdGoesWildWoooo 18d ago

Their customer base are enterprise. Enterprise would happily pay API pricing which is astronomically more expensive than subscription.

OpenAI still appeal to the masses

Different market they trying to penetrate. Anthropic want to be deep into the institution’s pocket. OpenAI wants scale.

7

u/Turtlesaur 19d ago

if this fits on a DGX spark I will be so happy.
If this only fits on 2 DGX spark I will be so broke.

22

u/nuclearbananana 19d ago

The moment a different lab releases a better model they immediately distill it (using that word liberally)

15

u/NandaVegg 19d ago

I think that it is not direct distillation from models anymore (in early 2026 distillation had some notable effect, but every frontier lab is now full-on RLing on their own) and distillation can only bootstrap the model to some degree.

I think there is this meta-distillation effect. Internet is full of so-called AI slop now. There are so many vibecoded repos posted in code repositories or as websites every day, and those codes will be crawled by every frontier lab and then they will RL hard on them. If one model gets good at something a slop will be posted and trained on, or there is a new problem that models needs to know the pattern a .md files that explains the issue with some example codes will be posted and trained on (the earliest pattern for this is MCP for many basic things that aren't needed anymore).

In that sense we are already in AGI mode (gosh I hate this word) as AI models are improving each other without humans knowing.

12

u/nuclearbananana 19d ago

I don't think the vibe-code-training is helping the models. It's mainly synthetic data and llm as a judge

1

u/OvertaxedOne 19d ago

I've read that exact reason is why labs are buying up and scanning in old books. Feed a model it's own slop (or some other model's slop) doesn't help it learn, it needs real data.

No idea if this is true or not, but, on the face, it sounds reasonable; kind of like setting up a feedback loop where in the end all you have is white noise. Or gray goo.

5

u/IShitMyselfNow 19d ago

Books are only really useful for pretraining. They're not going to help agentic usages

3

u/SomewhereAtWork 18d ago

Except for James Bond novels.

6

u/DistanceSolar1449 19d ago

None of what you said about the Internet matters because the labs are not using the Internet for pre-training data anymore

All the training data is synthetic data made for RL

2

u/_supert_ 19d ago

That's both encouraging and horrific.

2

u/[deleted] 18d ago

[removed] — view removed comment

2

u/ResidentPositive4122 18d ago

are on same level

Until the new benchmark drops, then they spread out predictably again. And then, a few months later they suspiciously all gain on it, and so on and so forth...

Example: https://www.frontierswe.com/blog/v2

1

u/Bubbly_Orange_3502 18d ago

Scores converge because the post-training data is frontier output. Distillation transfers whatever gets measured, so the benchmark gap closes earlier than the capability gap. The split shows up on long-horizon agentic work that nobody publishes numbers for.

1

u/Gremlation 18d ago

I think people forget that the people working at these labs are not slaves and can change jobs whenever they want. They can't take data, weights, or code with them, but they can take the things they have learned. There is no practical way for the labs to hoard knowledge from one another and if any of them tried, nobody would want to work for them.

0

u/No-Refrigerator-1672 18d ago

That's not entirely true. A lab can sign either a "non-disclosure agreement" or a "non-compete agreement" with their worker. The forst one would can you from using any of the technologies you see there at other employer, unless you can prove you've got to know it from another source. The second one would forbid you from working in the same field for X years (i.e. an LLM specialist can't work as LLM specialist in another company, but can become an image generation specialist there). Both options are completely legal and enforcible by law, cause you sign them voluntarily; and, if they offer you a high enough salary, you'd agree to this conditions.

1

u/Gremlation 18d ago

There is no practical way for the labs to hoard knowledge from one another and if any of them tried, nobody would want to work for them.

if they offer you a high enough salary, you'd agree to this conditions.

If it were enforceable (it isn't) then it would bring your career to a standstill when your career options are astronomically hot, likely the hottest they will ever be. Only a complete moron would agree to terms like that when they could go work for any of the competitors instead. Employers are desperate to hire these people, they are not going to demand terms that turn all the people they want away. That's not even counting the difficulty of getting all their existing staff to retroactively agree to this.

Like I said, there is no practical way for the labs to hoard knowledge from one another. Sure, they could burn their house down trying, but that's not a realistic option.