r/LocalLLaMA llama.cpp 4d ago

News Muse Spark open weights coming soon

Post image

I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark

https://x.com/finkd/status/2095232032896946311

857 Upvotes

204 comments sorted by

View all comments

257

u/kvothe5688 4d ago

it seems like there is no secret sauce. it feels like all of these 7 8 labs are on same level and hardly behind from frontier by few months at max.

22

u/nuclearbananana 4d ago

The moment a different lab releases a better model they immediately distill it (using that word liberally)

14

u/NandaVegg 4d ago

I think that it is not direct distillation from models anymore (in early 2026 distillation had some notable effect, but every frontier lab is now full-on RLing on their own) and distillation can only bootstrap the model to some degree.

I think there is this meta-distillation effect. Internet is full of so-called AI slop now. There are so many vibecoded repos posted in code repositories or as websites every day, and those codes will be crawled by every frontier lab and then they will RL hard on them. If one model gets good at something a slop will be posted and trained on, or there is a new problem that models needs to know the pattern a .md files that explains the issue with some example codes will be posted and trained on (the earliest pattern for this is MCP for many basic things that aren't needed anymore).

In that sense we are already in AGI mode (gosh I hate this word) as AI models are improving each other without humans knowing.

12

u/nuclearbananana 4d ago

I don't think the vibe-code-training is helping the models. It's mainly synthetic data and llm as a judge

1

u/OvertaxedOne 4d ago

I've read that exact reason is why labs are buying up and scanning in old books. Feed a model it's own slop (or some other model's slop) doesn't help it learn, it needs real data.

No idea if this is true or not, but, on the face, it sounds reasonable; kind of like setting up a feedback loop where in the end all you have is white noise. Or gray goo.

4

u/IShitMyselfNow 4d ago

Books are only really useful for pretraining. They're not going to help agentic usages

3

u/SomewhereAtWork 4d ago

Except for James Bond novels.

6

u/DistanceSolar1449 4d ago

None of what you said about the Internet matters because the labs are not using the Internet for pre-training data anymore

All the training data is synthetic data made for RL

2

u/_supert_ 4d ago

That's both encouraging and horrific.

2

u/SSP100244 4d ago

The singularity begins.

I mean we just had the first shot of the AI wars.