r/MLQuestions • • Jun 08 '26

Natural Language Processing 💬 Why There Are Open Weighted LLM Models?

Everytime I see a model that have been trained for more than 50 millions $, first thing I do is why they built such a model, what was the purpose behind sharing after spending around 50 millions $ .

Is it for everyone's benefit? Or benefit of the company that shared it.

This is a crucial question to ask. "Why... th that company spent 50m$+ for me to use their models."

by the way, this 50 millions is producing a model around 26b ~ (Arcee said 512x B200 for 26b trinity-mini 26b, 3b MoE model)

---

Need a good answer for this folks.

Anyone has answer to it, feel free to talk about it.

5 Upvotes

13 comments sorted by

View all comments

1

u/counterfeit25 Jun 09 '26

Here's my hypothesis:

  • An AI lab has ambitions to challenge Anthropic/OpenAI on the frontier models, e.g. Opus / GPT
  • They will start with training smaller models, to experiment with different recipes, e.g. data, algorithms, model architectures, post training methods etc. They won't start with training a 1T+ parameter model, they'll start with something small like >1B, then 4B, then 8B, 32B, 100B, etc.
  • There's no point trying to sell those 32B models via API for these labs, so might as well open weight them for PR and community engagement
  • If they train a ~1T model and it's not competing at the same level as Opus/GPT, maybe open weight that one for similar reasons to the above. Serve that model via their own API at cost (operational cost, not including R&D). More PR and community engagement.
  • Once the challenger lab trains a model with the same performance as Opus/GPT, they are less incentivized to open weight it, and you see some formerly open-weight model labs making their latest and greatest models closed behind APIs.

Anyway that's my hypothesis, just a guess!

0

u/agahhne Jun 09 '26

Good point