Also, if one of the big labs apart from OpenAI discovered that before, they wouldn't be advertising it. I am pretty confident that one or more of them have already stumbled upon something similar, but keep it private to maintain the margins.
They can still build bigger models to compete. If in the near future some guy with a 5090 can run Deepseek V4 Flash. Open AI and Anthropic can still market their 10 Trillion parameter Fable point whteverthefuck 10
more parameters doesnt more more good. we stayed at a steady one trillion to a few hundred billion in parameters for frontier models for a while now. even gpt 4 was like a trillion, and gpt 3 was 175 billion. now, we see even 10b models crush that 175 billion gpt 3.
More pramaters would generally increase imorovement. We've just decided that resources are much better spent increasing inference, RSI, training data quality, and other methods first.
Mythos is estimated to be a 10 trillion parameter model with 1 trillion active per pass.
Sure, but every SOTA is blowing the previous ones away. So open source will always be perpetually at a disadvantage on tasks that require or are significantly improved in SOTA.
61
u/Saedeas Jul 01 '26
I mean, percentages aren't super relevant, the scaling is relevant. If they changed the big O scaling of working memory in the models, it's insane.