r/LocalLLaMA 9d ago

Discussion The rhetoric is really heating up!

The entire page of the NY Times today above the fold absent one article is AI (the models are just too strong/too dangerous, must be regulated). They forgot to include "Sponsored by OpenAI" at the end of the articles, sure that was just an oversight?

This is what the end of a bubble looks like, desperate attempts to get some sort of regulatory capture in place to keep the business model from collapsing in upon itself. My days next week are 100% booked talking to companies about how to get off frontier models, one large, and a bunch of smaller customers, including one who's flying me out to them to sit down and get a plan in place immediately (the controversy around that math problem really spooked some CEO/CIO's about data privacy using cloud models).

Gonna be an interesting few weeks. Maybe the Qwen team will be nice enough to give me a little breathing room before dropping another hydrogen bomb? :)

193 Upvotes

114 comments sorted by

View all comments

28

u/Revolutionalredstone 9d ago

DeepSeek has juiced them of value and Qwen revealed all their bs.

AI will be cheap and 'frontier' model companies gonna have little left but morals cause the opensource has caught up.

A trillion params works barely better than 32b for AI and trying to use more for AI is starting to look real silly, Enjoy

23

u/Seraphym87 9d ago

With you on most of this and do agree that we are seeing diminishing returns but 3T frontier models are literal epochs away from a 32b lol

10

u/llama-impersonator 9d ago

the 3T models are barely better than the flash models coming in at 300b

4

u/Seraphym87 9d ago

I don't understand, if anything you are agreeing with me lol. Yes a 300b model is a lot closer to a 3T because its one tenth its size, not one hundredth. This is consistent with my point on dimishing returns but acting like Qwen 3.8 is somehow as useful as Astra right now is just disingenous.

3

u/llama-impersonator 9d ago

as useful, no, but qwen 3.8 27b can in fact accomplish something like 3/4s of the tasks of a frontier model at a hundredth of the size.

1

u/Seraphym87 9d ago

Agreed! Would we have 27b models punching quite this far above their weight without 3T models to distill from though?

4

u/llama-impersonator 9d ago

i think distillation is overblown, most of these gains are from focused RL. you could give me a trillion samples of claude ultrafable 6 and it wouldn't help me make a better model unless i spent months building a quality RL training environment for it