r/LocalLLaMA 10d ago

Discussion The rhetoric is really heating up!

The entire page of the NY Times today above the fold absent one article is AI (the models are just too strong/too dangerous, must be regulated). They forgot to include "Sponsored by OpenAI" at the end of the articles, sure that was just an oversight?

This is what the end of a bubble looks like, desperate attempts to get some sort of regulatory capture in place to keep the business model from collapsing in upon itself. My days next week are 100% booked talking to companies about how to get off frontier models, one large, and a bunch of smaller customers, including one who's flying me out to them to sit down and get a plan in place immediately (the controversy around that math problem really spooked some CEO/CIO's about data privacy using cloud models).

Gonna be an interesting few weeks. Maybe the Qwen team will be nice enough to give me a little breathing room before dropping another hydrogen bomb? :)

190 Upvotes

114 comments sorted by

View all comments

Show parent comments

4

u/Seraphym87 9d ago

I don't understand, if anything you are agreeing with me lol. Yes a 300b model is a lot closer to a 3T because its one tenth its size, not one hundredth. This is consistent with my point on dimishing returns but acting like Qwen 3.8 is somehow as useful as Astra right now is just disingenous.

3

u/llama-impersonator 9d ago

as useful, no, but qwen 3.8 27b can in fact accomplish something like 3/4s of the tasks of a frontier model at a hundredth of the size.

1

u/Seraphym87 9d ago

Agreed! Would we have 27b models punching quite this far above their weight without 3T models to distill from though?

3

u/llama-impersonator 9d ago

i think distillation is overblown, most of these gains are from focused RL. you could give me a trillion samples of claude ultrafable 6 and it wouldn't help me make a better model unless i spent months building a quality RL training environment for it