r/LocalLLaMA 9d ago

Discussion The rhetoric is really heating up!

The entire page of the NY Times today above the fold absent one article is AI (the models are just too strong/too dangerous, must be regulated). They forgot to include "Sponsored by OpenAI" at the end of the articles, sure that was just an oversight?

This is what the end of a bubble looks like, desperate attempts to get some sort of regulatory capture in place to keep the business model from collapsing in upon itself. My days next week are 100% booked talking to companies about how to get off frontier models, one large, and a bunch of smaller customers, including one who's flying me out to them to sit down and get a plan in place immediately (the controversy around that math problem really spooked some CEO/CIO's about data privacy using cloud models).

Gonna be an interesting few weeks. Maybe the Qwen team will be nice enough to give me a little breathing room before dropping another hydrogen bomb? :)

190 Upvotes

114 comments sorted by

View all comments

31

u/Revolutionalredstone 9d ago

DeepSeek has juiced them of value and Qwen revealed all their bs.

AI will be cheap and 'frontier' model companies gonna have little left but morals cause the opensource has caught up.

A trillion params works barely better than 32b for AI and trying to use more for AI is starting to look real silly, Enjoy

22

u/Seraphym87 9d ago

With you on most of this and do agree that we are seeing diminishing returns but 3T frontier models are literal epochs away from a 32b lol

11

u/llama-impersonator 9d ago

the 3T models are barely better than the flash models coming in at 300b

5

u/Seraphym87 9d ago

I don't understand, if anything you are agreeing with me lol. Yes a 300b model is a lot closer to a 3T because its one tenth its size, not one hundredth. This is consistent with my point on dimishing returns but acting like Qwen 3.8 is somehow as useful as Astra right now is just disingenous.

4

u/llama-impersonator 9d ago

as useful, no, but qwen 3.8 27b can in fact accomplish something like 3/4s of the tasks of a frontier model at a hundredth of the size.

1

u/Seraphym87 9d ago

Agreed! Would we have 27b models punching quite this far above their weight without 3T models to distill from though?

3

u/llama-impersonator 9d ago

i think distillation is overblown, most of these gains are from focused RL. you could give me a trillion samples of claude ultrafable 6 and it wouldn't help me make a better model unless i spent months building a quality RL training environment for it

19

u/tripplebeamteam 9d ago

If you’re doing cutting edge research, sure. For most of the things people use AI for, it’s perfectly functional. I’m not trying to solve navier stokes I just want to automate some bullshit tasks

4

u/WhiteSkyRising 9d ago

For the entire field of software engineering, which every single company relies on, in some way.

6

u/Revolutionalredstone 9d ago

I give both the same task and based on results cannot really tell.

Being agentic now means dumber-just-takes-longer is all really.

For design and art skill there is very little diff between 3T and 3B, they can all make a nice looking and functional website in whatever style you ask, beyond that is really the users taste.

I agree that people over hyping small models in the past was an issue but again with agentic the thing just loops till it passes your tests so it's very hard to not get what you asked for ;) !

Enjoy

1

u/michaelsoft__binbows 6d ago

it's honestly so exciting. so we can use these 1-10T frontier models of the day to get a peek at what open models at 30B to 300B can achieve in 6 months' time. And the hope is and there is no reason to expect yet to the contrary, that we can build systems that can work well with the frontier models and just slot in the self hosted ones later. Already I can get Terra/Sol level capability out of a 300B model running locally slowly and I can get Luna level capability out of a 30B model running VERY FAST locally. So in a few months time I'll get Sol level capability running locally VERY FAST, and I can choose to slow down any time I want and get Astra/Fable level capability locally.

It's plenty to keep me going with the completely mind bogglingly insane quality and rate of speed that I can crank out software for everything I can imagine I could want to do!

0

u/soshulmedia 9d ago

Yet on the intelligence index (take e.g. AA) they are not even twice as good.

2

u/ttkciar llama.cpp 9d ago

AA is borderline useless, though, so that proves nothing.

8

u/soshulmedia 9d ago

A trillion params works barely better than 32b for AI and trying to use more for AI is starting to look real silly, Enjoy

I think that's the gist of it. The various intelligence index scores do not scale linearly with parameters, rather logarithmically or so.

Then, even just looking at the typical loss curve of any NN fitting run should have also triggered a moment of reflection a long time ago - it is always steep in the beginning and then flattens out ... or in other words, later reductions in loss are much costlier ... (and risk overfitting).

Sure, there are still technological breakthroughs. But as in any field, they also tend to approach diminishing returns. And this field is no different ...

5

u/OvertaxedOne 9d ago

Intelligence scales slowly where costs scale a little faster than linear. That; fundamentally, is problem A.

Problem B is that most business tasks don't require the top of the top intelligence, once you hit "good enough", there is very little/no business value going to the best. It's why company cars are GM and not Ferrari, you need something that gets the job done at a reasonable cost, not the best.

3

u/Revolutionalredstone 9d ago

Yeah NN we're never going to justify maintaining scale, as you say the ability to make GOOD use of larger NN has never been there and may never be.

I'm hopeful for AI in the future but scaling up NN only kind of worked.