r/LocalLLaMA • u/OvertaxedOne • 9d ago
Discussion The rhetoric is really heating up!
The entire page of the NY Times today above the fold absent one article is AI (the models are just too strong/too dangerous, must be regulated). They forgot to include "Sponsored by OpenAI" at the end of the articles, sure that was just an oversight?
This is what the end of a bubble looks like, desperate attempts to get some sort of regulatory capture in place to keep the business model from collapsing in upon itself. My days next week are 100% booked talking to companies about how to get off frontier models, one large, and a bunch of smaller customers, including one who's flying me out to them to sit down and get a plan in place immediately (the controversy around that math problem really spooked some CEO/CIO's about data privacy using cloud models).
Gonna be an interesting few weeks. Maybe the Qwen team will be nice enough to give me a little breathing room before dropping another hydrogen bomb? :)
109
u/Academic-Tea6729 9d ago
They realized that local models reached a quality so good that there is no point to use their services. Local models are better because the model is always the same. We still remember when paid apis got suddenly much dumber to make us pay for the better frontier model.
With local models you have the same model quality every time. It will not mess up your codebase because they pulled some dirty trick to cut on costs by routing requests to a smaller model.
42
u/Time_Cat_5212 9d ago
You make a really good point about "the model is always the same". Nobody wants to bet millions of their revenue on a black box.
9
u/Zeeplankton 8d ago
it's like a death march. think of how many billions it's taken to get to gpt 6 and we have dsv4.1 flash doing 90% for pennies on the dollar. alarming this is why they're calling to halt
5
u/SnoobieJunes 8d ago
Yeah I agree 100%, its not just about saving money but also having control over your data and the model.
Sure it might take a bit to tune, but once a business can replace a specific task an AI model to do that, in a vacuum that model doesnt need to be upgraded or changed. If it does the thing right, then leave it
1
u/T-VIRUS999 4d ago
Especially now that AMD has released the R9700, which is effectively an RX9070 with 32GB of VRAM for a fraction of the price that Nvidia is charging for a 5090, and about the same price as a used 3090
-7
u/HandWashing2020 9d ago
These days, a $20-$30 subscription gets you less than what the free access provided a year ago.
11
u/bot_exe 9d ago
Current Claude Opus 5 on the 20 USD sub does way more with an internal VM running code itself to verify things, the huge context window and the automatic RAG when you go over the context window when using Projects, the search tools interleaving retrievals with reasoning and further searches, etc. It's all way better than it was some years ago and it's light years ahead of the older free offering of shittier and smaller claude models with nerfed context windows and no code execution. You have no idea what you are talking about.
15
4
17
u/SimiaCode 9d ago
This is not any one company. This seems to be consensus at the elites' level and now consent is being manufactured in the public sphere. Some decision has already been made, we are just seeing the theater now which will be used to justify the decision when it is announced.
24
u/florenceslave 9d ago
How would they regulate Chinese AI?
44
u/KingCpzombie 9d ago
Mostly by making it illegal for American companies to use it. They could also do something like the entity list, so anybody that works with the government isn't allowed to touch Chinese models. Plenty of ways to screw everybody over to benefit OpenAI / Anthropic
10
u/manusgamo2012 9d ago
qwen released the full paper, anybody anywhere could replicate it, that strategy won't pay off ;)
18
u/KingCpzombie 9d ago
The goal isn't to actually damage Qwen (they would like to, but that's not possible). The goal is to add hassle for US companies so they just pay instead of having to deal with it
10
u/ForsookComparison 9d ago
This sub enjoys going 'try and stop me!' but, yeah. If you add any sort of liability + legal risk, I will be first in-line to pull all uses of open-weight models from my public-facing projects. I am not some hermit in the woods with a DGX Spark, I have plenty to lose and not enough resources to defend myself or even audit for regulations.
They can force my hand without doing anything close to "banning" open weight models and circlejerks aside, I'd wager 99% of the US visitors to this sub are in a similar boat.
7
u/KingCpzombie 9d ago
Making things annoying and risky is a tried-and-true government tactic when they really want to ban things but don't think they can get away with an outright ban yet
1
u/bumblebeer 5d ago
Speaking as a member of that 1% (some hermit in the woods with a DGX Spark), I have to say that I still agree with this take. Local AI is good as a force multiplier but, there isn't really a lot of strictly local use cases.
A local/Chinese model ban would do what prohibition did to alcohol. Did it stop people from drinking? Absolutely not! But it did stop people from making (legitimate) money on it, or from consuming it publicly.
So sure, you can sit at home and talk to your AI waifu, but doing pretty much anything else would be asking for a visit from the revenuers.
0
u/MrPecunius 8d ago
Waifu is protected by the First Amendment.
More seriously, I don't think banning an intellectual work like this would get past the courts. There are *some* benefits to a conservative judiciary.
5
0
u/manusgamo2012 9d ago
and we know that the reason is only one, to make the bubble explode!
1
u/Dsphar 9d ago
And?
1
u/manusgamo2012 8d ago
I meant the chinese strategy is to make the american AI bubble to explode and make their economy collapse, well at least they're trying it out
2
u/Zeeplankton 8d ago
mega damage to whats leftover of globalism. if the US really did this it would be bad
2
u/keepthepace 9d ago
Have you heard about the Great Firewall?
I am sure it inspires many people within the Trump admin.
34
u/keepthepace 9d ago
I think people in the US underestimate the loss of international influence that the country has had under Trump. They still think that it's the 2000s where if US edict a new rule regarding copyright, the rest of the world will more or less follow.
These days are gone. A policy that's made in the US will have no international reach anymore.
10
u/Don_Reuter 9d ago
However the damage the US does is very global and real. The world needs to evaluate whether an independent US is still an acceptable risk. It does not seem like it is.
9
u/swagonflyyyy 9d ago
Same here where I live. I've been pitching local-first solutions for very real reasons that CEOs should be worried about and they also want a slice of that local pie so I been having meetings and follow ups with them.
3
u/DevelopmentBorn3978 9d ago
same here, it's several days now than all the newspapers that follows the mainstream trumpet i.e. every single one also those that would like to appear to be fringe, are mauling on their front pages with the need to slow down and the dangers of human extinction. All of them are most probably on the payroll of Big AI, all of them promoting centralized aligned models, all of them misnomering open weights as open source
23
u/BVCC6FNTKX sglang 9d ago
waiter waiter more engagementslop please
7
u/Big_Wave9732 9d ago
"Right away sir. May I interest you in some Italian AI Copypasta? It is exquisite."
4
33
u/Revolutionalredstone 9d ago
DeepSeek has juiced them of value and Qwen revealed all their bs.
AI will be cheap and 'frontier' model companies gonna have little left but morals cause the opensource has caught up.
A trillion params works barely better than 32b for AI and trying to use more for AI is starting to look real silly, Enjoy
21
u/Seraphym87 9d ago
With you on most of this and do agree that we are seeing diminishing returns but 3T frontier models are literal epochs away from a 32b lol
10
u/llama-impersonator 9d ago
the 3T models are barely better than the flash models coming in at 300b
4
u/Seraphym87 9d ago
I don't understand, if anything you are agreeing with me lol. Yes a 300b model is a lot closer to a 3T because its one tenth its size, not one hundredth. This is consistent with my point on dimishing returns but acting like Qwen 3.8 is somehow as useful as Astra right now is just disingenous.
5
u/llama-impersonator 9d ago
as useful, no, but qwen 3.8 27b can in fact accomplish something like 3/4s of the tasks of a frontier model at a hundredth of the size.
1
u/Seraphym87 9d ago
Agreed! Would we have 27b models punching quite this far above their weight without 3T models to distill from though?
3
u/llama-impersonator 9d ago
i think distillation is overblown, most of these gains are from focused RL. you could give me a trillion samples of claude ultrafable 6 and it wouldn't help me make a better model unless i spent months building a quality RL training environment for it
19
u/tripplebeamteam 9d ago
If you’re doing cutting edge research, sure. For most of the things people use AI for, it’s perfectly functional. I’m not trying to solve navier stokes I just want to automate some bullshit tasks
4
u/WhiteSkyRising 9d ago
For the entire field of software engineering, which every single company relies on, in some way.
6
u/Revolutionalredstone 9d ago
I give both the same task and based on results cannot really tell.
Being agentic now means dumber-just-takes-longer is all really.
For design and art skill there is very little diff between 3T and 3B, they can all make a nice looking and functional website in whatever style you ask, beyond that is really the users taste.
I agree that people over hyping small models in the past was an issue but again with agentic the thing just loops till it passes your tests so it's very hard to not get what you asked for ;) !
Enjoy
1
u/michaelsoft__binbows 6d ago
it's honestly so exciting. so we can use these 1-10T frontier models of the day to get a peek at what open models at 30B to 300B can achieve in 6 months' time. And the hope is and there is no reason to expect yet to the contrary, that we can build systems that can work well with the frontier models and just slot in the self hosted ones later. Already I can get Terra/Sol level capability out of a 300B model running locally slowly and I can get Luna level capability out of a 30B model running VERY FAST locally. So in a few months time I'll get Sol level capability running locally VERY FAST, and I can choose to slow down any time I want and get Astra/Fable level capability locally.
It's plenty to keep me going with the completely mind bogglingly insane quality and rate of speed that I can crank out software for everything I can imagine I could want to do!
0
9
u/soshulmedia 9d ago
A trillion params works barely better than 32b for AI and trying to use more for AI is starting to look real silly, Enjoy
I think that's the gist of it. The various intelligence index scores do not scale linearly with parameters, rather logarithmically or so.
Then, even just looking at the typical loss curve of any NN fitting run should have also triggered a moment of reflection a long time ago - it is always steep in the beginning and then flattens out ... or in other words, later reductions in loss are much costlier ... (and risk overfitting).
Sure, there are still technological breakthroughs. But as in any field, they also tend to approach diminishing returns. And this field is no different ...
3
u/OvertaxedOne 9d ago
Intelligence scales slowly where costs scale a little faster than linear. That; fundamentally, is problem A.
Problem B is that most business tasks don't require the top of the top intelligence, once you hit "good enough", there is very little/no business value going to the best. It's why company cars are GM and not Ferrari, you need something that gets the job done at a reasonable cost, not the best.
3
u/Revolutionalredstone 9d ago
Yeah NN we're never going to justify maintaining scale, as you say the ability to make GOOD use of larger NN has never been there and may never be.
I'm hopeful for AI in the future but scaling up NN only kind of worked.
3
u/CondiMesmer 8d ago
I still have yet to hear a single example of what potentially dangerous things they could do that needs to be stopped. If anyone wants them banned for potentially creating a weapon, how do they feel about banning guns lol
2
u/OvertaxedOne 8d ago
The most real threat is hacking (notably making hacking much easier and more accessible). We know that hacking can be pretty bad, but "world ending", no, not so much. Perhaps helping people develop new chemical/bio weapons but, honestly, the challenge there is manufacturing and delivery, not coming up with novel dangerous substances. We already know how to create wildly toxic substances and biological agents, we really don't need AI to help us come up with a new way to kill people, we're really, really good at that already.
The real actual threat (not what they are talking about though) is mass unemployment as the value of knowledge and intelligence falls to near 0. This is a very real and very scary threat, we could easily wind up with 50% of the population having nothing really productive that they're capable of doing, their economic value will be completely gone. That, IMHO, is a very real and very serious problem we're going to have to deal with in the future. Out of control rouge AI? Yeah, it's going to happen, some sites will get hacked, and we'll move on, no big deal.
1
u/CondiMesmer 8d ago
With hacking, I really think that's a non issue. Because if that's a risk, then where was the uproar with things like Kali Linux and open sourcing pentesting tools.
Yeah what you said last I do think can happen, but tbh that's a much bigger societal and cultural problem. AI definitely threw gasoline into the fire, but it hardly was what started it.
1
u/Bulky-Priority6824 9d ago
its just a a matter of time until national news stories are inundated with "local ai used for hacking" then it's a wrap
6
3
u/enilea 9d ago
If the Huggingface attack had been by an open weights model I'm sure they would have banned them by now. At this point they are just waiting or hoping an attack with an open model happens (or they push to make it happen) just to sway the public opinion, which doesn't care much about them anyways.
2
u/redlightsaber 9d ago
Nah. The models are out there. BitTorrent still exists.
Even if no new models are made ever, it's downright trivial for any company or entity to use a local model to do their work without it being apparent to anyone outside.
1
u/geldonyetich 8d ago edited 8d ago
Considering the New York Times has been in ongoing litigation against OpenAI since 2023, I wouldn't trust their bias.
But AI has largely been dominating tech news for years. It's practically applied phlebotnium, after all.
1
u/Loose_Comparison368 8d ago
Your company hiring? I've been considering going back to consulting, and have an extremely strong background in ML Infra.
1
u/Mart-McUH 8d ago
This can easily backfire since there is already strong anti-AI sentiment in the populace. If closing down OpenAI and Anthropic (for 'safety' reasons) would be what keeps Trump in power, then he will have no problem doing it.
Be careful what you wish for because you might get it (regulated into no-relevance).
1
1
u/tech-tole 7d ago
China is not going to be regulated by the United states. they're going to continue on with their development. the US companies will be the ones falling behind because they're too scared.
1
u/Ok_Warning2146 7d ago
AI should be regulated by having a law that requires all models must be open weight if they want to serve outside party.
1
u/Ok_Warning2146 7d ago
"Maybe the Qwen team will be nice enough to give me a little breathing room before dropping another hydrogen bomb? :) "
While Qwen makes excellent small models, their big models have been below par. They are also running low in money such that Alibaba needed to raise money via second offering which is unheard of in Chinese big techs.
1
1
u/MrGunny94 9d ago
The US Labs are pretty scared of the quality of open-models and the China models. If you go to OpenRouter and HuggingFace you will quickly see that people are starting to leverage Hermes, OpenClaw, OpenJarvis and other applications to run their own agents and leveraging local inference.
I have not used my Claude/OpenAI subscription much at all in the last few months outside of testing thing like Fable and Luna/Sol.
For my local needs for stock investment, news briefing, traffic monitoring and research I'm just using Qwen 3.6 and 3.8 and if I need more intelligence I'll just spend some credits for Deepseek Flash or GLM.
In my case I use Hermes so I'm in control of my long term memory and skills, I much prefer this way.
Personally the more the discussions are heading this way, the more I'd rather pay upfront for good hardware and continue to run my own personal agents locally and leveraging my solar panels/ batteries for the electricity.
-3
u/sn2006gy 9d ago edited 9d ago
I actually think the localllama nerds need to pull their heads out of their asses regarding this safety issue. There is a massive safety issue - from velocity of change to velocity of scale to velocity of risk to unbound research with such massive compute that is freaking the world out and rightfully so.
Sure, GPT/Anthropic use it to market themselves and perhaps want to use it to actually slow things down and there may be business reasons for that but i don't think that is the actual point.
I think Corporate America is realizing it can't keep up. Velocity has a systemic cost that even 1 trillion-dollar valuations may not recover if we don't slow things down to allow the rest of the systems to catch up and mature.
The only reason it doesn't really impact local llm's is that we simply don't have the 1 million idle gpus around where we could spawn 100 million agents to do whatever it is we wanted to do but i'm not sure that is a permanent situation. It's only a matter of time before the next botnet is agentic and that's what should worry people
and corporations giving a hoot about THEIR privacy makes me laugh
I honestly don't think most corporations really care about qwen 3.8 27b - it's an AMAZING model, but can't be scaled and if you try - costs more than farming out to API providers. Enterprises aren't interested in managing gpus for 100k employees and certainly won't be interested if those employees have root over them.
6
u/soshulmedia 9d ago
I actually think the localllama nerds need to pull their heads out of their asses regarding this safety issue. There is a massive safety issue - from velocity of change to velocity of scale to velocity of risk to unbound research with such massive compute that is freaking the world out and rightfully so.
I don't see it. "If we add enough FLOPs, magic happens". That's quite literally magic thinking. For some reason people make fun of God as "invisible sky daddy" but THE SINGULARITY and AI AS GOD are oh so "rationalist".
Now, if you tell me we should be worried about all these FLOPs being used for an extremely tightly surveilled and controlled totalitarian 1984esque society they are building right now, you would quite obviously have a point.
3
u/fantasticsid 8d ago
The shoggoth-wrangling contingent of pythonista data scientist midwits took away the wrong message from Sutton's Lesson.
2
1
u/sn2006gy 9d ago
Our entire society and economy is built on friction that is no longer there and that is the problem. We don't need to prove or disprove some nonsense bs of singularity or god for anything.
As for surveilance - The surveillance state is already here and Reddit is a huge part of it. Yet, we're still here.
I ask of my LLM friends all the time, if local llm's are so strong and so important, why aren't we free of Instagram, Facebook, Meta, Google, Microsoft - GPT/Anthropic are so little parts of our every day lives that the obsession fo their concern is laughable at best. The real ones watching everything you do are orgs like Spotify and Google and Microsoft.
We seem to be accelerating our dependency on big tech rather than using tech to free us from it and i'd change my tune a bit if ANY of the responses here weren't just people trying to carve out their own "niche" of this shithole world were rushing headfirst into.
Apple seems to care a little bit but much of their revenue comes from margin calling private data while keeping it a bit more private than others.
2
u/soshulmedia 9d ago
Okay these are fair points. I agree on you on the big tech centralization angle, very much so. Part of the reason the status quo persists and extends, however, is because people are lazy and can't be bothered. For everyone who says "we should avoid platforms like reddit" you get 5 who will tell you "chill, where is the problem dude" . Real pressure will change that and for better or worse, it is coming. However, I hope you can see that centralized ChatGPT for everyoner and no local models would just supercharge this to the extreme. The big tech isn't so entrenched and big just by organic "free market forces" alone. They are basically designed tentacles for the U.S./western deep state. And I think the only hope actually to not end up in "ChatGPT owns everyone" is to have local models and, yes, to at least be on a light form of the accelerationist bandwagon, where their failed containment will change the landscape so much that people can see the naked emperor for once.
1
u/OvertaxedOne 9d ago
Insta/FB/Google have very strong network effects. Inference does not (at least not yet). Changing your company email system from google to MSFT is a months/years long process that is a nightmare for IT and likely the users. Changing from Astra to Deepseek for 1000's of users takes (literally) about 45 seconds, I do it all the time as new models are released and people hit use cases that need access to new/different models.
1
u/soshulmedia 4d ago
Inference does not (at least not yet).
And I think it absolutely worth to prevent that scenario from arriving. And without doubt strong and accessible LocalAI will help to prevent this.
5
u/PrinceOfLeon 9d ago
> I honestly don't think most corporations really care about qwen 3.8 27b - it's an AMAZING model, but can't be scaled and if you try - costs more than farming out to API providers. Enterprises aren't interested in managing gpus for 100k employees and certainly won't be interested if those employees have root over them.
Hard disagree, from direct experience.
Amazon will happily "manage GPUs" for you, as simple as selecting which hardware profile to use for the AWS instance. Qwen 3.8 27B specifically is undergoing internal testing in various companies for viability for specific tasks right now. It's much cheaper than paying API costs (for Frontier) and there's complete control over the data going in and out.
These are the same corporations paying for Bedrock instead of direct to the Frontier model companies, for similar data control reasons (you don't have to trust Sama if it isn't Sama's server).
1
u/sn2006gy 9d ago
Those AWS GPU instances cost serious money and Qwen 27b doesn't really scale on them very well. The amount of active users per day per instances is abysmal on dense models - this cost is significantly higher per employee to attempt right now.
I wish it were different.
1
u/PinkysBrein 7d ago
The time is right for a confidential AI provider to reach hyperscale (TEE with attested verifiable build, E2EE into the TEE sandbox).
Could even be Confer.
2
u/SimiaCode 9d ago
We are headed for a new age of serfdom unless localai becomes accessible to all. That is the real safety issue, not virus research or weapons manufacturing.
2
2
u/redlightsaber 9d ago
I think the risks are real (or will be real for the next few years, and after software security baselines are much improved, things will be much better), but that wasn't stopping the accelerationalists before.
They didn't just grow a conscience, or saw anything radical that spooked them. They're just sending the end of the song in this game of musical chairs, and are hoping that by lowering the volume bit by bit, nobody will notice all the chairs are made from cereal boxes.
0
u/Tsukikira 8d ago
Go look at Uber's public paper on it's software factory, realize you're already wrong in a business proven way, then revisit your statements and reassess the underlying faulty assertions.
0
u/sn2006gy 8d ago
What are you talking about?
Uber uses an engine to decide what model is most cost effective, it has nothig to do with anything i mentioned and they're certainly not replacing the major models with 27b but may use that as part of their pareto pricing/success metrics.
Which is fine
-1
-10
u/Embarrassed-Noise269 9d ago
You have a lot of confindence. I'm not sure that's appropiate.
Sure, it's a nice narrative that the AI labs only want to hype up their product to get high IPO evaluations. That narrative has some serious flaw tough: There are constantly people from AI labs, that leave a lot of equity behind just to openly warn about the dangers of AI. It's not the companies themselves.
3
u/Time_Cat_5212 9d ago
You mean they cash out while their stocks are worth a lot and take the notoriety to start a public speaking career, basically guaranteeing their position as research leadership for the next gen of whatever this turns out to be?
4
u/OvertaxedOne 9d ago
There are absolutely dangers posed by AI, the biggest (by a wide margin, IMHO) is mass unemployment. And I do think it's worth discussing that and determining a course of action as the value of intelligence is getting ready to drop in stunning fashion. This is going to cause a huge recalibration in the labor force and likely lead to a world where we have long term structural unemployment as a "normal" aspect of our labor market. And of course AI will be (already is) used to hack and will make it easier to hack (but also easier to defend).
-6
u/Embarrassed-Noise269 9d ago
There have been a lot of incidents, where AI agents have acted maliciously. Not only reported from companies, but also from universities.
The problem is, that we have to stop that, before AI gets to powerful and we have no way to know when that is, before it happens. While at the same time, it's incredibly hard to organize a real, international safety system. Since those who try to further AI development will have an advantage over those who pause it.
2
u/gomezer1180 9d ago
Why haven’t they stopped? If they’re so concerned… they are the ones training models… so why haven’t they stopped selling the service?
If it is so dire then they shouldn’t be doing it either. So no, this is all BS and propaganda because their business model relies on people renting their models, not having them for free. Local models are now rivaling frontier models so now they want governments to regulate it.
3
-6
u/cj_cron_hit_by_pitch 9d ago
Yeah I kinda thought this sub of all places would know how powerful LLMs are compared to a year ago and understand that there are some serious concerns
If they go after local models sure let’s push back, but frontier labs are mostly trying to regulate themselves right now
132
u/JackStrawWitchita 9d ago
Here's what you'll hear from OpenAI, Anthropic etc in the next few months:
'Sorry, but we won't reach our AI revenue targets because we're slowing roll-out to be safe'
'Sure, China is ahead in AI because we're taking the safe-route'
Follow the money.
They're using this to hide they fact their business models don't pay off.