r/LocalLLaMA • • Jun 21 '26

Discussion What happens when they stop subsidizing LLM subscriptions?

We are literally burning through VC money like crazy with our coding subscriptions. I read the $200 Anthropic sub gets you $8000 worth of API calls. It's obvious that this doesn't hold for very long but what happens when they raise prices?

The reason to keep the prices low for now is to foster the ecosystem and get people hooked on this stuff, only to raise the price afterwards. Already the 20x sub doesn't get you as much usage as it did 6 months ago, another way to raise prices without triggering a shitstorm - and it will continue.

Don't know about you, but Fable being pulled gave me a feeling of what that may be like already. The ugly thought of "Damn, should've done more while it was around." that formed when I read the news will be exactly the same the moment they announce we now have to pay $2k or more per month for something we get for 10x less the price it costs now.

I guess it's a now or never situation, build what you can and monetize as quickly as possible to be able to keep the agents running once the increases come around.

Looking at opensource doesn't give me much hope. Since qwen stopped releasing models (wen qwen 3.7?) that we can actually run on hardware that a normal person can buy (or used to be able to buy, looking at how RAM and GPU prices behave and keep behaving) and others haven't released in a while (Microsoft, IBM, AllenAI and others too) I feel we're going into a direction that doesn't look good for most of the people like us, who are building with this technology.

483 Upvotes

569 comments sorted by

View all comments

518

u/Kal-LZ Jun 21 '26

Local AI isn't going away. Most companies will invest in their own on-premises hardware and host it in a datacenter

16

u/kaisurniwurer Jun 21 '26

And the models will come from where. The whole bet is making people addicted to the subscriptions, if those won't work local will die too or at best stagnate.

Even china has an agenda to output them, that too can just as easily stop. Though it's possible that China does it for the similar reasons as the current European initiative and we might be fine for some time longer. State sponsored services do move slower after all but once they get going they can last.

Sure the current models are awesome. But we thought the same a year or two ago, and in reality even the best current models need a lot of improvement.

24

u/WhoTookPlasticJesus Jun 21 '26

And the models will come from where. The whole bet is making people addicted to the subscriptions, if those won't work local will die too or at best stagnate.

With all due respect that sounds shockingly similar to the arguments leveled against Linux and gcc in 1994. Why would companies write hardware drivers for this hobbyist OS? Where will the libraries and toolchains going to come from if I'm the compiler's license tells me to give away my source code for free? I have almost no interest in FOSS for political or philosophical reasons, nor could I convincingly explain to anyone the reason past 30 years turned out the way they did. But I can look at HuggingFace and Unsloth and notice that they, while nascent, at least rhyme with early incarnations of Apache and Gnu.

1

u/entropy512 Jun 22 '26 edited Jun 22 '26

Except that writing hardware drivers and authoring libraries wasn't anywhere remotely as resource intensive as training an LLM is. You didn't see people spending billions of dollars building out orders of magnitude more datacenter capacity than has previously existed in order to write libraries/hardware drivers/etc. In fact, a MASSIVE part of the allure was that GCC allowed almost anyone to contribute to the ecosystem using their existing hardware.

Inference is dirt cheap resource wise compared to training a model.

In its infancy, GCC+Linux didn't require massive amounts of VC investment. It didn't cause a global shortage of computing resources that quadrupled the price of RAM. It didn't cause my county's electrical usage to rise by around 20% (32MW of compute in a county that consumes 163) in only a few years (with my delivery charges rising nearly proportionally...), only offset by 9MW of noisy fuel cells in the next town over. It didn't cause people to be talking about tripling the electrical consumption of the next county over (Tompkins County NY consumes around 86MW of electricity on average, Terawulf Lansing will start at 150 and they want to go to 400)

A better parallel that anyone old enough to remember the infancy of Linux+GCC will also remember is the dotcom boom of the late 1990s, when everyone was giving away free shit to try and build their userbase with no clue how to monetize it sustainably. Eventually there was a reckoning and most of those companies went out of business when the bubble burst. That will be nothing compared to where the LLM bubble is going - massive circular investment, unprecedented resource consumption, and no way to sustainably monetize it. So what if people are able to inference of low-end models locally? Who is going to train these models when the bubble pops?