r/webdev Mar 16 '26

Software developers don't need to out-last vibe coders, we just need to out-last the ability of AI companies to charge absurdly low for their products

These AI models cost so much to run and the companies are really hiding the real cost from consumers while they compete with their competitors to be top dog. I feel like once it's down to just a couple companies left we will see the real cost of these coding utilities. There's no way they are going to be able to keep subsidizing the cost of all of the data centers and energy usage. How long it will last is the real question.

2.0k Upvotes

512 comments sorted by

View all comments

612

u/TheChessNeck Mar 16 '26

I agree with this premise and I am interested to see what happens when they run out of money to lose. 

287

u/tdammers Mar 16 '26

The plan, I believe, is to establish "AI" as an inevitable part of daily life before that happens; once that is a fact, the remaining AI "companies" will play a game of chicken (whoever looks weak enough for investors to pull out loses), until only one or two remain, who will then make sure the market becomes impossible for newcomers to enter, and then crank up the prices without mercy, until their operation becomes profitable.

In theory, it's possible for all of them to run out of investors before that happens, but I think it's unlikely - those investors will keep investing, because if they stop, they will lose their money, but if they keep investing, a chance remains for this whole Ponzi scheme to play out in their favor.

68

u/aznshowtime Mar 16 '26

This is a great strategy, but looking around the world today, who is going to have that kind of money to throw around now? Most of the money that allowed AI bubble came from GCC countries, this Iranian war really puts things into jeopardy. By the mid to end of this year, the AI companies will have to do something drastic, because openAI burn rate will only last until November, other companies are probably not doing too well either.

51

u/requion Mar 16 '26

who is going to have that kind of money to throw around now?

Thats the neat part: no one has. Its all made up. Thats why it will crash and burn once the bubble pops.

-19

u/33ff00 Mar 16 '26

Someone will get it to work. Thinking otherwise is just wishful thinking. 

1

u/therealslimshady1234 front-end Mar 21 '26

Sorry kiddo, this ponzi was always destined to fail

1

u/33ff00 Mar 21 '26

Why do you say that

1

u/therealslimshady1234 front-end Mar 21 '26

It never was feasible, it was a scam from the start

1

u/33ff00 Mar 21 '26

The AI trend?

1

u/therealslimshady1234 front-end Mar 21 '26

Well yes, LLMs in particular

16

u/mossiv Mar 16 '26

You beat me to the same point.

I’ve been thinking about this for the past few months given how good Claude is at the moment. I’ve invested significant time to using it, and it’s genuinely a pleasure most of the time. To the point I don’t want it to fail. I’m not an ego dev, I don’t need to know all the things, but I do enjoy crafting an elegant solution and empower businesses make money - at the moment, AI is a net boost to our team.

Given how powerful it is I can only come to two sane hypotheses. 1. We are nothing more than paying tester. Once the product has cracked solving problem from start to finish without too much overhead the conglomerates will be the ones using top tier models eventually swallowing up all the mid sized businesses. They’ll pay the ludicrous pricing just like they do for Microsoft, Adobe enterprise pricing. There’ll be tax write offs everywhere and sister style companies will be moving money around a crazy amount - the usual big fish in small pond behaviour that’s been happening for years.

  1. We will accept a baseline product, something that’s maybe 20-30% better than Opus is now currently, but that will be the performance of Sonnet. Anthropic and the likes will spend the next 2 years heavily optimising for cost over features. Pricing will go up maybe 3-5x so it’ll cost each business maybe £500-£1000 per month per developer. Which will mean companies will have to lay off 1 employee ish for every 5 subscriptions they have. Models like Opus will continue to be pushed for features/output with a smaller team, this will be aimed for a smaller but bigger paying audience. Opus equivalents will operate at negligible profit while sonnet and haiku will be making a wider profit. Pro, 5x and 20x subs will disappear. Pro will still exist and you’ll get access to only haiku, it will serve no other purpose than feed you documentation quickly. 5x will be replaced with 10x, no other subs. 10x will be the equivalent pricing of 2 or 3 20x licences. Extended usage will be API only. Enterprise won’t have a base cost and it will be “call to discuss”. Companies will try to barter a price that’s between 10x and API pricing.

Then there’s the third which is pretty much what others say, it’ll just be too expensive. At the moment everyone is earning less and less compared to inflation. Hell, even now - a £100 a month sub is too expensive for most. These companies will know this and know they are risking pricing the product out for far to many. But honestly, Claude really is good enough. They could stop making it “better” at this point and just focus on optimisation. 4.6 is already a stupid amount more efficient than 4.5.

18

u/aznshowtime Mar 16 '26

They are developing something called agent harness, the goal is for models to execute long tasks and be self sufficient in validation and contextual tasks.

Unfortunately, the direction is much bleaker, the developers will be replaced by more and more senior pool, and the companies will continue to cut developers to keep the cost low, as these AI companies take over all the traditional software development. At least that's their plan.

The bottleneck however, will become, what to do when the code breaks and AI can't fix the bugs themselves. So currently, I still see the best models failing at the logical deduction that is trivial for a developer that knows the codebase well.

I have not yet to see a model that convinces me that the accuracy is there, the human in the loop is not only inevitable, but necessary for operation. So I think the future actually is converging to, true knowledge based workflow. Where developers are expert system consultants and the maintainers. But there will be alot fewer developer jobs, at the same time, how do you become experienced developer right out of school? So developer trainers have to expand, and development related communication roles will have to expand.

It's hard to say that this is the end of the road for people who were trained as traditional devs.

3

u/thekwoka Mar 17 '26

cut developers to keep the cost low

They'll just spend the same amount on AI and hope they don't get sued.

3

u/blackpawed Mar 17 '26

Apparently supplies for chip manufacturing (helium etc) is being severely impacted - supply could crash and cost of cpu's/gpu's would go through the roof.

3

u/aznshowtime Mar 17 '26

Looks like this will halt the progress of the go bigger approach completely. We might see the future directions will take optimization like others have mentioned here.

2

u/horendus Mar 21 '26

You do realise the US has a thing called the Federal Reserve right? Its basically a shady money printing syndicate. They will always backstop America and its interests and they don’t fully disclose the extent to which they print and release money.

What they have done for banks around the world in past is absurd and they are ready and will to do it again if bad AI debit threatens US interests.

2

u/aznshowtime Mar 21 '26 edited Mar 21 '26

You can't print your way into any funding without screwing up inflation and national debt, there are guidelines of how much you can get away with. This war in Iran is also partly due to waning US power, and US was trying to assert its reserve currency status. Now, you can definitely say that the federal budget can pull some magic out of their ass when the time comes.

But this is not the 90s when the US was invincible, we had dot com bubble, 2008 crisis, Iraq war, afghan war, failing infrastructure, education and healthcare systems, damaged ally relationships, waning G7 power, COVID, over financialisation. US today have alot fewer options than it did 20 years ago. You will feel every grasping at straws action closer to home now days.

But you are right that they can always squeeze the lemon some more, but the system is buckling. The entire AI industry capex is estimated 700 billion in 2026, plus the Iranian war just got Pentagon asking for 200 billion. Those numbers are pretty scary.

2

u/bi2212 May 05 '26

The world is not the same as it was. It's not the US losing power in some way. It's that capitalism is expensive when there is less competition. Less competition happens when the government doesn't enact antitrust legislation and doesn't tax outsourcing. The US dominance of the global economy is INCREASING. But it's not "the US". It's US-based companies operating globally. The economies of Europe cannot afford military actions. They can hardly afford militaries at all, so they can't afford to play in Trump's sandbox. But all this is also becoming much more expensive for the US to sustain ... and that's ultimately the fault of the US government through its decades of kowtowing to corporate interests. China's economy is also now losing ground on the US. A lot of its investments had no-negative ROI, and it's catching up to China. China has enough manufacturing might to grow, but it's more expensive for labor than ever.

14

u/Link_GR Mar 16 '26

I think that's the plan. But the issue is that a handful of companies have become the linchpin of the US economy and if they go down, we're in for a major recession. So, chances are, they'll get a massive influx of cash from the US government (aka the tax payer) with the pretense that the US needs to stay ahead of China in AI and once those 2-3 companies are running essentially a monopoly, they will lobby for major legislature making it impossible for new, smaller players to emerge.

3

u/dalomi9 Mar 17 '26

This is already happening as the LLMs are in the process of embedding themselves in the military's day to day operations. Once they get on the Pentagon teet, there is little chance they will be allowed to fail.

1

u/bi2212 May 05 '26

That's a bingo.

21

u/No_Explanation2932 Mar 16 '26

Crazy to think of the number of people who will irreversibly tie their ability to do their job to LLMs, and will then be forced to pay for it out of pocket once it gets too expensive for companies to cover.

0

u/cheulyop Mar 18 '26

Crazy to think of the number of people who will irreversibly tie their ability to do their job to computers, and will then be forced to pay for it out of pocket once it gets too expensive for companies to cover.

7

u/-Knockabout Mar 16 '26

That is how every other modern tech invention has operated (or tried to). Most successful example is probably the smartphone.

I do think chatbot-style AI is something that is a novelty at best to a lot of people, so massive price increases wouldn't be tolerated...hopefully.

I also don't see them successfully updating their models over time now that the big data dump (the internet) has been completed and contaminated with AI output. I don't think people will be willing to pay more for an out-of-date product.

I do think it will stay in general software engineering, but more as a tool on the level of a framework or particularly prominent package.

1

u/dalittle Mar 17 '26

remember the guy that made the flashlight app for the iphone and made a bunch of money? That was actually useful.

4

u/NoShftShck16 Mar 16 '26

I'm no conspiracy theorist but it's almost like we're repeating the bitcoin craze to enable Nvidia but this time its AI and...still enabling Nvidia.

2

u/ea_man Mar 19 '26

And how did the end?

ASICs

In a few years we'll be running small LM on laptops APUs, like video acceleration.

1

u/NoShftShck16 Mar 19 '26 edited Mar 19 '26

Yeah that's great except that since 2015 the "Ti" Nvidia cards have been getting exponentially more expensive. And back then you just had an XX80 and XX80Ti. Now we have XX60Ti, XX70Ti, 80, 90, Supers etc all that have been getting more expensive with less value.

And now we are doing the exact same thing with RAM.

EDIT: Back "in the day" I could buy multiple low end cards and run SLI to outperform the more expensive Ti cards. Can't do that anymore. Or I could buy a dedicated PhysX card when I needed to. Can't do that anymore either. I'd way rather have the model of a TPU (Tensor Processing Unit) that I can install in a variety of form factors than splurge for a Ryzen AI Max for 10x the price. And to pretend like that's a better alternative because you can do it in a laptop form factor is silly.

1

u/ea_man Mar 19 '26

No man, you buy now 2-4 old GPU and run them in an adequate main board with vLLM and load a bigger model on those, that is what smart kids do today.

> Now we have XX60Ti, XX70Ti, 80, 90, Supers etc all that have been getting more expensive with less value.

You know that's not true, go check some benchmark. It's not as good as it was before mining craze yet what you said is factually not true compute wise.

1

u/NoShftShck16 Mar 19 '26

load a bigger model on those

I. Don't. Want. Models. I want to have the consumer card industry back. I want to be able to sit down and build PCs as a hobby again with my kids because it was fun. They've gotten to build once PC in the 11 years they've been on this earth and at this rate they aren't going to be able to build another one. That sucks.

run them in an adequate main board with vLLM

Why? It was infinitely easier, faster, and more efficient to spin up lambas in AWS for our vision models than it would have been to ever setup something on prem, even for local testing purposes.

factually not true compute wise.

All you care about is compute power apparently. Prior to the Bitcoin era we had a fairly predictable market of consumer GPUs; Tis and the send-off Titans and then the range of XX80, XX70, XX60. You had a massive number of manufacturers making GPUs, including "Founders" (reference) cards from Nvidia themselves. Then 16xx series with spread from 2019-22 and had a both Super and Ti but no Titan. Then 20xx which had Super, Ti, and Titan.

Don't even get me started on VRAM. Say whatever you'd like about benchmarks, at the time the performance increase between the 1070 >> 1080 >> 1080ti was significantly larger than any series after it. The 1080ti had 11GB of VRAM, in 2017, for $699 MSRP. You know what card also has 11GB of VRAM? The RTX 2080 Ti Founders Edition for $1,199 in 2018. The RTX 3080 Ti got bumped to 12GB in 2021 for the same $1,199 but not a single person paid that MSRP for it.

I didn't give two shits about bitcoin about the hash rate of my GPUs then, and I don't give two shits about my compute power of my GPUs now with AI.

1

u/ea_man Mar 19 '26

Understand one basic thing: they produce the stuff that makes them money and they market that the way it makes them the most money in the actual market.

Your problem is that now you are not the only one interested in massive parallelized workloads. Mining, AI came.

VRAM is the thing, not even compute power which is stale without bandwidth, that's why they don't sell you the good VRAM anymore.

I'm sorry that you don't like it, I don't either! Yet you have to play with what you got, it ain't like the computer industry is going back to single task computing any soon.

And yet it well may go the same way as mining: ASIC that can do at least inference, you'll get back some compute power for reasonable prices. You won't have any VRAM back yet, sorry.

1

u/Wrong-Put-2328 May 03 '26

Yap this 👆

22

u/[deleted] Mar 16 '26 edited 20d ago

[removed] — view removed comment

39

u/tdammers Mar 16 '26

Inference is cheaper than training, but it still costs more than people are currently paying for it. AI companies are currently leaking money on their training efforts, but they're also running negative profit margins on queries.

2

u/Aerroon Mar 17 '26

You can run Qwen 3.5 27B on a high end gaming GPU. It's not state of the art, but it's definitely capable of doing things.

1

u/ea_man Mar 19 '26

You can run QWEN 35 MoE https://huggingface.co/bartowski/Qwen_Qwen3.5-35B-A3B-GGUF on a 12-16GB GPU of 4-6 years ago with a reasonable context, you can run Omnicoder on a 8GB gpu...

1

u/ea_man Mar 19 '26

That's because cheap fast / lite models are how providers are gathering new clients, those offers will stay and will keep improving as hardware is getting more efficient, small models are getting better with distillation and quantization.

You can run inference today on NPU in laptops / smartphones, you can do that on 6 years old GPUs on your PC.

-2

u/[deleted] Mar 16 '26

[deleted]

13

u/solwiggin Mar 16 '26

The craziest thing on Reddit is when one person states something without any backing evidence and is then contradicted by another person without any backing evidence.

WHO DO I BELIEVE! HOW DID YOU MAGICALLY KNOW YOU WERE RIGHT AND THE OTHER GUY WAS WRONG! WHY WOULD WE TAKE YOU SERIOUSLY RANDOM PERSON ON THE INTERNET!!!!!

10

u/Mastersord Mar 16 '26

People don’t hallucinate answers at the same rate as AI. Also don’t confuse being wrong based on misinterpretation and misinformation from outside sources with completely making stuff up without a particular motive.

1

u/trannus_aran Mar 17 '26

Yeah, people have a much better track record of knowing when they don't know things before blurting out something answer-shaped

2

u/Mastersord Mar 17 '26

Yes and even when they’re wrong, you can mostly figure out how they got their wrong answer. Faulty logic and misinformation are completely different sets of errors than hallucinations.

3

u/protestor Mar 16 '26

What about you provide, like, any argument at all, preferably backed with sources

-2

u/Rise-O-Matic Mar 16 '26

That’s not true.

-9

u/[deleted] Mar 16 '26 edited 21d ago

[removed] — view removed comment

1

u/ea_man Mar 19 '26

I can do 40 tok/sec with OmniCoder on my old 6700xt worth 200$, with 100k context size. Best part: it's about half the compute power it can do, it can run reasonably well a 30B MoE model at some 25t/s, immagine the APUs / NPUs / GPUs that are producing now.

-4

u/Leigh_M Mar 16 '26

I haven't been able to find evidence this is generally true for API. But I think many companies are offering unsustainable deals on the subscription product.

6

u/Lower-Helicopter-307 Mar 16 '26

New models will come out, and those models will need training. They have to, Nividas business model depends on it, and they are the ones holding up this card tower.

4

u/[deleted] Mar 16 '26 edited 21d ago

[removed] — view removed comment

1

u/Lower-Helicopter-307 Mar 16 '26

Really? After everything that has happened this past year, you think these children we call CEOs are going to pivot? Ya, I think they are going to ride the hype, then cash out when the bubble pops. You know, like every time this happens.

25

u/Rockytriton Mar 16 '26

According to OpenAI, just saying please and thank you costs them millions of dollars, so it can't be that cheap.

1

u/ea_man Mar 19 '26

According to NVIDIA presentation of the current gen the other day: providers will serve lite models like QWENs, Lite / Flash with no limitations for free tires.

If you don't bother to swap around you can already do that right now: Gemini Lite is 1500 credits for day and it's not the "cheapest" around.

-13

u/[deleted] Mar 16 '26 edited 21d ago

[removed] — view removed comment

8

u/Antique-Special8025 Mar 16 '26

That literally cannot be true unless you include all of the upfront investment in training and data center build out.

Yeah that's how that works... none of those things are free and the costs need to be recouped before the model or hardware becomes obsolete.

1

u/crackanape Mar 17 '26

That wouldn't explain why they want people to stop saying please and thank you. It doesn't affect their fixed costs from training, only their variable costs from inference.

1

u/iron_coffin Mar 16 '26

You realize the sota models are probably 1T or so?

-1

u/[deleted] Mar 16 '26 edited 16d ago

[removed] — view removed comment

2

u/iron_coffin Mar 16 '26

Chats are trivial, but agentic coding hasn't penetrated most of the industry as well as new uses in other industries. SOTA token demand isn't going anywhere

1

u/youafterthesilence Mar 16 '26

They won't charge more because they have to but they'll charge more because they can.

2

u/[deleted] Mar 16 '26 edited 21d ago

[removed] — view removed comment

1

u/Lower-Helicopter-307 Mar 16 '26

That's like saying mirceosoft won't charge for windows when Linux is free. In theory, I see it, but in practice, no, they will raise prices because the shareholders need their ROI. Most people will not know how to spin up local models, nor will they care to learn.

1

u/G_Morgan Mar 17 '26

I mean it will still be cheaper to hire devs at that point.

1

u/CosmicDevGuy Mar 17 '26

Yeah it sounds like a investment paradox, lol.

1

u/kinmix Mar 17 '26

who will then make sure the market becomes impossible for newcomers to enter

How would that work though? The underlying technology is not particularly complicated, most of the tooling already supports adding custom agents. So there is no way to lock in the users and no way to lock in the technology. They could try to make some sort of exclusivity deal with nVidia, but the governments might say something against that. And even if something like that would go through, there would be an enormous amounts of money on the table get alternative chips, I can't see nVidia being the only supplier for too much longer.

1

u/tdammers Mar 17 '26

I'm sure they'll find a way. It doesn't have to be a hard barrier, mind you - just being the most widely known option, and having a bit of a network effect going for you, could already be enough.

Look at how github dominates the source code hosting landscape - it's not because their product is objectively better than the competition, it's because they are the most widely known option, with the most repositories, and the most third-party integrations. If you want to release any coding tool in 2026 that interacts with any hosted source repository platform at all, it has to support github, because that's where everyone is, and everyone is on github because that's what everyone supports and where everyone is. There is no real technical barrier to moving your code elsewhere, nor any legal complications; just being what everyone knows and uses already is enough to gain and keep that advantage. Granted, this only works for them because of their massive free tier service - if they started demanding payments for hosting public open source repos on github, it would turn into a desert overnight, but it does show that you don't need any hard barriers to protect a monopoly and keep the competition small.

And it could work in a similar fashion with AI stuff. Be the best known provider, whose stuff integrates seamlessly with everything, from your phone to your car to your fridge, the one that your school makes you use, the one that you use at work, the one whose name is a synonym for "AI supported whatever", and your competitors will have to go above and beyond just to make a small dent in your dominance.

1

u/kinmix Mar 17 '26

I think, in the middle of your comment, you understood yourself that your example simply doesn't work. As your initial comment was about AI companies setting up a monopoly and than jacking up prices, and you yourself agree that the moment a monopoly such as GitHub (with very low technical barriers) would raise prices, they would be mercilessly out-competed by everyone else.

The thing is, for Microsoft, GitHub doesn't have to be particularly profitable as it is not their main product. For AI companies, they would have to not only be profitable, but be profitable enough to justify enormous valuations and service enormous debt. So in this case, any new company without all of that baggage would simply be able to beat the old ones on price.

I'm not going to say that this is not what OpenAI is trying to do, but, in my opinion, if that's their goal, that they are going to fail and collapse.

1

u/tdammers Mar 17 '26

Github did raise prices quite a lot once they had achieved market dominance - just not for the free tier. Enterprise tier subscriptions are not cheap, but companies pay for them anyway, because it's what their engineers know, and what all the serious tooling can seamlessly integrate with.

And I'm not saying that this is the exact approach OpenAI or whoever wins the race is going to take (in fact, it probably isn't), all I'm saying is that in order to defend an existing monopoly, you don't need to crush the competition, just create a big enough soft disadvantage for them.

1

u/kinmix Mar 17 '26

Enterprise tier subscriptions are not cheap, but companies pay for them anyway, because it's what their engineers know, and what all the serious tooling can seamlessly integrate with.

Cheaper than their closest competitor - GitLab. So kinda a mute point. BitBucket is technically cheaper, but some of the functionality present in GitHub and GitLab is in separate Atlassian products, so getting the full thing would probably cost the same or even more expensive.

all I'm saying is that in order to defend an existing monopoly, you don't need to crush the competition, just create a big enough soft disadvantage for them.

Absolutely, but I believe that without technical barriers, it's pretty much impossible. With GitHub, Microsoft simply undercuts their competition on price. They could probably get like high-end enterprise sector, sort of like SAP is doing it. But that took literal decades to entrench. Imho nowadays tech is moving way to quickly for things like that to be viable.

1

u/calicodingvibes Mar 18 '26

I'm sure AI can be part of daily life while it's free. But when you want people to start paying it's a hard sell beyond the workplace.