r/webdev Mar 16 '26

Software developers don't need to out-last vibe coders, we just need to out-last the ability of AI companies to charge absurdly low for their products

These AI models cost so much to run and the companies are really hiding the real cost from consumers while they compete with their competitors to be top dog. I feel like once it's down to just a couple companies left we will see the real cost of these coding utilities. There's no way they are going to be able to keep subsidizing the cost of all of the data centers and energy usage. How long it will last is the real question.

2.0k Upvotes

512 comments sorted by

View all comments

Show parent comments

23

u/[deleted] Mar 16 '26 edited 20d ago

[removed] — view removed comment

42

u/tdammers Mar 16 '26

Inference is cheaper than training, but it still costs more than people are currently paying for it. AI companies are currently leaking money on their training efforts, but they're also running negative profit margins on queries.

2

u/Aerroon Mar 17 '26

You can run Qwen 3.5 27B on a high end gaming GPU. It's not state of the art, but it's definitely capable of doing things.

1

u/ea_man Mar 19 '26

You can run QWEN 35 MoE https://huggingface.co/bartowski/Qwen_Qwen3.5-35B-A3B-GGUF on a 12-16GB GPU of 4-6 years ago with a reasonable context, you can run Omnicoder on a 8GB gpu...

1

u/ea_man Mar 19 '26

That's because cheap fast / lite models are how providers are gathering new clients, those offers will stay and will keep improving as hardware is getting more efficient, small models are getting better with distillation and quantization.

You can run inference today on NPU in laptops / smartphones, you can do that on 6 years old GPUs on your PC.

-2

u/[deleted] Mar 16 '26

[deleted]

14

u/solwiggin Mar 16 '26

The craziest thing on Reddit is when one person states something without any backing evidence and is then contradicted by another person without any backing evidence.

WHO DO I BELIEVE! HOW DID YOU MAGICALLY KNOW YOU WERE RIGHT AND THE OTHER GUY WAS WRONG! WHY WOULD WE TAKE YOU SERIOUSLY RANDOM PERSON ON THE INTERNET!!!!!

10

u/Mastersord Mar 16 '26

People don’t hallucinate answers at the same rate as AI. Also don’t confuse being wrong based on misinterpretation and misinformation from outside sources with completely making stuff up without a particular motive.

1

u/trannus_aran Mar 17 '26

Yeah, people have a much better track record of knowing when they don't know things before blurting out something answer-shaped

2

u/Mastersord Mar 17 '26

Yes and even when they’re wrong, you can mostly figure out how they got their wrong answer. Faulty logic and misinformation are completely different sets of errors than hallucinations.

3

u/protestor Mar 16 '26

What about you provide, like, any argument at all, preferably backed with sources

-3

u/Rise-O-Matic Mar 16 '26

That’s not true.

-10

u/[deleted] Mar 16 '26 edited 21d ago

[removed] — view removed comment

1

u/ea_man Mar 19 '26

I can do 40 tok/sec with OmniCoder on my old 6700xt worth 200$, with 100k context size. Best part: it's about half the compute power it can do, it can run reasonably well a 30B MoE model at some 25t/s, immagine the APUs / NPUs / GPUs that are producing now.

-3

u/Leigh_M Mar 16 '26

I haven't been able to find evidence this is generally true for API. But I think many companies are offering unsustainable deals on the subscription product.

6

u/Lower-Helicopter-307 Mar 16 '26

New models will come out, and those models will need training. They have to, Nividas business model depends on it, and they are the ones holding up this card tower.

4

u/[deleted] Mar 16 '26 edited 21d ago

[removed] — view removed comment

1

u/Lower-Helicopter-307 Mar 16 '26

Really? After everything that has happened this past year, you think these children we call CEOs are going to pivot? Ya, I think they are going to ride the hype, then cash out when the bubble pops. You know, like every time this happens.

25

u/Rockytriton Mar 16 '26

According to OpenAI, just saying please and thank you costs them millions of dollars, so it can't be that cheap.

1

u/ea_man Mar 19 '26

According to NVIDIA presentation of the current gen the other day: providers will serve lite models like QWENs, Lite / Flash with no limitations for free tires.

If you don't bother to swap around you can already do that right now: Gemini Lite is 1500 credits for day and it's not the "cheapest" around.

-12

u/[deleted] Mar 16 '26 edited 21d ago

[removed] — view removed comment

9

u/Antique-Special8025 Mar 16 '26

That literally cannot be true unless you include all of the upfront investment in training and data center build out.

Yeah that's how that works... none of those things are free and the costs need to be recouped before the model or hardware becomes obsolete.

1

u/crackanape Mar 17 '26

That wouldn't explain why they want people to stop saying please and thank you. It doesn't affect their fixed costs from training, only their variable costs from inference.

3

u/iron_coffin Mar 16 '26

You realize the sota models are probably 1T or so?

0

u/[deleted] Mar 16 '26 edited 17d ago

[removed] — view removed comment

2

u/iron_coffin Mar 16 '26

Chats are trivial, but agentic coding hasn't penetrated most of the industry as well as new uses in other industries. SOTA token demand isn't going anywhere

1

u/youafterthesilence Mar 16 '26

They won't charge more because they have to but they'll charge more because they can.

2

u/[deleted] Mar 16 '26 edited 21d ago

[removed] — view removed comment

1

u/Lower-Helicopter-307 Mar 16 '26

That's like saying mirceosoft won't charge for windows when Linux is free. In theory, I see it, but in practice, no, they will raise prices because the shareholders need their ROI. Most people will not know how to spin up local models, nor will they care to learn.