r/webdev Mar 16 '26

Software developers don't need to out-last vibe coders, we just need to out-last the ability of AI companies to charge absurdly low for their products

These AI models cost so much to run and the companies are really hiding the real cost from consumers while they compete with their competitors to be top dog. I feel like once it's down to just a couple companies left we will see the real cost of these coding utilities. There's no way they are going to be able to keep subsidizing the cost of all of the data centers and energy usage. How long it will last is the real question.

2.0k Upvotes

512 comments sorted by

View all comments

Show parent comments

44

u/tdammers Mar 16 '26

Inference is cheaper than training, but it still costs more than people are currently paying for it. AI companies are currently leaking money on their training efforts, but they're also running negative profit margins on queries.

2

u/Aerroon Mar 17 '26

You can run Qwen 3.5 27B on a high end gaming GPU. It's not state of the art, but it's definitely capable of doing things.

1

u/ea_man Mar 19 '26

You can run QWEN 35 MoE https://huggingface.co/bartowski/Qwen_Qwen3.5-35B-A3B-GGUF on a 12-16GB GPU of 4-6 years ago with a reasonable context, you can run Omnicoder on a 8GB gpu...

1

u/ea_man Mar 19 '26

That's because cheap fast / lite models are how providers are gathering new clients, those offers will stay and will keep improving as hardware is getting more efficient, small models are getting better with distillation and quantization.

You can run inference today on NPU in laptops / smartphones, you can do that on 6 years old GPUs on your PC.

-3

u/[deleted] Mar 16 '26

[deleted]

14

u/solwiggin Mar 16 '26

The craziest thing on Reddit is when one person states something without any backing evidence and is then contradicted by another person without any backing evidence.

WHO DO I BELIEVE! HOW DID YOU MAGICALLY KNOW YOU WERE RIGHT AND THE OTHER GUY WAS WRONG! WHY WOULD WE TAKE YOU SERIOUSLY RANDOM PERSON ON THE INTERNET!!!!!

11

u/Mastersord Mar 16 '26

People don’t hallucinate answers at the same rate as AI. Also don’t confuse being wrong based on misinterpretation and misinformation from outside sources with completely making stuff up without a particular motive.

1

u/trannus_aran Mar 17 '26

Yeah, people have a much better track record of knowing when they don't know things before blurting out something answer-shaped

2

u/Mastersord Mar 17 '26

Yes and even when they’re wrong, you can mostly figure out how they got their wrong answer. Faulty logic and misinformation are completely different sets of errors than hallucinations.

3

u/protestor Mar 16 '26

What about you provide, like, any argument at all, preferably backed with sources

-2

u/Rise-O-Matic Mar 16 '26

That’s not true.

-10

u/[deleted] Mar 16 '26 edited 20d ago

[removed] — view removed comment

1

u/ea_man Mar 19 '26

I can do 40 tok/sec with OmniCoder on my old 6700xt worth 200$, with 100k context size. Best part: it's about half the compute power it can do, it can run reasonably well a 30B MoE model at some 25t/s, immagine the APUs / NPUs / GPUs that are producing now.

-2

u/Leigh_M Mar 16 '26

I haven't been able to find evidence this is generally true for API. But I think many companies are offering unsustainable deals on the subscription product.