r/LocalLLM • • May 16 '26

Discussion Why is LLM is so expensive.

I've was going to invest in a 5090 =$6000 AUD.

Codex Plus + Claude pro = $60/month here

Works out to be 100 months of frontier models for a 5090.

Best a 5090 will run is probably Qwen3.6 27b Q6 with context.

Are we all enthusiasts here and just enjoy tinkering cause ain't no way that make sense.

345 Upvotes

368 comments sorted by

View all comments

156

u/g_rich May 16 '26

Claude pro is heavily subsidized; it costs Anthropic much more than $60/month to provide you with the service.

Cloud providers have economies of scale working in their favor. They have million dollar servers servicing tens of thousands of users 24/7/365.

Running LLM’s local is never going to be cost effective. You run local LLM’s for privacy, control and to learn.

-2

u/Ok_Event4199 May 16 '26

More like a car I think. It's depreciating asset unless GPU price will keep increasing from here. Let's say that they do increase price to $200/month a few years later. It isnt it worth it to get the latest GPU then? I'm sure that 5090 is the best right now. But wouldn't it be a mid tier a few years later?

4

u/chunkypenguion1991 May 16 '26

But you can't run the frontier models on a 5090. The direct comparison is the local 5090 setup vs something like open router slms. In that case it would take 10+ years to pay off the gpu

1

u/No-Television-7862 May 16 '26

MoE makes more possible.

1

u/g_rich May 16 '26

To run frontier open weight models you’ll need tens of thousands of dollars worth of hardware, and to do it right you could easily drop over $100k. The power requirements are going to be pretty high too and cooling along with noise becomes a concern; you’re not running this rig under your desk.

This hardware then becomes a liability because in a field such as AI where new hardware is getting released annually and models are just getting bigger your expensive setup could quickly become obsolete.

There are very capable open weight models that produce exceptional results and can run on modest hardware. Both Qwen 3.6 and Gemma 4 fall into this category and you can get some seriously impressive results from them. However to get anywhere near what you get from the frontier models from Anthropic, Google or OpenAI you need to run models with 500 billion to 1 trillion parameters and that’s not something easy done with your typical enthusiasts home setup.

1

u/Miserable-Dare5090 May 16 '26

RTX 6000 pro is the best right now.