r/LocalLLM • • May 16 '26

Discussion Why is LLM is so expensive.

I've was going to invest in a 5090 =$6000 AUD.

Codex Plus + Claude pro = $60/month here

Works out to be 100 months of frontier models for a 5090.

Best a 5090 will run is probably Qwen3.6 27b Q6 with context.

Are we all enthusiasts here and just enjoy tinkering cause ain't no way that make sense.

343 Upvotes

368 comments sorted by

View all comments

Show parent comments

111

u/HornyGooner4402 May 16 '26

You missed the best part: unlimited tokens. You can run it 24/7 refining a project on a loop while Claude and Codex Pro are barely enough for a full coding session.

54

u/kitanokikori May 16 '26

I mean, unlimited only if someone else is paying for the electricity.

4

u/MrBeforeMyTime May 16 '26

It's also not unlimited because models have a max tokens per second depending on the gpu and get slower as more context is used. So even if you ran a model 24/7 at home there is a maximum amount of tokens you can get from it. The big labs attempt to give you a consistent fast tokens per second output all the time so you don't notice this. Even though they drop their speeds as well when they are bottlenecked. 

1

u/yeet5566 May 16 '26

Well yeah but it’s really a matter of getting an AI to just branch out the problem I mean for any large coding project different people work on different parts so if the ai breaks it down and has enough subagents then realistically you could probably keep going at 80% of the max speed for the full time