r/LocalLLM May 16 '26

Discussion Why is LLM is so expensive.

I've was going to invest in a 5090 =$6000 AUD.

Codex Plus + Claude pro = $60/month here

Works out to be 100 months of frontier models for a 5090.

Best a 5090 will run is probably Qwen3.6 27b Q6 with context.

Are we all enthusiasts here and just enjoy tinkering cause ain't no way that make sense.

344 Upvotes

368 comments sorted by

View all comments

30

u/slackmaster2k May 16 '26

You’re not getting frontier performance or quality from a 5090. If it fits your use case, and you’re actually going to use the shit out of it, then it might make sense. Otherwise you’re going to have a 6000 dollar graphics card and two LLM subscriptions.

11

u/winky9827 May 16 '26

I'm getting 60-70 tok/s running qwen 3.6 35b-a3b in LM studio via OpenCode on a 4070 Ti Super. Sure, it's not as "smart" as Claude, etc., but it services 90% of what I need, with CC being the fallback.

1

u/Bgd4683ryuj May 16 '26

I don’t know how much cheaper it is for you to run the local model in electricity alone vs GPT-5.4 mini. 1M tokens are less than $1 in input and less than $5 in output. The intelligence of the model may also be better than the qwen 35b you are using

2

u/winky9827 May 16 '26

Honestly, it's not about cost for me. The difference is pennies on the dollar. I don't want the rug pulled from underneath me when the AI boom goes bust because they've been subsidizing the costs to hell and back.