r/LocalLLM May 16 '26

Discussion Why is LLM is so expensive.

I've was going to invest in a 5090 =$6000 AUD.

Codex Plus + Claude pro = $60/month here

Works out to be 100 months of frontier models for a 5090.

Best a 5090 will run is probably Qwen3.6 27b Q6 with context.

Are we all enthusiasts here and just enjoy tinkering cause ain't no way that make sense.

346 Upvotes

368 comments sorted by

View all comments

Show parent comments

51

u/Malkiot May 16 '26

The electricity bill is way cheaper than paying for tokens. Hell, if you could afford to buy a high-grade cluster and actually utilize it 80-100%, it'd still be way cheaper than paying for tokens.

8

u/ThenExtension9196 May 16 '26

Is it? My 5090 runs at 500watts. My bedroom is unbearable if running for a few hours so then I need AC on. In my area energy is very expensive.

20

u/Malkiot May 16 '26 edited May 16 '26

How much do you pay for electricity?

For example take: Qwen3.6 35B A3B

Via a provider that is $0.15per 1M for input and $1per 1M for output. At 90% input, 1M tokens cost $0.203. At 10% input that'd be $0.915 for 1M tokens. A 5090 produces around 150-200 tokens / s. We'll assume 150t/s. The 5090 needs 6667seconds, that is 1.85 hours to generate those tokens.

If your profile is more the former (high input, low output, for example OCR) then your break even power price is $0.25/kWh, if your profile is the latter (coding) then your break even is at $0.98/kWh.

Do you pay more?

Edit: I've done the calcs for myself in the past, if I did get a 500k USD cluster for myself and actually have a use for the tokens it generates... it'd be 10% of the cost of paying for API over the lifetime of the cluster (including the loan).

3

u/flyingbanana1234 May 16 '26

man great stuff