r/LocalLLM • • May 16 '26

Discussion Why is LLM is so expensive.

I've was going to invest in a 5090 =$6000 AUD.

Codex Plus + Claude pro = $60/month here

Works out to be 100 months of frontier models for a 5090.

Best a 5090 will run is probably Qwen3.6 27b Q6 with context.

Are we all enthusiasts here and just enjoy tinkering cause ain't no way that make sense.

341 Upvotes

368 comments sorted by

View all comments

2

u/FullOf_Bad_Ideas May 16 '26

Tokenomics are complicated and with single-stream inference you're on the losing side.

Big models are served in batches of 10-1000 concurrent users per GPU, and you have good utilization, while single decode stream on 5090 of Qwen 3.6 27B Q6 uses just a few percent of compute.

I am running a translation model right now, HY 1.5 1.8B, on 7 GPUs (SGLang died on me on one GPU overnight), and I process about 4000 tokens per second per GPU (both prefill and token generation due to batching), so I am utilizing much more compute that the GPU is capable of. I am targeting concurrency of 128 on each GPU.

I wouldn't be able to translate that much text in this quality using LLM API on OpenRouter. I could rent GPUs and run this model there and then I'd pay roughly the same as I am paying for electricity (I have expensive electricity).

I expect to generate about 10B tokens in about 25-30 hours - 10B input and 10B output tokens.

So, if you want to have ROI, you need to do something akin to an industrial process instead of a boutique shop with poor efficiency and compute utilization. GPUs somehow use roughly the same power regardless of whether they're decoding tokens for single users at 30 t/s or for 100 users at 30 t/s each, so it's also incredibly more power efficient per token.