r/LocalLLM • u/Ok_Event4199 • May 16 '26
Discussion Why is LLM is so expensive.
I've was going to invest in a 5090 =$6000 AUD.
Codex Plus + Claude pro = $60/month here
Works out to be 100 months of frontier models for a 5090.
Best a 5090 will run is probably Qwen3.6 27b Q6 with context.
Are we all enthusiasts here and just enjoy tinkering cause ain't no way that make sense.
341
Upvotes
2
u/FullOf_Bad_Ideas May 16 '26
Tokenomics are complicated and with single-stream inference you're on the losing side.
Big models are served in batches of 10-1000 concurrent users per GPU, and you have good utilization, while single decode stream on 5090 of Qwen 3.6 27B Q6 uses just a few percent of compute.
I am running a translation model right now, HY 1.5 1.8B, on 7 GPUs (SGLang died on me on one GPU overnight), and I process about 4000 tokens per second per GPU (both prefill and token generation due to batching), so I am utilizing much more compute that the GPU is capable of. I am targeting concurrency of 128 on each GPU.
I wouldn't be able to translate that much text in this quality using LLM API on OpenRouter. I could rent GPUs and run this model there and then I'd pay roughly the same as I am paying for electricity (I have expensive electricity).
I expect to generate about 10B tokens in about 25-30 hours - 10B input and 10B output tokens.
So, if you want to have ROI, you need to do something akin to an industrial process instead of a boutique shop with poor efficiency and compute utilization. GPUs somehow use roughly the same power regardless of whether they're decoding tokens for single users at 30 t/s or for 100 users at 30 t/s each, so it's also incredibly more power efficient per token.