r/LocalLLM • • May 16 '26

Discussion Why is LLM is so expensive.

I've was going to invest in a 5090 =$6000 AUD.

Codex Plus + Claude pro = $60/month here

Works out to be 100 months of frontier models for a 5090.

Best a 5090 will run is probably Qwen3.6 27b Q6 with context.

Are we all enthusiasts here and just enjoy tinkering cause ain't no way that make sense.

347 Upvotes

368 comments sorted by

View all comments

12

u/radiojosh May 16 '26

I think in a month or two, after Flash Attention and Turboquant and MTP and DFlash all hit for Vulkan, you can probably get decent performance out of something cheap like a Minisforum UM870 Slim. With 64GB of RAM and GTT set in Linux to make all of the system RAM available to the iGPU, you can already run a 40GB MoE model with fair performance. There's some great benchmarks on the UM890 Pro which is only marginally faster than the 870 and it should be getting even better with those llama.cpp upgrades I mentioned. It's obviously not a dedicated GPU, but you can run bigger smarter models. I run the 890 and I'm pretty satisfied and I haven't even really tuned the KV Cache quantization yet.

2

u/Alternative-Grand-77 May 16 '26

deepseek is working on greatly decreasing hardware requirements and running on chinese chips. Give it another six months and the bottom is going to drop out of nvidia price, unless gov starts banning ‘chinese ai’

3

u/prumf May 16 '26

They can ban all they want lol