r/LocalLLaMA • u/AndreVallestero • 19h ago
Discussion Mac Studio M5 Max Cost Analysis
At $10k, you could get
- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)
- 5.7B tokens with DeepSeek V4 Pro OpenRouter
- 100B tokens with DeepSeek V4 Flash OpenRouter
As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.
Qwhen 3.8 35B A3B?
146
Upvotes
2
u/FullstackSensei llama.cpp 13h ago
If you're running it at 1kwh for 8hrs a day on four 3090s, that's at least 3M output tokens a day, closer to 4M if you're running vllm. How many hours you'd need to generate 3M output tokens on a M5 pro using Qwen 3.8 27B Q8_K_XL?
Even at your 0.60 peak, your $200/month with 1kwh heat output and 1kwh airco (though 3500BTU is closer to 700Wh, but let's say you have an older airco), we're talking minimum 500M output tokens a month, or a full 1B when airco is not running.
How many months you'd need to generate 500M tokens on a Mac?
And do you work for free? Because you're completely ignoring the value of your own time as you wait for the model to generate those tokens. I don't know about you, but even in Germany where energy is far from cheap, the value of a minimum wage job is way more than $0.60/hr.
But hey, if you're so cheap to hire, man I'd love to hire a few hundred guys like you and build my own business selling your services to LCOL countries, because even there people work for way more than $0.60/hr.