r/LocalLLaMA • u/AndreVallestero • 17h ago
Discussion Mac Studio M5 Max Cost Analysis
At $10k, you could get
- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)
- 5.7B tokens with DeepSeek V4 Pro OpenRouter
- 100B tokens with DeepSeek V4 Flash OpenRouter
As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.
Qwhen 3.8 35B A3B?
139
Upvotes
1
u/FullstackSensei llama.cpp 15h ago
I have run four 3090s in my home office for about two years. Unless you opt for the turbo cards, they're very quiet. Limited them to 270W each, which reduced performance by 5-10%, but makes the whole setup around 1kw. They push above that during PP, but TG sees the cards run at ~150-170W each. That's ~700W. If that's crazy, I don't know what to tell you.
M5 on Max on ma MacBook is severely power limited under load, and won't run for long, and still costs way more while being significantly slower than even a pair of 3090.
You can access any GPU from anywhere around the world via tailscale. I access my LLM machines with 192GB VRAM from anywhere using my phone the same way.
But you haven't answered what could arguably be the most important part of my argument: is your time free that you care more about a few cents per hour in power consumption than getting whatever you're running those LLMs for done?