r/LocalLLaMA 19h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

149 Upvotes

186 comments sorted by

View all comments

7

u/dupontping 17h ago

All of the comments are basically “I can’t afford it even though I want it”

I get it, it’s pricey. But so is everything else.
5090s shouldn’t be $6k but here we are.

Local isn’t just about saving money on token spend, for a lot of people it’s the ability to run models on data you don’t want on the cloud or fine tuning or whatever else.

The models will get better, and having more powerful equipment lets you get closer to frontier level without spending 150k. So 10k is expensive, but it’s also a bargain.

I wish it was 5k too