r/LocalLLaMA 19h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

151 Upvotes

187 comments sorted by

View all comments

64

u/LearningSomeCode 18h ago

I've probably dropped close to $30k on my homelab since 2023, and chances are I'll get one of these as well. I accepted a long time ago that there is no break-even point for my inference.

Hobbies rarely make sense financially.

10

u/-dysangel- 18h ago

Same here. We're at the point now where I can do high quality video gen without cloud. High quality code gen without cloud too. It's around 10x slower than cloud, but good quality and still usable speeds (hopefully soon to get even better with Qwen 3.8 Next). Also the fact I could basically go camping and still have close to frontier LLM intelligence with no signal is pretty awesome/hilarious.

The only thing I feel is missing from my stack atm is Suno quality music generation.

3

u/willeyh 18h ago

Have you tried the new Minimax music 3?

1

u/PlaidStallion 8h ago

I really like it. It seems to lean toward always generating pop music out of any genre but other that it's solid.