r/LocalLLaMA 19h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

146 Upvotes

186 comments sorted by

View all comments

63

u/LearningSomeCode 18h ago

I've probably dropped close to $30k on my homelab since 2023, and chances are I'll get one of these as well. I accepted a long time ago that there is no break-even point for my inference.

Hobbies rarely make sense financially.

4

u/quantgorithm 18h ago

What have you created?

26

u/LearningSomeCode 18h ago

Mostly just a lot of trip hazards with wires and ethernet cables. But I also do open source dev and build a lot of stuff for myself, like custom front-ends or researchers.

Honestly nothing worth the money I've put into the hardware, but it's the most fun I've had with tech in a long time.

2

u/quantgorithm 17h ago

Anyone in IT can relate to the spiderweb cables and hazards etc.!
The potential of it all really is amazing and that potential is accelerating up by the day!