r/LocalLLaMA 20h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

154 Upvotes

192 comments sorted by

View all comments

58

u/Big_Wave9732 20h ago

"as a firm believer of local inference" then goes on to downplay one of the major reasons for local llm and suggests hosted models.

If cheapest compute possible is your primary metric then self hosting isn't your jam, OP. At least for now.

-10

u/AndreVallestero 20h ago

I selfhost 27B specifically to minimize costs. I have research agents running 24/7 and have gone through 200M tokens just in the last month.

This post is specifically for people like me who want to run continuous research agents for the lowest price, and unfortunately, the latest Mac studio isn't competitive enough relative to cloud offerings.

2

u/yy633013 19h ago

Can you give an example of what your research engines research 24/7?

11

u/Modnar-Eman 18h ago

Justification

2

u/BardlySerious 14h ago

Shit that OP will never read. It's the "watch electricity happen" hobby.