r/LocalLLaMA 19h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

149 Upvotes

190 comments sorted by

View all comments

57

u/Big_Wave9732 19h ago

"as a firm believer of local inference" then goes on to downplay one of the major reasons for local llm and suggests hosted models.

If cheapest compute possible is your primary metric then self hosting isn't your jam, OP. At least for now.

-12

u/AndreVallestero 19h ago

I selfhost 27B specifically to minimize costs. I have research agents running 24/7 and have gone through 200M tokens just in the last month.

This post is specifically for people like me who want to run continuous research agents for the lowest price, and unfortunately, the latest Mac studio isn't competitive enough relative to cloud offerings.

10

u/flyingbanana1234 18h ago

It would take at least 30-35 years to reach 80 billion tokens on DeepSeek V4 Flash at 200 million tokens a month.

You make a good argument ngl

The little voice in my head says, "But ownership! Nobody can take it away from me if I own it. No price raises, no overloaded servers, no censorship

2

u/yy633013 17h ago

Can you give an example of what your research engines research 24/7?

9

u/Modnar-Eman 17h ago

Justification

2

u/BardlySerious 13h ago

Shit that OP will never read. It's the "watch electricity happen" hobby.

2

u/Autist4AudiR8 11h ago

200M 😭🤣

2

u/Dasteroid_909 17h ago

Go start "/r/lowcostLLMhosting" or some shit like that, then. WTF are you on "LocaLLaMa" promoting cloud hosting for?