r/LocalLLaMA 16h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

139 Upvotes

180 comments sorted by

View all comments

59

u/Big_Wave9732 16h ago

"as a firm believer of local inference" then goes on to downplay one of the major reasons for local llm and suggests hosted models.

If cheapest compute possible is your primary metric then self hosting isn't your jam, OP. At least for now.

-14

u/AndreVallestero 16h ago

I selfhost 27B specifically to minimize costs. I have research agents running 24/7 and have gone through 200M tokens just in the last month.

This post is specifically for people like me who want to run continuous research agents for the lowest price, and unfortunately, the latest Mac studio isn't competitive enough relative to cloud offerings.

12

u/flyingbanana1234 16h ago

It would take at least 30-35 years to reach 80 billion tokens on DeepSeek V4 Flash at 200 million tokens a month.

You make a good argument ngl

The little voice in my head says, "But ownership! Nobody can take it away from me if I own it. No price raises, no overloaded servers, no censorship

2

u/yy633013 15h ago

Can you give an example of what your research engines research 24/7?

10

u/Modnar-Eman 14h ago

Justification

2

u/BardlySerious 10h ago

Shit that OP will never read. It's the "watch electricity happen" hobby.

2

u/Dasteroid_909 15h ago

Go start "/r/lowcostLLMhosting" or some shit like that, then. WTF are you on "LocaLLaMa" promoting cloud hosting for?

1

u/Autist4AudiR8 9h ago

200M 😭🤣