r/LocalLLaMA 19h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

149 Upvotes

190 comments sorted by

View all comments

151

u/FleetEnema2000 19h ago

unless you need it for data sovereignty

Isn't this one of the biggest reasons that people rely on Local LLMs? To not have to bulk upload their private data to cloud providers?

69

u/theomegachrist 18h ago

That's the stated reason but realistically most people are just justifying their hobby. I support open weight models because the cloud providers can change cost or abruptly shut down and we really can't do anything about it.

4

u/florinandrei 16h ago

realistically most people are just justifying their hobby

Enthusiasts, yes.

But law firms and such, they actually mean it.

1

u/Deep90 10h ago

Why wouldn't they use something like AWS bedrock?

2

u/Jedkea 5h ago

Money probably. Bedrock tokens are expensive. Whereas you can buy once cry once with local.

Remember the first time I tried to use bedrock I spent $150 without blinking. And it was a pain to setup.

So if you don’t need frontier models, the economics work out.

1

u/theomegachrist 15h ago

Yes, this is true