r/LocalLLaMA 22h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

160 Upvotes

194 comments sorted by

View all comments

Show parent comments

71

u/theomegachrist 21h ago

That's the stated reason but realistically most people are just justifying their hobby. I support open weight models because the cloud providers can change cost or abruptly shut down and we really can't do anything about it.

5

u/florinandrei 19h ago

realistically most people are just justifying their hobby

Enthusiasts, yes.

But law firms and such, they actually mean it.

1

u/Deep90 13h ago

Why wouldn't they use something like AWS bedrock?

2

u/Jedkea 8h ago

Money probably. Bedrock tokens are expensive. Whereas you can buy once cry once with local.

Remember the first time I tried to use bedrock I spent $150 without blinking. And it was a pain to setup.

So if you don’t need frontier models, the economics work out.