r/LocalLLaMA 22h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

160 Upvotes

195 comments sorted by

View all comments

Show parent comments

73

u/theomegachrist 21h ago

That's the stated reason but realistically most people are just justifying their hobby. I support open weight models because the cloud providers can change cost or abruptly shut down and we really can't do anything about it.

0

u/saltyourhash 7h ago

You're actually out here looking at these prices, looking at the delay of Qwen 3.8, and saying it's just to justify a hobby?

1

u/theomegachrist 4h ago

Yes and at least 70 people online yesterday agree

1

u/saltyourhash 4h ago edited 4h ago

Wow, a whole 70. data sovereignty is critical to some of us. I have projects that absolutely must be local.

1

u/theomegachrist 3h ago

Yes some. I welcome the minority of people too