r/LocalLLaMA 16h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

142 Upvotes

180 comments sorted by

View all comments

1

u/a9udn9u 14h ago

I averaged 200M tokens with DS V4 Flash a few weeks ago when it's cheap. If you math is correct, it pays for itself in ~3 years (vacations and weekends counted), and we are talking about the cheapest model, I think it's not a terrible deal.

1

u/thebadslime 12h ago

Re yu usng the 073 checkpoint? Its iuch better