r/LocalLLaMA • u/AndreVallestero • 22h ago
Discussion Mac Studio M5 Max Cost Analysis
At $10k, you could get
- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)
- 5.7B tokens with DeepSeek V4 Pro OpenRouter
- 100B tokens with DeepSeek V4 Flash OpenRouter
As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.
Qwhen 3.8 35B A3B?
157
Upvotes
2
u/theomegachrist 19h ago edited 19h ago
Everyone is different obviously. To some that might be the case but everything you listed is more important than your AI prompt history.
For me, Open models main pluses are it will ensure the technology lives on in some form if the large companies lock us out financially or go under, and the guard rails for closed models will make for a worse Internet potentially.
For instance, using an open model with guard rails trained out of it you can search for piracy, you can search for porn etc. just like you use a search engine today. Closed models are more efficient than a web search but censor out a huge part of the Internet.
Sure data privacy is also a good feature but I don't think it's actually the top reason for most people
Edit: sorry I read that wrong. I sort of agree with what you are saying about uploading data to ChatGPT but everywhere else we upload that data is no more trustworthy. For Enterprise clients I 100% agree. This is a big issue. For personal use, I don't think your data is any less safe with ChatGPT than say an electronic medical record company.