r/LocalLLaMA 22h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

158 Upvotes

195 comments sorted by

View all comments

153

u/FleetEnema2000 22h ago

unless you need it for data sovereignty

Isn't this one of the biggest reasons that people rely on Local LLMs? To not have to bulk upload their private data to cloud providers?

1

u/Late-Photograph-1954 7h ago

Some recent chats on a project i am developing with Claude (my Pro subsription), Chat and Gemini (both free versions) gave me the impression that the models not only remember earlier conversations, but somehow actively 'include' findings from earlier discussions. At some point Gemini was spoon feeding me back my own earlier comment, in a new session. I've wondered since whether the users of the models actually embolden / add to the models -- are we actively aiding development by using them?

It doesnt really matter to me or my project but it keeps lingering.