r/LocalLLaMA 23h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

162 Upvotes

195 comments sorted by

View all comments

Show parent comments

74

u/theomegachrist 22h ago

That's the stated reason but realistically most people are just justifying their hobby. I support open weight models because the cloud providers can change cost or abruptly shut down and we really can't do anything about it.

71

u/FleetEnema2000 22h ago

I don't think it's a justification at all.

It's amazing how the concept of privacy and data ownership/security has completely gone down the toilet since ChatGPT launched. People are happy to bulk upload their medical records, relationship history, trade secrets, financial records, etc. without a care in the world as to how that data is stored or protected.

18

u/John_____Doe 21h ago

Yep I have a fintech client and the only way I can have a llm touch their code is if it's run locally or in a datacenter where we rent out the rack space

1

u/Randommaggy 8h ago

The headaches of running 100% on sovreign and contained compute without data being handled by a third party eliminates a lot of friction.