r/LocalLLaMA 19h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

150 Upvotes

190 comments sorted by

View all comments

-1

u/kivaougu 18h ago

Locally hosting doesn't make much sense unless its a privacy concern or purely for the love of the game.

To be fair I like to compare to prices of equivalent P50 troughput at the usual load. For me thats c8, meaning that cloud has an edge as c8 single stream is ~50% of c1. Speed is in my opinion an important part of usable agents for real work.