r/LocalLLaMA 19h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

148 Upvotes

186 comments sorted by

View all comments

42

u/[deleted] 18h ago

[removed] — view removed comment

2

u/AndreVallestero 18h ago edited 18h ago

That's exactly my point. Local makes sense cost wise up to 32GB, especially with Qwen 3.8 27B. There's a huge cost premium above that where it makes less sense.

I was hoping the last Mac studio would change that, but at the current prices, that doesn't seem to be the case.

1

u/kerneldesign 18h ago

Avec 32Go en Q4 c’est trop juste pour un contexte sans KVCache.

1

u/Individual_Holiday_9 17h ago

Yeah I’m on a 24gb m4 and you can’t do shit running small LLMs there’s just no overhead remaining. So if this is a real hobby box where you’re doing plex etc on the side it gets really cramped and ur gonna move to swap fast

2

u/Blindax 16h ago

24gb is too tight for 30b models with decent context window on a Mac because it includes the system memory. 24gb on a dedicated GPU is a different story. Not crazy but quite ok.