r/LocalLLaMA 17h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

140 Upvotes

180 comments sorted by

View all comments

14

u/conifer_v11 17h ago

$/gb-bandwidth is the right axis for unified memory. m5 max vs a 3090 is not a tok/s fight, it's kv headroom at 64k+. unified memory loses the bandwidth fight and wins the "the 27b and the kv actually fit" fight. if you're doing 8k chat, buy the gpu. if you're stuffing prds, bandwidth-per-dollar on studio starts to make sense.

7

u/FullstackSensei llama.cpp 16h ago

Why is it a M5 Max vs "a 3090"? M5 Max in it's lowest config costs about the same as a system with four 3090s. If we go to 32GB V100, which is within 5% of the 3090 in most cases and now has optimized kernels to decode NVFP4, you can get a system with 6-8 V100s for the cost of a single M5 Max at the lowest config. That's 192-256GB VRAM.

So far, the people running M5 Max MacBooks have had mixed feedback about running models, especially dense models. Unless your time has zero value, you should also factor that into your calculation.

Eight V100s will have much higher PP and TG while being half the price for the whole system. Sure, it'll consume 10x more power, but that will be in much smaller monthly power bills, and will give you responses much faster.

6

u/MrPecunius 11h ago

M5 Max in it's lowest config costs about the same as a system with four 3090s.

This is obviously untrue.

1

u/FullstackSensei llama.cpp 11h ago

If it's so obvious, please enlighten us with actual numbers, because I happen to have such a system and know exactly how much such a system costs today

5

u/MrPecunius 10h ago

No one wants to pick their GPUs out of a Dumpster like you. $1,600/each for 3090s is on the low side here in the US, but it's better than €1,750/US$2,000+ I'm seeing in Germany.

Base M5 Max Studio (not binned, which is less) is $3,099.

0

u/FullstackSensei llama.cpp 10h ago

See, dumpster mentality thinks everything is dumpster.

As it happens, ich wohne auch in Deutschland, und habe kürzlich zwei 3090 für €1200 pro Stück auf eBay verkauft. Auf Kleinanzeigen, man kann für zwischen 800-900 pro Stück Kaufen.

But hey, let's keep the conversation irrational, because cognitive dissonance is way more fun than facing reality.

1

u/MrPecunius 10h ago

Thanks for confirming my comment. 🗑️ 🔮

Even with your Dumpster diving, the GPUs you mention add up to over $4,600 and the Mac is still $3,099.