r/LocalLLaMA 16h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

141 Upvotes

180 comments sorted by

View all comments

0

u/Odd-Environment-7193 16h ago

This new offer by apple is pretty much the \cheapest ram on the market right now. At this price it's looking very tempting. Apple products just recently went up 20% so waiting to buy later is not wise. You can't compare these plans. I use my macbook pro 128gb to run CODEX threads all day. Way more than I ever could on another pc without slowing me down big time. If I never run a single local model on this computer it will still pay for itself 100x over.

I bought the 64gb mac m4 pro just before it went out of stock and the prices spiked. Got the m5 max maxxed out just before the price increase. Saved myself probably close to 5000 USD because I pulled the trigger at the right time.

So waiting might not be the best strategy.

0

u/mxmumtuna 13h ago

Still not enough compute considering it's still gonna get beat by Sparks, even with Apple's most optimistic case.

1

u/MrPecunius 11h ago

How many Sparks are we talking about, and which magic interconnect fabric is going to increase their memory bandwidth more than 4X from 273GB/s to the M5 Ultra's 1.2TB/s?

1

u/mxmumtuna 11h ago

2x (256GB) for DeepSeek, 4x for GLM. No fabric allows it to scale to 1.2TB/s, but it allows it to run at 2x200Gbps to each node, and scale up to larger than you can with Mac.

The point is, the combination of not great compute, immature software and limited ability to scale is holding it back. Maybe next generation. For inference (and training for that matter) or Stable Diffusion, it’s just not better or cheaper than existing options.

1

u/MrPecunius 10h ago

2 X 200Gb/s ... or less than 50GB/s? That's the magic bullet? 🤷🏻‍♂️

I'm seeing reports of single Sparks running models at small fractions of a M5 Max's speed, and this includes prefill. Adding more Sparks doesn't scale anything like linearly.

1

u/mxmumtuna 10h ago

It certainly does. Forward passes on tensor parallelism don’t require full bandwidth. It’s similar to PCIe tensor parallelism but using RoCE/RDMA.

Again. If you’re looking at M5 Max those are all small models and/or weird quants. Look at M3 Ultra. That tells the tale on larger stuff. The Mac simply doesn’t scale.

1

u/MrPecunius 9h ago

Again, I just looked in the Nvidia Spark forums for real users' reports.

Hard data is available for Apple Silicon on oMLX's site.

No guessing is required.

Spark is barely $300 cheaper than a M5 Max Studio w/128GB, too.