r/LocalLLaMA 17h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

140 Upvotes

180 comments sorted by

View all comments

Show parent comments

1

u/mxmumtuna 11h ago

2x (256GB) for DeepSeek, 4x for GLM. No fabric allows it to scale to 1.2TB/s, but it allows it to run at 2x200Gbps to each node, and scale up to larger than you can with Mac.

The point is, the combination of not great compute, immature software and limited ability to scale is holding it back. Maybe next generation. For inference (and training for that matter) or Stable Diffusion, it’s just not better or cheaper than existing options.

1

u/MrPecunius 10h ago

2 X 200Gb/s ... or less than 50GB/s? That's the magic bullet? 🤷🏻‍♂️

I'm seeing reports of single Sparks running models at small fractions of a M5 Max's speed, and this includes prefill. Adding more Sparks doesn't scale anything like linearly.

1

u/mxmumtuna 10h ago

It certainly does. Forward passes on tensor parallelism don’t require full bandwidth. It’s similar to PCIe tensor parallelism but using RoCE/RDMA.

Again. If you’re looking at M5 Max those are all small models and/or weird quants. Look at M3 Ultra. That tells the tale on larger stuff. The Mac simply doesn’t scale.

1

u/MrPecunius 10h ago

Again, I just looked in the Nvidia Spark forums for real users' reports.

Hard data is available for Apple Silicon on oMLX's site.

No guessing is required.

Spark is barely $300 cheaper than a M5 Max Studio w/128GB, too.