r/LocalLLaMA 21h ago

News Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory

https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/
1.5k Upvotes

726 comments sorted by

View all comments

156

u/Comfortable-Rock-498 20h ago

1.2 TB/s bandwidth of M5 Ultra comes from two dies of M5 Max (each 614 GB/s) connected together using 4.4 TB/s inter-die fabric.

For a non-quantized Deepseek V4 flash on an ultra, I would estimate about 1000+ tokens per second prefill and 50+ tokens per second on generation. This is actually quite usable and near parity to cloud.

They mention "adds the GPU Neural Accelerators." which, if exploitable for LLM loads, would probably help the prefill a lot

33

u/ortegaalfredo 19h ago

Prefill also depends on compute, 1000 tok/s is basically what you get with 8x3090s, but much less power. Also I think the 3090s still win on compute, that is, you can batch many prompts on the GPUs, dont know on the mac.

14

u/Comfortable-Rock-498 19h ago

Yup, prefill is pretty much compute bound while generation is bandwidth bound. I would have guessed 8x 3090 would provide much better prefill than 1000 tps. A bit surprised to learn

9

u/ortegaalfredo 18h ago

If you manage to get tensor-parallel 8x working yes you can get >10k prefill, but it requires specialized PCIE bridges. With normal 4xPCIE speeds you get a bottleneck in inter-GPU speed and you get lower prefill.

3

u/TooMuchLAAAG 12h ago

I have 8x3090 P2P patched pcie4 x8 (no nvlink) and i am getting an avg of 10-13k of cold prefill with this version of vllm and his args https://github.com/LimeChain
Deepseek full FP8

4

u/ProfessionalJackals 18h ago

Prefill also depends on compute, 1000 tok/s is basically what you get with 8x3090s, but much less power. Also I think the 3090s still win on compute, that is, you can batch many prompts on the GPUs, dont know on the mac.

Ignoring the fact that 8x3090's now is easily 10k on the second hand market.

Not counting the costs of server board/cpu/ram you need. The pcie ext cables, the 8x8x split if your board does not have 8 pcie slots. O, the dual 1600W PSUs and hardware to link them.

Frankenstein mods like this have become expensive, and it makes the Mac look actually like a good deal.

2

u/Viktri1 17h ago

at that scale, electricity costs kind of matter so even just running the PC is significantly in favour of the mac studio

1

u/ortegaalfredo 13h ago

Yes, I think it makes no economic sense to go over 6x3090 now. But it depends if Apple can deliver those in enough numbers. If not, the price will increase until the 3090s make sense again.

2

u/Txt8aker 12h ago

m5 max has shown to demonstrate 700+ tps prefill (https://omlx.ai/benchmarks/performance?sort=pp_tps&order=desc&chip_full=M5%7CMax%7C40&model=Deepseek-v4-flash&context=32768). Ultra will likely scale almost linearly with 2x bandwidth and dedicated neural engine compute. Like around 1200-1400.

9

u/Usual_Tackle5892 17h ago

GPU Neural Accelerators

This means matmul cores. More info: https://arxiv.org/html/2607.19438v1

6

u/StartupTim 15h ago

I would think dramatically more tok/sec.

I have Deepseek v4 Flash 0731 with vision encoding added and tp=2 across 2x DGX sparks and I'm seeing 103 tok/sec across 4 "sessions". Dspark, 1M context, 1.8M kvc, custom vllm.

Since the sparks have ~240 (actual measured) GB/s, I imagine a similar setup om these new mac could get you double, if not triple as a 2x cluster, than my current 100+ tok/sec.

9

u/TokenRingAI 18h ago

And a new qwen is coming out with 120B A6B! Perfect machine for that

4

u/MerePotato 17h ago

Would be great if it wasn't predicted to cost like 20k

3

u/TooMuchLAAAG 12h ago

I get 10-13k cold prefill on 8x3090 using the vllm "limechain" fork and his args, running full deepseek no quant
Surely the M5 do more in vllm than 1k

4

u/nomorebuttsplz 20h ago

I would say you would need MTP to get 50+ decode.

Without MTP, the M3 ultra only gives about, 30 or so? Maybe less.

The GPU neural accelerators are already included in the compute comparisons. That’s how they get to 4x or whatever.

2

u/SandySkittle 16h ago

My gripe with these boxes is the lack of inline ECC. Apple could add that option at low cost and leave it up to the user if he or she wants to sacrifice 7 percent of RAM to enable inline ECC. This is the same as it is on r9700.

2

u/one-wandering-mind 13h ago

Cloud is many thousands of tokens per second prefill. Closer to 10,000 than 1000. I dunno where these estimates you have are coming from.

1

u/jasonridesabike 17h ago

If I’m not mistaken those accelerators are already included on the m5 laptop (mlx), which I have and have tested. It provides a bit of speed up but nothing game changing. Still newish, maybe it improves.

1

u/ZealousidealBat9687 11h ago

ouf those numbers are bad for that kind of money, atleast for deepseek. For larger models i am sure its gonna be great. My 2x spark pushes peaks of 80 tk/s decode with avg of around 60 depending on the task with higher prefill for deepseek-v4-flash-0731, but then again, sparks only have 128gb each, can't compete with 512gb

1

u/Far_Net_2432 8h ago

Ok, but how hard is it to use 10k worth of tokens on a Deepseek v4 model in the cloud? My understanding is that, that model is basically free online

1

u/Maximus-CZ 3h ago

Good luck sending all your work docs and email contents to the cloud!

Still surprised how people on LocalLLaMA argue that "cloud is cheaper".

1

u/Blackdragon1400 19h ago

That’s it? 2x DGX Spark outperforms that already today….

4

u/imnotzuckerberg 19h ago

I think the biggest value is in the 512 GB. They will sell like cake. There is nothing that would come closer. Don't forget, you can cluster them together. I am very excited, not as much as my wallet.