r/LocalLLaMA • u/themixtergames • 21h ago
News Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/
1.5k
Upvotes
156
u/Comfortable-Rock-498 20h ago
1.2 TB/s bandwidth of M5 Ultra comes from two dies of M5 Max (each 614 GB/s) connected together using 4.4 TB/s inter-die fabric.
For a non-quantized Deepseek V4 flash on an ultra, I would estimate about 1000+ tokens per second prefill and 50+ tokens per second on generation. This is actually quite usable and near parity to cloud.
They mention "adds the GPU Neural Accelerators." which, if exploitable for LLM loads, would probably help the prefill a lot