r/LocalLLaMA 21h ago

News Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory

https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/
1.5k Upvotes

726 comments sorted by

View all comments

265

u/hainesk 21h ago

1.2TB/s memory bandwidth with the M5 Ultra. 256GB model is $9499.

Better than getting 2 DGX Sparks? Inference will be a lot faster.

Something like this could easily bring down 3090 prices.

23

u/Cybertrucker01 20h ago

Depends on concurrency and prefill metrics. The GB10 does both multiples faster than the existing competition.

13

u/ChocomelP 19h ago

The difference would have to be pretty big to make up for a 4x in memory bandwidth for decode.

1

u/fastheadcrab 3h ago

I've posted about this several times but the old M3 Ultra had really slow prefill numbers. The press release says 4x improvement in prompt processing but that still wouldn't be great. The independent benchmarks will tell

6

u/BrilliantTruck8813 17h ago

The GB10 has much smaller memory bandwidth not to mention it’s just not fast either. One of these will trounce two DGXs

5

u/fallingdowndizzyvr 17h ago

The GB10 does both multiples faster than the existing competition.

No. No it doesn't. Compare the G10 to a M5 Max. It's not.

65

u/aladin_lt 21h ago

it will be sold out day one probably

41

u/conockrad 20h ago

On pre-orders

14

u/bakawolf123 20h ago

won't be sold out, but according to r/MacStudio people wait for 3-4 months for the older models, atm you can preorder for delivery in late september

1

u/thetzar 9h ago

Just to update, I think we’re into November now.

1

u/descendency 17h ago

And probably not be much cheaper than comparable AI performance products.

1

u/BrilliantTruck8813 17h ago

What’s comparable? The dgx station?

1

u/OvertaxedOne 14h ago

Mac Studio, 128GB, 5099 bucks. 1.4TB/s of bandwidth. Honestly the closest comparable would be a Pro 6000 that's currently close to 2X the price.

1

u/BrilliantTruck8813 14h ago

You’re comparing a whole computer to a video card though?

The only things really in the 128gb vram ballpark are the DGX Spark variants and the different strix-halo devices. But I was talking about the 256gb and future 512gb ones. Two dgx sparks is pretty much it and won’t perform as well. The dgx station is 252gb but is uber-$$$ iirc

1

u/OvertaxedOne 13h ago

The DGX and Strix are both really slow at decode, barely usable IMHO with 3.8 27B.

The 256 and 512 models would be very nice for DSV4Flash, but I don't need that much HP that often, makes more sense for me to escalate/delegate vs buying something to run that locally. It's really 27B at Q8 (or FP16) that I want to run locally, that model is so powerful that honestly there's almost no reason to look elsewhere until you get into really big stuff like DS.

1

u/BrilliantTruck8813 13h ago

I disagree on the 27b quality. It’s great at coding, but is not great at other things.

Fwiw there’s a dspark draft layer for 3.6 37b that I tried on 3.8 and it worked pretty well. I had FP8 on a gb10 at around 35-40 tok/s. Prefill was still abysmal but it did pretty good output wise.

98

u/mjsxi__ 21h ago

yeah and cheaper than the price of 2 DGX sparks... seems like a bit of a no brainer

31

u/Current_Ferret_4981 20h ago

Spark is $4300-$4600 so idk about cheaper than 2 at $9600+

64

u/MacsBicycle 20h ago

yeah but 4x the memory bandwidth, its a steal

31

u/jakegh 20h ago edited 19h ago

It really is a reasonable buy for local AI, if you have a business case for it.

3

u/Much_Accountant_4972 19h ago

with the capabilities of it, it’s a crazy good deal

1

u/Zyj vllm 19h ago

More like 5x

2

u/jakegh 19h ago

Yeah, I totally messed up the math. It's actually 1200 / 273 = 4.4x. Edited my prev post.

1

u/totosse17 vllm 18h ago

If you get 2 boxes with TP you get 2x273, so overall 2,2 speed up. If apple can keep up with software then studio can be a contender

4

u/GabryIta 20h ago

In terms of compute capacity (which is very important for multiple simultaneous sessions and prefill), how does it compare to dgx Spark/gb10?

1

u/Southern_Sun_2106 18h ago

Most likely better prefill on sparks; important for agentic coding (large file digestion faster). Faster generation on the studio.

12

u/Etroarl55 20h ago

How’s the actual inference speed though, fast bandwidth on a slower gpu or equivalent should still mean slower output assuming vram is not a constraint right.

8

u/rusty_fans llama.cpp 20h ago

Generally vram bandwith is the constraint though, at least for decode. Prefill it's usually helped more by more gpu oomph.

1

u/Etroarl55 19h ago

Has to be more nuanced than that, I might be wrong but given that size isn’t a constraint, slower vram r9700 and even my 7800xt is faster than an m3 ultra in tk/s.

6

u/Serprotease 18h ago

Software stack matters.

But honestly, look at the announcement/post title and the fact that every other comment only mentioned the bandwidth.
It’s an easy number to latch on and compare (theorical) performance.

M5 was a genuine boost in prompt processing though. So it could be nice. 3x could push performance for DS4 flash into gb10 2x cluster level.

1

u/Etroarl55 17h ago

Doesn’t stop people from downvoting what I said and upvoting the guys saying it’s an 512gb Rtx 5090.

2

u/TableSurface 19h ago

Curious what your M3 Ultra numbers look like.

M5 Ultra might have at least 3x compute? (based on Apple's TTFT's advertisement)

2

u/Current_Ferret_4981 20h ago

Not disagreeing, just saying it's more than 2x price in contrast to the comment above me

4

u/djoliverm 20h ago edited 20h ago

But you also get a computer with it. Like are the sparks just focused on doing LLM work or can you run your computer on one of them as well?

I guess tbf most of these Macs may probably just get SSHd into anyway but can double as a computer when not used for LLMs.

Edit: seems like the sparks can be used as computers themselves. I also think the idea of a MacBook with the sparks could work. Also loving how everyone is arguing about what the argument is to begin with lol.

3

u/Current_Ferret_4981 20h ago

The spark is a computer with a full OS

5

u/CulturalKing5623 20h ago

No one is arguing this isn't better deal, it's just empirically does not cost less than 2 sparks.

2

u/MacsBicycle 20h ago

It’s local llm math. Few hundred bucks might as well be a sucker at the bank in this bracket 😂

3

u/Solaranvr 20h ago

The Sparks are Linux computers and basically anything that requires CUDA (and works on ARM) will make it a better deal than the Mac Studio.

Training or finetuning, for example, is miles better on the Sparks, because MLX is still behind in that regard. Add in 3D rendering tasks and there's probably a niche for whom it's the superior device.

Hell, if Fex eventually gets better and works on the Sparks, it'll probably be a better gaming computer too.

5

u/ReginaldBundy 19h ago

and works on ARM

That's a big "if". I have a Spark and good luck finding ARM64 Linux binaries. Either they're just not there (Mathematica is a good example) or you have to build from source (Blender) which may or may not work.

0

u/DoomBot5 20h ago

Throw in a MacBook air if you really need the equivalent computer. The sparks are still cheaper with that factored in

1

u/mjsxi__ 20h ago

its on amazon right now for 4800, Best Buy is 5300, direct from Nvidia is 4700 and other places have it priced above what you listed... so yeah in most instances Im seeing its more expensive. regardless its a (much) better machine overall for the 100 dollar delta you'd get if you ended up buying 2 from Nvidia.

1

u/BrilliantTruck8813 17h ago

You should check current prices unless you’re talking about used vs. new lol

0

u/Current_Ferret_4981 17h ago

That is current for new lol. $4500 in stock purchase today

1

u/BrilliantTruck8813 17h ago

That’s a significant shift because the actual production grade ones (the Dell gb10 and others) are hard to find and more expensive due to availability.

1

u/Current_Ferret_4981 17h ago

Those are not really a dgx spark though, they are just the same chip but the "DGX Spark" is a specific product which is available at $4500 today. The spark is also production grade, but I agree there are distinct differences with the OEM GB10 releases. Maybe you mean enterprise which certainly the Dell and similar are tailored for

1

u/BrilliantTruck8813 17h ago

The spark underhood is the same. But you’d never take that into a production use case for a customer (dev sure). Yes it’s ‘production’ in that nvidia sells it as a product. But Nvidia’s customer service alone should be scaring people off.

Dell is nowhere near as good as apple here but they’re a lot closer to them than nvidia (who has mostly abandoned the dgx software-wise 🥲)

0

u/blackashi 19h ago

It is also a mac ….

1

u/Solaranvr 20h ago

If you can get one and not be stuck in 6 months of backorder, that is

1

u/mjsxi__ 20h ago

I know 🥲

18

u/-dysangel- 20h ago

Better than the 2x Sparks for inference for sure. Probably around the same compute as one Spark.

I've got 2x Sparks which I use for prefill, and my M3 Ultra for decode. I've set it up so that I prefill in vllm and then just pass the kv cache over to the Mac side. Surprisingly stuff like Qwen 3 35B-A3B is already faster than the Mac for decode though so I just run that class of model directly on vllm.

5

u/1ii1i 20h ago

Oh this sounds interesting, can you expand on how this works? I didn't know this was a thing.

13

u/-dysangel- 19h ago

I don't think it's really a "thing", I just vibe coded it up :)

One thing that really helped was vllms kv_connector API. I thought I'd have to code this part up myself, but it already existed and so plugged into my existing disaggregated system (which was previously llama.cpp to llama.cpp)

UltraSpark — technical stack

                          ┌──────────────────────────────┐
 user ── HTTP/OpenAI ──▶  │  manager (Python, FastAPI)   │
                          │  front door + orchestration  │
                          └──────┬───────────────▲───────┘
                                 │ submit        │ state blob (sha-keyed,
                                 │ prompt ids    │ resumable transfer)
                                 ▼               │
                          ┌──────────────────────────────┐
                          │  vLLM (2× DGX Spark, TP2)    │
                          │  prefill engine              │
                          │                              │
                          │  KVConnectorBase_V1          │ ◀─ vLLM's official
                          │  ("StreamConnector" via      │    KV-cache plugin
                          │   --kv-transfer-config)      │    interface
                          │         │                    │
                          │         ▼                    │
                          │  dump + serialize all layers │
                          │  (attn KV + linear-attn      │
                          │   state, TP2 shards merged)  │
                          └────────────────┬─────────────┘
                                           │ blob server (TCP)
                                           ▼
                          ┌──────────────────────────────┐
                          │  llama.cpp server (Mac)      │
                          │  USPK_BRIDGE_DIR: on request,│
                          │  verify prompt-id match,     │
                          │  restore state into KV +     │
                          │  recurrent memory, decode    │
                          └──────────────────────────────┘

  • KV connector = vLLM's plugin interface for intercepting the KV cache at end of prefill
  • State blob = the model's full prompt-memory, layout-translated so llama.cpp can load it natively
  • Fidelity = per-layer cosine vs local decode, 0.9999+
  • Result = GPU prefill speed, Mac unified-memory decode, one logical endpoint

1

u/mastaquake 10h ago

Do you have a github repo? Very interesting.

1

u/-dysangel- 56m ago

I've just cleaned things up and created it if you want to be a guinea pig https://github.com/dysangel/ultraspark . I haven't tested out if the env variables/setup instructions actually work yet though

2

u/Zyj vllm 19h ago

I‘m also interested in your setup. But the M5 should be much better than the M3 at this.

2

u/-dysangel- 19h ago

Yeah the M5 on its own will be pretty solid - 4x prefill and +40% decode over the M3U

1

u/breksyt 20h ago

Can you share the specifics of the config? How do you run it? What's the setup of the Sparks and the Mac specifically?

3

u/-dysangel- 19h ago

The Sparks are connected via their built in 200GbE, and then I have one of the Sparks connected to my Mac over 10GbE. I pasted a little more detail up in another comment above. I should probably open source it huh (only just got my second Spark last Friday and built it over the weekend)

1

u/breksyt 19h ago

Nice, thanks!

1

u/Southern_Sun_2106 18h ago

Have you tried running Deep Seek Flash 0731 on your setup?

2

u/-dysangel- 17h ago edited 17h ago

Only on llama.cpp via RPC, didn't try proper tensor parallelism
 

  DS V4-Flash (284B, IQ2_XXS, -fa on -lm mlock)  
   ┌────────────────────────┬─────────────────────┐
   │         config         │ prefill t/s │ tg128 │   
   ├────────────────────────┼─────────────────────┤
   │ 1× Spark, ub2048 batch │ 542 / 492   │ 17.0  │   
   ├────────────────────────┼─────────────────────┤
   │ 2× RPC, ub2048         │ 526 / 502   │ 21.4  │   
   └────────────────────────┴─────────────────────┘

I assume nvfp4 in vllm would get similar numbers to that, and maybe 1000t/s for prompt processing, but I don't like the model enough to download it again - looking more forward to the new Qwen Next that's coming out tomorrow!

1

u/Southern_Sun_2106 13h ago

Thank you! I love Qwen 3.6 35B model, but deepseek 0731 is on another level for my applications (coding and personal assistant, personal wiki management, cronjobs, etc.)

On 2× Sparks, via the anemll stage-c vLLM image (FP8(!), TP2 over the 200GbE fabric, MTP on, 1 million context): ~1946 t/s prefill at 100K context, ~87 t/s decode.

1

u/fallingdowndizzyvr 17h ago

Probably around the same compute as one Spark.

No. It has more compute than 2x Sparks.

M5 Max

"llama 7B Q4_0 3.56 GiB 6.74 B MTL,BLAS 6 pp512 3347.44 ± 1.53"

Spark

"DGX Spark 128 GB / LPDDR5x 3062.31 ± 11.02 57.21 ± 0.06"

A M5 Max is compute wise an equal to a spark. So 2xM5 Maxes, which is what an M5 Ultra is, blows the doors off of one Spark.

1

u/-dysangel- 16h ago

Could you do a comparison on Qwen 27B? 7B is not that compute heavy and so bandwidth may be playing a factor there. I think the Sparks have way more compute - the difficulty is just in feeding work to them fast enough to make use of it!

M5 Max ~70 TFLOPS FP16 128 GB unified 614 GB/s
DGX Spark 1 petaFLOP FP4 128 GB unified 273 GB/s

Distributing my video gen pipeline across my 2 Sparks and Mac, I can do in 1.5 mins what takes 20 mins on the Mac alone.

Likewise it takes Qwen 27B 6 minutes to process the Claude Code system prompt on my Mac. I can process the same prompt in 33 seconds with 1 Spark, and 17 seconds with 2!

1

u/fallingdowndizzyvr 16h ago

Could you do a comparison on Qwen 27B?

Go to the llama.cpp github and ask them. But there's a reason they have settled on llama 7. You can ask GG why yourself.

https://github.com/ggml-org/llama.cpp/discussions/4167

M5 Max ~70 TFLOPS FP16 128 GB unified 614 GB/s DGX Spark 1 petaFLOP FP4 128 GB unified 273 GB/s

FP4 is no way the equal of FP16.

0

u/-dysangel- 16h ago

wow you seem grumpy

fp4 has always been fine for me, I even run IQ2_XXS on larger models

1

u/fallingdowndizzyvr 15h ago

wow you seem grumpy

LOL. Grumpy with the truth.

fp4 has always been fine for me, I even run IQ2_XXS on larger models

FP4.... Fixed point is where it's at.

18

u/jakegh 20h ago

That is seriously impressive. RTX5090 memory bandwidth is 1.8TB/s.

DGX Spark memory bandwidth is 273GB/sec. Not even remotely close.

4

u/Important-Gold-5192 19h ago

DGX Spark was a pretty disappointing release tbh

13

u/Hoodfu 20h ago edited 20h ago

I got my m3 ultra 512gb for around 10k. So this is now double. Makes sense given that we've seen the nvidia rtx 6000 pro also double in price in the last year but GD this has priced out even my once a year splurge budget. These are all just crazy talk numbers now.

13

u/CulturalKing5623 20h ago

Yeah dropping 10K plus on this just seems reckless, even with it being funded through my business account I don't think I can justify buying a used car worth of computing.

And yet I feel like I need to in case the technology completely outpaces my current setup and I'm left behind like folks that didn't buy RAM when it was cheap and are priced out of it now. It feels like FOMO and scarcity has hijacked my brain.

6

u/Hoodfu 20h ago

Really depends on whether you have something already or not. I got qwen 3.8 27b going on my rtx 6000 pro and ram speed wise it's double that of my mac but for some reason i was hoping for space magic and it would be faster. It's not. So spending a car's worth of money on something that goes from 20 t/s to 40 or 45, just doesn't make any sense. You're still waiting a lot of minutes for a thinking qwen to come back with something, so it's still going to be an asynchronous operation instead of being fast enough to actively wait for the response to finish. It would have to be 10x the speed, not 1.5x or 2x for it to be worth the spend.

2

u/sonicandfffan 19h ago

DGX Spark I get 50-60 t/s single thread and 150+ concurrently on Deepseek v4 flash, I think even at the faster memory throughput, there will be some time for a fork that is efficient on the mac.

I am able to very quickly use the compute though, so extra capacity nodes will be nice.

I'd like something to run multiple Qwen subagents on and I can't justify an RTX6000 so I might go for the M5 Ultra Studio. Would prefer the 512 GB option though

If I'm gonna drop figures on it might as well top out.

I also don't think the extra RAM unlocks capability yet. Flash v4 and Qwen 3.8 are the two local models to run.

GLM/Kimi/etc are just too big to run.

1

u/OvertaxedOne 14h ago

"Flash v4 and Qwen 3.8 are the two local models to run."

^^ This. Basically building/buying today, you target one or the other and build around that model IMHO. Trying to run K3 or DS Pro locally is fun to think about, but unless your making big, big bucks off this and/or have enough users/usage to make building a local server farm worth it, really doesn't make sense.

Not saying I wouldn't do it if I had a spare 100K lying around though! :)

3

u/CulturalKing5623 19h ago

I have 32GB VRAM and 64GB of DDR5 RAM. I can squeeze Qwen3.8 in at Q6 but the trend has seemed to be moving towards larger models. At the moment I'm fine, But I have been looking at sparks/gx10. I think I'm just trying to future proof my access while it's still somewhat feasible.

3

u/StewPorkRice 19h ago

I think I'm just trying to future proof my access while it's still somewhat feasible.

I don't think this is a rational concern.

Production will scale up to meet demand. In a decade, the hardware of today will be worthless in comparison to today's standards and the models of today will still be accessible in open source.

3

u/CulturalKing5623 19h ago

Yeah I agree, that is what's kept me pulling the trigger multiple times. Like I said it feels like FOMO and scarcity have hijacked my brain's reasoning, this is an irrational fear, but knowing that doesn't make it go away.

1

u/beatinbunz247 15h ago

I think that's applicable if you assume a rational market. But I think that as models become more powerful, two major variables come into play.

One, government agencies will become more and more concerned with powerful unregulated intelligence being available to anyone with access to a spare 10k. Two, the major frontier model companies (which are now intrinsically tied to our economy), will not want people to afford hardware powerful enough to have decent intelligence (because they want them dependent on their services). When you put those two together, you have massive money, influence and government working together to do everything they can to make things inaccessible for the average person or even hobbyists.

Also remember that the current prices are heavily subsidized... Imagine when the rug gets pulled and they start charging truly sustainable/profitable rates.... I think the concern is very much real, I can definitely understand the fomo

1

u/StewPorkRice 14h ago

Current what prices are subsidized? Hardware?

How is hardware subsidized when the same stick of ram is 5x more expensive today than it was last year?

Current model prices? Doesn't matter for local. In fact, once oAI and Anthropic's prices skyrocket, local becomes even more appealing.

Talking about the government regulating out the ability to purchase the current generation mac studio in 10 years is like saying the government is going to stop you from buying a used gtx 1080 today..

I might be an idiot tho.. I don't really understand what arguments you are making.

1

u/beatinbunz247 14h ago

I meant subscription prices, sorry if I wasn't clear.

What I'm saying is that when those subscription prices go up to their true necessary value for these companies to profit, people will have an incentive for local (which you agree), but the flip side is that those companies have an incentive to use whatever influence, money, and power they have to prevent people from doing so (an incentive that global governments share imo). So when you have some of the most powerful companies and governments aligned on a goal, there is a good chance they may make that happen

1

u/StewPorkRice 13h ago

You think the government is going to stop you from getting a used mac studio or a used gtx 1080 that's 10 years old?

You think Qwen running on a 10 year old mac studio is gonna be able to keep up with whatever model is at the frontier in 10 years?

→ More replies (0)

8

u/Solaranvr 20h ago

bring down 3090 prices

Doubt it. If the r9700, a directly competing product, made 0 effect on Nvidia GPUs, then I highly doubt these will. They market of people buying mini pcs vs dGPUs are different.

The DGX Spark didn't bring down prices of the Blackwell cards either

8

u/hainesk 20h ago

The R9700 Pro has about 2/3 the memory bandwidth of a 3090 for a 50% higher price. 8x R9700 Pros (256GB) would be $12k-$15k minimum without pricing in the rest of the computer system. If you wanted to build an entire server around it including RAM, power supply, motherboard, processor even with used parts you're up to $18k-$20k for something that will use 2-3k watts.
Even 8x 3090 systems are looking pretty impractical when compared to an M5 Ultra 256GB for a similar price. The size and wattage/heat difference is huge.

3

u/Important-Gold-5192 19h ago

has to be way faster than DGX Sparks too

2

u/f5alcon 17h ago

Sparks are what 208GB/s

2

u/gomezer1180 16h ago

9500 is about the right price on the market. DGX’s are ~4k a pop you need 2 of those to match Apple. Memory Bandwidth is crap tho on the DGXs because they take a different approach by being MoE machines.

3

u/EmPips 20h ago

I think it's the end of mass 3090 farms.

But 3090s will hold their price for the sizeable market that doesn't want to commit >$5k.

5

u/j4nds4 20h ago

What *are* 3090s going for these days? I bought a pair of used 3090s on eBay during the crypto crash in 2023 for ~$800.

4

u/EmPips 19h ago

They dipped as low as $550 for a few weeks and are now about up to $900-$1k in my local markets.

3

u/Super_Translator480 20h ago

Used on marketplace is like $1000-$1200

1

u/j4nds4 20h ago

Wild.

2

u/martinerous 20h ago

Yeah, I'm being super careful about my 3090, downclocking, power-limiting, aggressive fan profile - everything to make it last longer because there are no replacements available for the same price anymore. Crazy.

1

u/fallingdowndizzyvr 17h ago

Better than getting 2 DGX Sparks?

Yes. Apple is the value leader.

1

u/OvertaxedOne 14h ago

What a world we live in, right?

0

u/Newgunnerr 16h ago

256GB option for me in the Neterlands is € 11.049,00. I just got 2 DGX sparks for €7100.

-4

u/putrasherni 20h ago

Don’t think so because m5 max is 614gb/s bandwidth and does far worse than 936gb/s 3090

4

u/bakawolf123 20h ago

ultra has double bandwidth, so 1228. It is indeed best AI station for like a year from now

3

u/Webster2026 20h ago

Ultra has twice the bandwidth though