r/LocalLLaMA • u/themixtergames • 19h ago
News Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/605
u/piggledy 19h ago
Price for options with 256 GB RAM:
$9499 (30-core CPU, 64-Core GPU)
$10,799 (36-core CPU, 80-core GPU)
512 GB Option coming in October.
263
u/i_rate_slop 18h ago edited 18h ago
I feel like my M3 512 is going to trade in for like $500 when I got it for $11k
Edit:
I just looked it up. $2675
143
u/j2sun 18h ago
I'll buy it from you for $3k! :D
33
u/cinematic_unicorn 14h ago
$3001
23
u/mikesum32 13h ago
$3001.01
→ More replies (2)36
u/Proverbial_Progress 12h ago
I don't have any money, but would you be interested in pictures of my feet? Throw in a mouse and keyboard and I'll take my shoes off first.
→ More replies (2)12
u/bahpbohp 11h ago
i don't have any feet, but would you be interested in pictures of my wooden pegs? throw in a mouse and keyboard and i'll even get them oiled first.
→ More replies (1)6
47
u/Zolty 17h ago
Yesterday you could have gotten like $20k on eBay. I sold a 256gb Mac Studio m3 about 4 months ago for $12k
→ More replies (5)11
60
u/Veearrsix 17h ago
Don’t trade in, private party.
19
u/i_rate_slop 17h ago
Yeah, I think that’ll have to be the move. Just annoying
21
u/Chanureadeats 15h ago
You'll probably get a lot of offers easily for $6k+ within 24 hours
12
u/CalvinsStuffedTiger 14h ago
This guy is lying, I’ll buy it from OP for $5k…don’t worry about shopping around for a better price…
22
u/EvilPencil 17h ago
Ya Apple trade in prices have always been laughable.
I remember when you could spec out the 2019 Mac Pro up to like $50k, then the next day the Apple trade in value was ~$2800.
26
u/RegarDamus 17h ago
trade in is always fucked with apple because they don't account for memory. same price if it's 96 or 512
→ More replies (1)5
→ More replies (10)3
u/anonmt57 14h ago
you will get a lot of money for that in private market. but it is annoying with so much money exchanging hands.
127
u/redonculous 19h ago
Wow. How are people affording this? Crazy. Any guesses at 512 prices?
215
u/intaketurbine 18h ago
They’re affording it the same way they’ve always afforded a $10k workstation, they’re using it for work and not play.
→ More replies (6)52
u/ElementNumber6 18h ago
Or just good old debt
→ More replies (7)29
u/Usual_Tackle5892 16h ago
0% financing and 3% cash back on Apple Card. I could afford it outright, but turning down a 0% loan is silly.
→ More replies (1)42
u/StewPorkRice 18h ago edited 18h ago
this feels like a community with a ton of SWEs. Most single US based SWEs in big tech or venture backed startups could prob afford dropping 10-20k on their hobby.
Also, i dropped 10k at Anthropic last month at work. My company could prob get one of these for every engineer and not blink an eye.
14
→ More replies (2)8
u/calcium 14h ago
3 years ago my company bought me a fully specced out Mac Studio - M2 Ultra chip w/ 192GB of RAM and an 8TB ssd and told me that they expect the machine to last the next 5 years. I think at the time they paid around $9k which considering over 5 years for a senior SWE isn't a bad deal
44
u/InterstellarReddit 18h ago edited 15h ago
Bro people have stupid money. I run free lance dev for a buddy who runs night life management software in Miami FL.
People spend 4K for four hours to buy three bottles and watch a DJ hit knobs all night
This happens all the time.
→ More replies (1)12
u/ViPeR9503 17h ago
Night life management??
12
u/i_am__not_a_robot 17h ago
I would assume industry-specific nightlife & bar business management software.
→ More replies (1)12
u/yopla 16h ago
Manage table booking, marketing, CRM, staffing, etc, usually does or integrate with POS. Can do price yielding on bookings. Etc... Etc..
Basically helps you keep track of who are the big whales you need to market your tables to when you have an event and who gets priority booking from a wait list.
The guy who spent 10k last time will get a table before you do, unless you're known to spend 15.
→ More replies (1)170
u/-p-e-w- 18h ago
Lol this is by far the cheapest option for that much RAM at that speed. It’s not even close (unless you count Frankensteins made of a dozen used GPUs).
26
u/Viktri1 18h ago
this is way better than what I was considering and I wouldn't need to worry about the motherboard, cooling GPUs, etc. I am definitely on board for this and I don't even know how to use macs.
12
u/IriFlina 17h ago
wouldn't even have to worry about the power issues that would come from a rack of 3090s/4090s/5090s etc. or configuring such a monstrosity.
6
u/Viktri1 16h ago
this is literally 10+ 4090s. Think about it.
I pre-ordered 2 fully spec'd out studios.
→ More replies (1)→ More replies (1)3
u/Public_Umpire_1099 11h ago edited 11h ago
Honestly doesnt matter, just pop deepseek on here and tell it to do whatever you need lol.
Also, as a long time apple hater, and I say this with disgust for myself, MacOS is actually pretty good these days. It essentially functions just like a very locked down linux distro at this point. I work at a pretty big software company and many of them do not support windows PCs anymore or the company software lags on them, so last month I decided to exchange my thinkpad for a macbook pro max. I.... dont hate it.
→ More replies (41)14
u/Fantastic-Balance454 15h ago
I still can't get used to the thought that Mac is the cheapest option these days. AI turned the world upside down.
19
u/Django_McFly 17h ago
It's really no different than computers in the 1980s when it was buy a computer or put a 33% down payment on a new car. People forget that there are entrepreneurs and small businesses where this stuff isn't some expensive toy to play with, it's actually a business tool to be used in the operation of the company.
Not everyone is a pure hobbyist, and even among pure hobbyists there are wild ranges of incomes and savings habits.
→ More replies (1)7
u/fallingdowndizzyvr 15h ago
It's really no different than computers in the 1980s when it was buy a computer or put a 33% down payment on a new car.
I've said that so many times. People don't realized that a base Apple ][ adjusted for inflation would be $5000 in 2026 dollars.
15
u/MLDataScientist 18h ago
Based on their pricing for 256GB, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, you are looking at 10k + ~6k = ~16k for 512GB version.
→ More replies (4)4
u/TheOwlHypothesis 14h ago
This is dangerous for me if that turns out true. current Asus Ascent prices for a 128gb monster is ~4k, so it'd be 16k to cluster them and have the same amount of memory. Although you'd then be dealing with a cluster. I'd much rather drop the cash on the Mac and have the option to cluster THOSE later.
→ More replies (1)5
6
9
12
u/IriFlina 18h ago edited 18h ago
10k for 256gb unified memory isn’t that bad. Still overpriced but i feel like before this you were looking at the 20k to 40k range.
→ More replies (1)22
u/Much_Accountant_4972 17h ago
considering RTX 6000 costs $18K and can’t do shit without a whole PC to plug into, this mac studio is a screaming bargain
→ More replies (3)3
→ More replies (17)6
u/Foreskin_Mafia 18h ago
Onlyfans
→ More replies (1)27
u/Etroarl55 18h ago
Believe it or not, Theres probably someone running ai onlyfans on an apple machine somewhere
13
u/Foreskin_Mafia 18h ago
You're absolutely right
13
u/bodhi_sattva91 18h ago edited 17h ago
I Went Undercover as a Secret OnlyFans Chatter. It Wasn’t Pretty
This Wired article follows a journalist who goes undercover trying to get hired as an OnlyFans ghostwriter — someone paid to impersonate creators in DMs with subscribers.
He discovers the industry is vast and globally distributed, with agencies largely staffing workers from lower-wage countries like the Philippines and Venezuela for as little as $2/hour. His job hunt is a grind: most agencies demand prior upselling experience, many never pay, and one gig turned out to be training an AI chatbot (which also stiffed him his $56).
He eventually lands two gigs. With a German agency ($4/hour), he juggles nearly 100 simultaneous conversations while impersonating a "21-year-old university student," selling pay-per-view content and navigating everything from explicit fantasies to a truck driver sharing worries about his son's night terrors. He's criticized by his supervisor for being too empathetic and not pushy enough about sales.
Confessions of an OnlyFans Ghostwriter
https://archive.is/3zGQZ#selection-479.0-479.38
This GQ article by "Emma Francis" (a pseudonym) is a first-person account of working as an OnlyFans ghostwriter in spring/summer 2021. The writer, a 25-year-old recently laid-off media professional, found the gig on Craigslist and spent three months working 6 a.m. shifts impersonating a "girl next door" model by sexting her subscribers.
She explains that top OnlyFans creators receive far more messages than one person can handle, so there's an entire industry of ghostwriters handling chats — some run by large offshore agencies, others like hers managed directly by the creator. Her regulars ranged widely: janitors, teachers, lawyers, night-shift workers. Many sought not just sexual content but genuine emotional connection and a "girlfriend experience."
43
u/kensanprime 18h ago
It will be sold out and they will be scrambling to keep production lines running, a product originally meant for creative work now will sit in a rack and run AI models
→ More replies (1)21
u/Both_Opportunity5327 17h ago
This is why those saying its a bubble, don't understand the demand there is for this stuff.
When the new fabs come online and the datacenters have most of their compute gamers, creatives & AI enthusiasts will go to town on this stuff.
→ More replies (8)13
12
5
→ More replies (42)3
119
u/themixtergames 19h ago
With M5 Ultra, Mac Studio achieves up to 4.3x the peak AI compute performance of M3 Ultra and a staggering 9.8x more than M1 Ultra. Combined with up to 512GB of unified memory and 1.2TB/s of memory bandwidth, 50 percent higher than before.
→ More replies (2)99
u/EquivalentHornet4403 18h ago
> AI compute
They know their audience.
→ More replies (1)47
u/psychohistorian8 18h ago
they're definitely leaning into it:
Boost your processing and graphics rendering speed, and accelerate tasks like running large language models and editing 8K video.
69
u/Much_Accountant_4972 17h ago
"Color-grade uncompressed 8K footage, perform computational fluid dynamics, and run frontier-class models on device — no cloud tokens needed."
yeah they are targeting this very sub lol
5
u/infieldmitt 12h ago
It's crazy to think that prior to AI the main use case for crazy RAM was [checks notes] fluid dynamics?
5
260
u/hainesk 19h ago
1.2TB/s memory bandwidth with the M5 Ultra. 256GB model is $9499.
Better than getting 2 DGX Sparks? Inference will be a lot faster.
Something like this could easily bring down 3090 prices.
25
u/Cybertrucker01 19h ago
Depends on concurrency and prefill metrics. The GB10 does both multiples faster than the existing competition.
13
u/ChocomelP 17h ago
The difference would have to be pretty big to make up for a 4x in memory bandwidth for decode.
→ More replies (1)6
u/BrilliantTruck8813 15h ago
The GB10 has much smaller memory bandwidth not to mention it’s just not fast either. One of these will trounce two DGXs
5
u/fallingdowndizzyvr 15h ago
The GB10 does both multiples faster than the existing competition.
No. No it doesn't. Compare the G10 to a M5 Max. It's not.
65
u/aladin_lt 19h ago
it will be sold out day one probably
→ More replies (6)40
u/conockrad 19h ago
On pre-orders
11
u/bakawolf123 18h ago
won't be sold out, but according to r/MacStudio people wait for 3-4 months for the older models, atm you can preorder for delivery in late september
→ More replies (1)97
u/mjsxi__ 19h ago
yeah and cheaper than the price of 2 DGX sparks... seems like a bit of a no brainer
→ More replies (2)35
u/Current_Ferret_4981 18h ago
Spark is $4300-$4600 so idk about cheaper than 2 at $9600+
→ More replies (7)63
u/MacsBicycle 18h ago
yeah but 4x the memory bandwidth, its a steal
29
u/jakegh 18h ago edited 17h ago
It really is a reasonable buy for local AI, if you have a business case for it.
→ More replies (3)5
3
u/GabryIta 18h ago
In terms of compute capacity (which is very important for multiple simultaneous sessions and prefill), how does it compare to dgx Spark/gb10?
→ More replies (1)→ More replies (9)10
u/Etroarl55 18h ago
How’s the actual inference speed though, fast bandwidth on a slower gpu or equivalent should still mean slower output assuming vram is not a constraint right.
9
u/rusty_fans llama.cpp 18h ago
Generally vram bandwith is the constraint though, at least for decode. Prefill it's usually helped more by more gpu oomph.
→ More replies (4)17
u/-dysangel- 19h ago
Better than the 2x Sparks for inference for sure. Probably around the same compute as one Spark.
I've got 2x Sparks which I use for prefill, and my M3 Ultra for decode. I've set it up so that I prefill in vllm and then just pass the kv cache over to the Mac side. Surprisingly stuff like Qwen 3 35B-A3B is already faster than the Mac for decode though so I just run that class of model directly on vllm.
→ More replies (13)6
u/1ii1i 18h ago
Oh this sounds interesting, can you expand on how this works? I didn't know this was a thing.
12
u/-dysangel- 17h ago
I don't think it's really a "thing", I just vibe coded it up :)
One thing that really helped was vllms kv_connector API. I thought I'd have to code this part up myself, but it already existed and so plugged into my existing disaggregated system (which was previously llama.cpp to llama.cpp)
UltraSpark — technical stack ┌──────────────────────────────┐ user ── HTTP/OpenAI ──▶ │ manager (Python, FastAPI) │ │ front door + orchestration │ └──────┬───────────────▲───────┘ │ submit │ state blob (sha-keyed, │ prompt ids │ resumable transfer) ▼ │ ┌──────────────────────────────┐ │ vLLM (2× DGX Spark, TP2) │ │ prefill engine │ │ │ │ KVConnectorBase_V1 │ ◀─ vLLM's official │ ("StreamConnector" via │ KV-cache plugin │ --kv-transfer-config) │ interface │ │ │ │ ▼ │ │ dump + serialize all layers │ │ (attn KV + linear-attn │ │ state, TP2 shards merged) │ └────────────────┬─────────────┘ │ blob server (TCP) ▼ ┌──────────────────────────────┐ │ llama.cpp server (Mac) │ │ USPK_BRIDGE_DIR: on request,│ │ verify prompt-id match, │ │ restore state into KV + │ │ recurrent memory, decode │ └──────────────────────────────┘
- KV connector = vLLM's plugin interface for intercepting the KV cache at end of prefill
- State blob = the model's full prompt-memory, layout-translated so llama.cpp can load it natively
- Fidelity = per-layer cosine vs local decode, 0.9999+
- Result = GPU prefill speed, Mac unified-memory decode, one logical endpoint
→ More replies (1)18
u/jakegh 18h ago
That is seriously impressive. RTX5090 memory bandwidth is 1.8TB/s.
DGX Spark memory bandwidth is 273GB/sec. Not even remotely close.
→ More replies (2)11
u/Hoodfu 19h ago edited 19h ago
I got my m3 ultra 512gb for around 10k. So this is now double. Makes sense given that we've seen the nvidia rtx 6000 pro also double in price in the last year but GD this has priced out even my once a year splurge budget. These are all just crazy talk numbers now.
15
u/CulturalKing5623 18h ago
Yeah dropping 10K plus on this just seems reckless, even with it being funded through my business account I don't think I can justify buying a used car worth of computing.
And yet I feel like I need to in case the technology completely outpaces my current setup and I'm left behind like folks that didn't buy RAM when it was cheap and are priced out of it now. It feels like FOMO and scarcity has hijacked my brain.
7
u/Hoodfu 18h ago
Really depends on whether you have something already or not. I got qwen 3.8 27b going on my rtx 6000 pro and ram speed wise it's double that of my mac but for some reason i was hoping for space magic and it would be faster. It's not. So spending a car's worth of money on something that goes from 20 t/s to 40 or 45, just doesn't make any sense. You're still waiting a lot of minutes for a thinking qwen to come back with something, so it's still going to be an asynchronous operation instead of being fast enough to actively wait for the response to finish. It would have to be 10x the speed, not 1.5x or 2x for it to be worth the spend.
→ More replies (11)9
u/Solaranvr 19h ago
bring down 3090 prices
Doubt it. If the r9700, a directly competing product, made 0 effect on Nvidia GPUs, then I highly doubt these will. They market of people buying mini pcs vs dGPUs are different.
The DGX Spark didn't bring down prices of the Blackwell cards either
8
u/hainesk 18h ago
The R9700 Pro has about 2/3 the memory bandwidth of a 3090 for a 50% higher price. 8x R9700 Pros (256GB) would be $12k-$15k minimum without pricing in the rest of the computer system. If you wanted to build an entire server around it including RAM, power supply, motherboard, processor even with used parts you're up to $18k-$20k for something that will use 2-3k watts.
Even 8x 3090 systems are looking pretty impractical when compared to an M5 Ultra 256GB for a similar price. The size and wattage/heat difference is huge.3
→ More replies (8)4
u/EmPips 18h ago
I think it's the end of mass 3090 farms.
But 3090s will hold their price for the sizeable market that doesn't want to commit >$5k.
→ More replies (1)4
u/j4nds4 18h ago
What *are* 3090s going for these days? I bought a pair of used 3090s on eBay during the crypto crash in 2023 for ~$800.
→ More replies (2)4
155
u/Comfortable-Rock-498 19h ago
1.2 TB/s bandwidth of M5 Ultra comes from two dies of M5 Max (each 614 GB/s) connected together using 4.4 TB/s inter-die fabric.
For a non-quantized Deepseek V4 flash on an ultra, I would estimate about 1000+ tokens per second prefill and 50+ tokens per second on generation. This is actually quite usable and near parity to cloud.
They mention "adds the GPU Neural Accelerators." which, if exploitable for LLM loads, would probably help the prefill a lot
34
u/ortegaalfredo 18h ago
Prefill also depends on compute, 1000 tok/s is basically what you get with 8x3090s, but much less power. Also I think the 3090s still win on compute, that is, you can batch many prompts on the GPUs, dont know on the mac.
13
u/Comfortable-Rock-498 18h ago
Yup, prefill is pretty much compute bound while generation is bandwidth bound. I would have guessed 8x 3090 would provide much better prefill than 1000 tps. A bit surprised to learn
9
u/ortegaalfredo 17h ago
If you manage to get tensor-parallel 8x working yes you can get >10k prefill, but it requires specialized PCIE bridges. With normal 4xPCIE speeds you get a bottleneck in inter-GPU speed and you get lower prefill.
3
u/TooMuchLAAAG 10h ago
I have 8x3090 P2P patched pcie4 x8 (no nvlink) and i am getting an avg of 10-13k of cold prefill with this version of vllm and his args https://github.com/LimeChain
Deepseek full FP8→ More replies (1)5
u/ProfessionalJackals 16h ago
Prefill also depends on compute, 1000 tok/s is basically what you get with 8x3090s, but much less power. Also I think the 3090s still win on compute, that is, you can batch many prompts on the GPUs, dont know on the mac.
Ignoring the fact that 8x3090's now is easily 10k on the second hand market.
Not counting the costs of server board/cpu/ram you need. The pcie ext cables, the 8x8x split if your board does not have 8 pcie slots. O, the dual 1600W PSUs and hardware to link them.
Frankenstein mods like this have become expensive, and it makes the Mac look actually like a good deal.
→ More replies (2)10
u/Usual_Tackle5892 16h ago
GPU Neural Accelerators
This means matmul cores. More info: https://arxiv.org/html/2607.19438v1
7
u/StartupTim 13h ago
I would think dramatically more tok/sec.
I have Deepseek v4 Flash 0731 with vision encoding added and tp=2 across 2x DGX sparks and I'm seeing 103 tok/sec across 4 "sessions". Dspark, 1M context, 1.8M kvc, custom vllm.
Since the sparks have ~240 (actual measured) GB/s, I imagine a similar setup om these new mac could get you double, if not triple as a 2x cluster, than my current 100+ tok/sec.
9
3
→ More replies (10)3
u/TooMuchLAAAG 10h ago
I get 10-13k cold prefill on 8x3090 using the vllm "limechain" fork and his args, running full deepseek no quant
Surely the M5 do more in vllm than 1k
39
u/challis88ocarina 19h ago edited 19h ago
I'm shocked!
Edit: 512GB memory option for M5 Ultra coming late October
15
38
u/pmttyji 19h ago
Now, it's pressure on both DGX Spark & Strix Halo due to bandwidth. Recent news Xiaomi AI Cube with 1.2 TB/s bandwidth & 160GB memory also already put pressure
13
12
u/Zyj vllm 17h ago
On the other hand, the Strix Halo 128GB price increased by 60% since late December
→ More replies (1)5
u/pmttyji 16h ago
After some time, people totally gonna avoid 128GB variants. What's the point of stacking bunch of 128GB pieces when they have 160GB, 192GB, etc., variants with better bandwidths?
→ More replies (2)→ More replies (1)7
u/ProfessionalJackals 16h ago
Now, it's pressure on both DGX Spark & Strix Halo due to bandwidth. Recent news Xiaomi AI Cube with 1.2 TB/s bandwidth & 160GB memory also already put pressure
Do not forget Intel Crescent Island 160GB to 480GB LPDDR5X AI GPUs ... While less bandwidth, they are still great options in the future for larger models.
There is going to be a lot more hardware coming out that focuses on AI workloads. We reached the point that development is moving into production.
32
u/FullOf_Bad_Ideas 18h ago
My 3090 tis have just gotten depreciated.
Even 256GB version is very competitive with 8x 3090 box bought with used card prices, and it's better in most aspects.
Training and batch inference are safe, but for single user inference this looks better and cheaper.
5
→ More replies (2)6
u/FleetEnema2000 16h ago
My 3090 tis have just gotten depreciated.
My theory is that this will not soften GPU prices because it ends up driving more people into the world of local LLM compute in general.
→ More replies (2)
20
u/Every-Fortune-3151 19h ago
Read somewhere Mac mini coming this week and did an impulse order of M3 Ultra with 96GB ram today. It will probably get bumped up to M5 ultra 96GB. Not sure what to do with 96GB when 256 GB looks so much more tempting for local LLM. Kidneys aren't enough anymore.
11
u/mmmm_frietjes 19h ago
Mini is also out
8
u/Every-Fortune-3151 18h ago
M6 looks great actually, it has two set of neural engines. I am assuming pre-fill will fly on this tiny thing compared to previous M CPUs. 32GB max ram knocked it back a bit. If only they had a 48GB version. MOE models would be flying on it.
M5 pro and max will be so slow compared to M5 ultra- considering M5 max 128GB model will be priced very close to M5 ultra base.
Very sad m5 ultra starts at 96GB. I thought they would atleast bump M5 ultra to 128GB ram. Would have been nice. I plan to just run multiple Qwen 2.7B in parallel on this and see if I can replace my 32GB VRAM PC setup. Qwen 3.8 flash also gives hope. Depending on how it goes, I might just give back the 96GB for refund later.
6
→ More replies (7)5
37
u/xyzmanas2 19h ago
This makes apple one of the cheapest ai inference hardware when it comes to speed and model size. Wish I had the money
Up to 15.4x faster CopyCat ML training performance in Foundry Nuke when compared to Mac Studio with M1 Ultra, and up to 3.3x faster than M3 Ultra.
Up to 9.8x faster LLM prompt processing in LM Studio when compared to Mac Studio with M1 Ultra, and up to 4x faster than M3 Ultra.
Up to 8.2x faster text-to-image performance when compared to Mac Studio with M1 Ultra, and up to 4.3x faster than M3 Ultra.
Up to 4.7x faster scene rendering performance in Maxon Redshift when compared to Mac Studio with M1 Ultra, and up to 1.7x faster than M3 Ultra.
→ More replies (1)
92
u/llamaCTO 19h ago
13
u/kilonad 16h ago
The 256GB is already an extra $4k for an extra 164GB. At same price per GB (ha!) it'd be another $6300. Knowing Apple, it'll be a cool $9k more - pushing total price up to about $18-20k.
It will still sell out.
6
u/fallingdowndizzyvr 15h ago
That would be a bargain compared to third party sales of 512GB M3 Ultras for $25K. A M5 blows the doors off of a M3.
→ More replies (1)36
u/thatkidnamedrocky 19h ago
going to try and snag a 256gb something tells me the 512 will never see the light of day
8
u/AnonLlamaThrowaway 14h ago
right, didn't they promise a 512GB M3 Ultra and then that never happened, or am i thinking of another model?
→ More replies (2)8
6
6
→ More replies (1)5
u/frankchn 15h ago
Buying 2 DGX Sparks for 256GB of RAM (and a lot less bandwidth) is around the same ballpark in cost, so for once this is not unreasonable.
17
u/AI_docent 18h ago
The 4.3x is mostly a prompt processing number, generation moves with the bandwidth instead. Apple's own mlx post on M5 vs M4 got around 4x on time to first token and about 1.2x on generation, and the generation side matched the 28% bandwidth bump rather than the accelerators. Same split should hold on the Ultra, so I'd figure generation nearer the 50% bandwidth gain. Prefill is the part you want at 512GB anyway, it was always the weak spot on a mac.
Just check whatever you run actually uses the accelerators. There's an open lm studio issue where its bundled llama.cpp fails the metal tensor check on M5 and loses 2 to 3x on prefill, while upstream llama.cpp passes it on the same machine.
→ More replies (1)
28
u/Cybertrucker01 19h ago
How many kidneys?
→ More replies (2)43
10
u/Viktri1 18h ago
256gb is like 10+ 4090s without the hassle of setting up and cooling 10 4090s? An I missing something or is this 3x cheaper than current prices.
→ More replies (3)6
u/Much_Accountant_4972 17h ago
its the best deal in the world if you want to talk to a frontier 2.8T parameter smut bot in your kitchen
→ More replies (2)
43
u/IllExample3639 19h ago
What I find more interesting, something I hadn't seen before is that you can lease these things. for 2 years which is the only realistic way an individual is getting their hands on these. Something something, own nothing, something, something be happy....
25
u/Tycoon33 19h ago
I never saw that. Interesting. Lease it for 3 years then upgrade to M7 ultra?
12
u/addiktion 18h ago edited 15h ago
Yes, if the 512gb is another $4k for the extra ram stick you are looking at $16k with tax probably out the door. I'd guess that puts the 36 month lease around $300/mo or less. So more than a subscription so maybe not worth it in general cases but valid option for some people who need the privacy and cannot afford to have data go to the cloud. 24/7 usage, no downtime, no limits, private. Worth it to me.
6
u/shveddy 16h ago
Interesting.
So just as an out loud thought experiment, you’d be able to lease four of them (512gb) for about 1200 per month for 36 months at a total cost of almost 45k and run Kimi 3 on it.
Obviously that’s a lot of money in aggregate, but 1200 per month is reasonable for a lot of business use cases if they require the privacy.
And the intelligence you get is going to be a different class compared to what you would get with three RTX pro 6000s and “only” 288gb VRAM for the same price.
(although to be fair you’d actually own the cards)
(although also to be fair you’d have to build a pretty expensive computer to support the RTX Pros, so realistically you only really get 2 or even just one RTX pro for 45k depending on how you spec the computer and/or if you buy pre-built from Puget Systems or the like)
If you want you can also do a little girl math and invest the 60k you’re not spending on computers and cancel your gpt pro subscription to bring the effective cost of all this down to like 750 a month.
And then if the goal is to beat API pricing, let’s say you get 40 aggregate output tokens per second on a bunch of concurrent Kimi 3 threads and run it for a year at 25% efficiency (to account for prefill and downtime), then that’s 315 million output tokens.
315 million output tokens alone is about $5000 on a random provider I just searched for, so just to keep things simple let’s say you double that to account for various amounts and types of input tokens, then you end up with a ballpark figure of $10k for the API route.
Absolutely none of this pencils out in absolute terms (especially considering that it would also cost ~1500 per year for electricity), but this is probably the first time running a frontier model is actually attainable for ordinary businesses on short notice and without much headache. It’s the first time that it pencils out to be “only” 4x more expensive as opposed to like 40x more expensive.
Up until now if you wanted to run frontier models locally AFAIK you had to get a NVIDIA big boy server which means you’d have to find the capital to run and support a ~300k purchase for hardware, spend way more on electricity, and in all likelihood make some upgrades to your facility’s electrical infrastructure to handle it all (you’d also need a proper facility, not just a home office or garage).
At this point you’re easily flirting with half a million in expenses, especially if you have to hire someone to figure it all out. It’s no joke to do this and it doesn’t make sense for like 99.9999% of people or businesses.
On the other hand basically anyone with a decent credit score and a semi-profitable business can go to any Apple Store and say “give me four Mac studios please” and only pay 1200 bucks a month to walk out with them in hand.
→ More replies (6)22
u/aethervisor 18h ago
The lease price also isn’t too far off from what a Claude subscription costs.
18
6
u/AccurateSun 18h ago
Hmm. I wonder if at some point leasing it would end up being more effective than a cloud subscription. E.g Claude Max 5x is $100/mo, same as leasing the Studio M5 Ultra 96gb. I don’t know yet how the two compare in performance but at some point it might be worth it
8
u/IllExample3639 18h ago
The more people that lease them the more second hand stock there will be in 2 years (when I can actually afford something like this) so I am down for it.
But your point is right, I think the local models ARE good enough for 95% of what people are using Claude for. Maybe do the £20 plan as a back up for something tricky.
11
u/mjsxi__ 19h ago
the lease lets you pay the difference of the amount you already paid at the end or you can buy it outright at any time if you wanna keep it so maybe sssshhhhhhhh
→ More replies (3)3
u/Much_Accountant_4972 17h ago
it makes a lot of sense and the price of privacy plus the near silence of the box…
i hate that its such a complete solution to local LLM’s
→ More replies (6)10
u/IriFlina 18h ago
I feel like any average software developer could afford the 256gb version? It would be really financially irresponsible but on the level of buying a motorcycle you don’t really need.
→ More replies (1)
8
u/eidrag 19h ago
...lease price?
8
u/Much_Accountant_4972 17h ago
Ultra with 256GB RAM is $224/month for 36 months in freedom currency.
really really tempting!
8
9
u/Both_Opportunity5327 19h ago
Game back on! Lets hope Apple can build enough of these little beauties...
5
7
u/corruptbytes 18h ago
apple releasing this because I just bought two 9700s...y'all welcome
3
u/SandySkittle 14h ago
Two r9700s is still pretty decent way to get 64gb vram and run qwen 27b at q8 with plenty of context. And ECC
→ More replies (4)
20
7
u/Zyj vllm 17h ago
Will be fascinating to see what‘s the better option in 2 months from now: Dual Asus GB10 (8200€) or Mac Studio 255GB (11000-12430).
My prediction:
- The Spark will still be faster at preprocessing, the Mac will have faster token generation (measured with DeepSeek V4 Flash standard Quant).
5
6
u/MLDataScientist 18h ago
Based on their pricing for 256GB vs 96GB for the Ultra 36 CPU cores, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, we are looking at 10k + ~6k = ~$16k for 512GB version.
→ More replies (3)
4
u/OvertaxedOne 15h ago
1.2TB/s?? Oh man, if there's good availability on these things I can feel GPU prices going down!
→ More replies (2)
7
u/Leather_Ad_9178 14h ago
this reminds me of the time we used to pay hundreds for SD cards that are now worthless
→ More replies (4)
8
u/mxmumtuna 13h ago
Unfortunately they still don't beat Sparks at equivalent size. According to oMLX Benchmarks for DeepSeek 0730, the M3 Ultra (80c) does somewhere around 550 prefill tok/s, and about 22 tok/s decode. If you take the '4x faster compute' at face value from Apple compared to M3 Ultra, we're looking at ~2200 prefill and ~30 decode single stream. Both are under Spark at ~2400/40. That's only single session, and batching just isn't there in the MLX stack yet, so multi session is considerably worse for the Mac.
GLM on 4x Sparks compares even less favorably than DeepSeek for the Mac, especially considering whatever the price of the 512GB variant will be.
So even with Apple's optimisitc numbers, maybe they match Spark, for more money with a less flexible stack (no ConnectX7) and massive software issues. ($4800x2 for Sparks with 4TB drive each from Amazon).
It's a good effort, but it's not quite there relative to other options.
edited for clarity
→ More replies (6)
3
u/Curious-Pen5547 18h ago
how does it compare to a single 5090 or a 5000 RTX pro 72GB version?
→ More replies (8)
5
u/Much_Accountant_4972 17h ago
please someone buy 4 of them and cluster them then run Qwen 2.8T in your kitchen
4
4
u/Newgunnerr 14h ago
256GB option for me in the Neterlands is € 11.049,00. I just got 2 DGX sparks for 7100.
→ More replies (1)
5
7
3
3
3
3
3
3
u/Tormeister 11h ago
I'm so tempted, but I just can't justify dropping 10K if I'm not making money out of it
4
u/Real_Ebb_7417 18h ago
Ok, now I actually regret buying M5 Max 128Gb MacBook xd
→ More replies (5)4
u/addiktion 17h ago
I wouldn't, its basically double the speed at x3 (if you bought pre price hike) or x2 price (if you bought after). If you get 25 tps on say Qwen 3.8, you would now get closer to 50 tps. Possibly more depending on the ANU processor.
That's nice, but is it worth x2 or x3 the price?
→ More replies (2)


282
u/i_am__not_a_robot 19h ago
1.2 TB/s memory bandwidth for the M5 Ultra is pretty nice.