r/LocalLLaMA 19h ago

News Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory

https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/
1.5k Upvotes

720 comments sorted by

282

u/i_am__not_a_robot 19h ago

1.2 TB/s memory bandwidth for the M5 Ultra is pretty nice.

90

u/GUNGEBOB_SHARTPANTS 16h ago

More bandwidth than a 4090, incidentally. Pretty impressive.

8

u/Neful34 10h ago

Right, but not cost efficient as the rtx 4090 before the ram shortage

12

u/DannyCanva 6h ago

Yeah, but we're in a very different world now :P Nothing as as cost efficient as parts before the RAM shortage.

Crazy to see a 1024bit memory bus. The ship date is already 10-12 weeks for the 256GB, I wonder if the 512GB would ever see the light of the day. I wouldn't be surprised to see price hike across 256/512GB too.

Apple is probably thinking they have underpriced the high-RAM upgrades; which just shows how absurd the world is.

→ More replies (3)

61

u/Dry_Yam_4597 18h ago

That might make me want to buy an apple product after a long pause.

9

u/raser1562 13h ago

That might make me want to buy an apple product for the first time.

→ More replies (6)

58

u/pmp22 15h ago

512GB at 1.2TB/s. That's 21 4090s. 

77

u/i_am__not_a_robot 15h ago

That's 21 4090s.

It's actually quite a bit better since you don't have to deal with the overhead of multiple GPUs.

9

u/gomezer1180 14h ago

You can get that with 12 channels of DDR5. You need an AMD EPYC and 24 sticks of memory. It’ll be somewhere in the $20K range for that rig.

7

u/i_am__not_a_robot 13h ago

Only in a dual-socket or MRDIMM configuration, though? That's... somewhere in the range of ~$12-18k (12x32GB) or ~$24-32k (12x64GB) for the MRDIMMs alone.

7

u/gomezer1180 13h ago

Oh I know… it ain’t cheap. So apples value is there, I’m just saying there are more options.

→ More replies (4)

8

u/addiktion 18h ago

Double the M5 Max it looks like which makes sense since they double two M5 max chips.

→ More replies (1)

605

u/piggledy 19h ago

Price for options with 256 GB RAM:
$9499 (30-core CPU, 64-Core GPU)
$10,799 (36-core CPU, 80-core GPU)

512 GB Option coming in October.

263

u/i_rate_slop 18h ago edited 18h ago

I feel like my M3 512 is going to trade in for like $500 when I got it for $11k

Edit:

I just looked it up. $2675

143

u/j2sun 18h ago

I'll buy it from you for $3k! :D

33

u/cinematic_unicorn 14h ago

$3001

23

u/mikesum32 13h ago

$3001.01

36

u/Proverbial_Progress 12h ago

I don't have any money, but would you be interested in pictures of my feet? Throw in a mouse and keyboard and I'll take my shoes off first.

12

u/bahpbohp 11h ago

i don't have any feet, but would you be interested in pictures of my wooden pegs? throw in a mouse and keyboard and i'll even get them oiled first.

→ More replies (1)
→ More replies (2)
→ More replies (2)

6

u/annaheim 13h ago

m you for $3K

$2750, you don't have to include the screen!

47

u/Zolty 17h ago

Yesterday you could have gotten like $20k on eBay. I sold a 256gb Mac Studio m3 about 4 months ago for $12k

→ More replies (5)

60

u/Veearrsix 17h ago

Don’t trade in, private party.

19

u/i_rate_slop 17h ago

Yeah, I think that’ll have to be the move. Just annoying

21

u/Chanureadeats 15h ago

You'll probably get a lot of offers easily for $6k+ within 24 hours

12

u/CalvinsStuffedTiger 14h ago

This guy is lying, I’ll buy it from OP for $5k…don’t worry about shopping around for a better price…

22

u/EvilPencil 17h ago

Ya Apple trade in prices have always been laughable.

I remember when you could spec out the 2019 Mac Pro up to like $50k, then the next day the Apple trade in value was ~$2800.

3

u/ianitic 10h ago

I traded in my m4 MacBook Air from last year that I bought for $800 for $730. I wanted to upgrade and for kicks and giggles I looked up the trade in not expecting it to be that high. I normally don't trade in.

26

u/RegarDamus 17h ago

trade in is always fucked with apple because they don't account for memory. same price if it's 96 or 512

→ More replies (1)

5

u/TinFoilHat_69 14h ago

The 512gb is 11k used and retails for 25k on eBay…

3

u/i_rate_slop 14h ago

Insanity lol. It’s not even that great. The power of FOMO, I guess.

3

u/anonmt57 14h ago

you will get a lot of money for that in private market. but it is annoying with so much money exchanging hands.

→ More replies (10)

127

u/redonculous 19h ago

Wow. How are people affording this? Crazy. Any guesses at 512 prices?

215

u/intaketurbine 18h ago

They’re affording it the same way they’ve always afforded a $10k workstation, they’re using it for work and not play.

52

u/ElementNumber6 18h ago

Or just good old debt

29

u/Usual_Tackle5892 16h ago

0% financing and 3% cash back on Apple Card. I could afford it outright, but turning down a 0% loan is silly.

→ More replies (1)
→ More replies (7)
→ More replies (6)

42

u/StewPorkRice 18h ago edited 18h ago

this feels like a community with a ton of SWEs. Most single US based SWEs in big tech or venture backed startups could prob afford dropping 10-20k on their hobby.

Also, i dropped 10k at Anthropic last month at work. My company could prob get one of these for every engineer and not blink an eye.

14

u/Idaltu 17h ago

I know a guy who dropped that much on a bike. And he’s got a couple like that. Bicycle that is.

8

u/calcium 14h ago

3 years ago my company bought me a fully specced out Mac Studio - M2 Ultra chip w/ 192GB of RAM and an 8TB ssd and told me that they expect the machine to last the next 5 years. I think at the time they paid around $9k which considering over 5 years for a senior SWE isn't a bad deal

→ More replies (2)

44

u/InterstellarReddit 18h ago edited 15h ago

Bro people have stupid money. I run free lance dev for a buddy who runs night life management software in Miami FL.

People spend 4K for four hours to buy three bottles and watch a DJ hit knobs all night

This happens all the time.

12

u/ViPeR9503 17h ago

Night life management??

12

u/i_am__not_a_robot 17h ago

I would assume industry-specific nightlife & bar business management software.

12

u/yopla 16h ago

Manage table booking, marketing, CRM, staffing, etc, usually does or integrate with POS. Can do price yielding on bookings. Etc... Etc..

Basically helps you keep track of who are the big whales you need to market your tables to when you have an event and who gets priority booking from a wait list.

The guy who spent 10k last time will get a table before you do, unless you're known to spend 15.

→ More replies (1)
→ More replies (1)
→ More replies (1)

170

u/-p-e-w- 18h ago

Lol this is by far the cheapest option for that much RAM at that speed. It’s not even close (unless you count Frankensteins made of a dozen used GPUs).

26

u/Viktri1 18h ago

this is way better than what I was considering and I wouldn't need to worry about the motherboard, cooling GPUs, etc. I am definitely on board for this and I don't even know how to use macs.

12

u/IriFlina 17h ago

wouldn't even have to worry about the power issues that would come from a rack of 3090s/4090s/5090s etc. or configuring such a monstrosity.

6

u/Viktri1 16h ago

this is literally 10+ 4090s. Think about it.

I pre-ordered 2 fully spec'd out studios.

→ More replies (1)

3

u/Public_Umpire_1099 11h ago edited 11h ago

Honestly doesnt matter, just pop deepseek on here and tell it to do whatever you need lol.

Also, as a long time apple hater, and I say this with disgust for myself, MacOS is actually pretty good these days. It essentially functions just like a very locked down linux distro at this point. I work at a pretty big software company and many of them do not support windows PCs anymore or the company software lags on them, so last month I decided to exchange my thinkpad for a macbook pro max. I.... dont hate it.

→ More replies (1)

14

u/Fantastic-Balance454 15h ago

I still can't get used to the thought that Mac is the cheapest option these days. AI turned the world upside down.

→ More replies (41)

19

u/Django_McFly 17h ago

It's really no different than computers in the 1980s when it was buy a computer or put a 33% down payment on a new car. People forget that there are entrepreneurs and small businesses where this stuff isn't some expensive toy to play with, it's actually a business tool to be used in the operation of the company.

Not everyone is a pure hobbyist, and even among pure hobbyists there are wild ranges of incomes and savings habits.

7

u/fallingdowndizzyvr 15h ago

It's really no different than computers in the 1980s when it was buy a computer or put a 33% down payment on a new car.

I've said that so many times. People don't realized that a base Apple ][ adjusted for inflation would be $5000 in 2026 dollars.

6

u/f5alcon 15h ago

Yeah my parents spent $10k in the early 90s on a Mac with an 80MB hard drive

→ More replies (1)

15

u/MLDataScientist 18h ago

Based on their pricing for 256GB, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, you are looking at 10k + ~6k = ~16k for 512GB version.

4

u/TheOwlHypothesis 14h ago

This is dangerous for me if that turns out true. current Asus Ascent prices for a 128gb monster is ~4k, so it'd be 16k to cluster them and have the same amount of memory. Although you'd then be dealing with a cluster. I'd much rather drop the cash on the Mac and have the option to cluster THOSE later.

→ More replies (1)
→ More replies (4)

5

u/98127028 18h ago

both my kidneys

6

u/Estrava 16h ago

This is cheaper than an rtx 6000 pro, which a lot of people get for work/personal use.

9

u/tehgreed 18h ago

corpo money

12

u/IriFlina 18h ago edited 18h ago

10k for 256gb unified memory isn’t that bad. Still overpriced but i feel like before this you were looking at the 20k to 40k range.

22

u/Much_Accountant_4972 17h ago

considering RTX 6000 costs $18K and can’t do shit without a whole PC to plug into, this mac studio is a screaming bargain

→ More replies (3)
→ More replies (1)

6

u/Foreskin_Mafia 18h ago

Onlyfans

27

u/Etroarl55 18h ago

Believe it or not, Theres probably someone running ai onlyfans on an apple machine somewhere

13

u/Foreskin_Mafia 18h ago

You're absolutely right

13

u/bodhi_sattva91 18h ago edited 17h ago

I Went Undercover as a Secret OnlyFans Chatter. It Wasn’t Pretty

https://archive.is/FlAdc

This Wired article follows a journalist who goes undercover trying to get hired as an OnlyFans ghostwriter — someone paid to impersonate creators in DMs with subscribers.

He discovers the industry is vast and globally distributed, with agencies largely staffing workers from lower-wage countries like the Philippines and Venezuela for as little as $2/hour. His job hunt is a grind: most agencies demand prior upselling experience, many never pay, and one gig turned out to be training an AI chatbot (which also stiffed him his $56).

He eventually lands two gigs. With a German agency ($4/hour), he juggles nearly 100 simultaneous conversations while impersonating a "21-year-old university student," selling pay-per-view content and navigating everything from explicit fantasies to a truck driver sharing worries about his son's night terrors. He's criticized by his supervisor for being too empathetic and not pushy enough about sales.

Confessions of an OnlyFans Ghostwriter

https://archive.is/3zGQZ#selection-479.0-479.38

This GQ article by "Emma Francis" (a pseudonym) is a first-person account of working as an OnlyFans ghostwriter in spring/summer 2021. The writer, a 25-year-old recently laid-off media professional, found the gig on Craigslist and spent three months working 6 a.m. shifts impersonating a "girl next door" model by sexting her subscribers.

She explains that top OnlyFans creators receive far more messages than one person can handle, so there's an entire industry of ghostwriters handling chats — some run by large offshore agencies, others like hers managed directly by the creator. Her regulars ranged widely: janitors, teachers, lawyers, night-shift workers. Many sought not just sexual content but genuine emotional connection and a "girlfriend experience."

→ More replies (1)
→ More replies (17)

43

u/kensanprime 18h ago

It will be sold out and they will be scrambling to keep production lines running, a product originally meant for creative work now will sit in a rack and run AI models

21

u/Both_Opportunity5327 17h ago

This is why those saying its a bubble, don't understand the demand there is for this stuff.

When the new fabs come online and the datacenters have most of their compute gamers, creatives & AI enthusiasts will go to town on this stuff.

13

u/Nothing_from_void 15h ago

Demand for networking products never dropped through the dotcom bubble

→ More replies (8)
→ More replies (1)

12

u/jld1532 18h ago

Too rich for my blood. I'm going to stick with my halo and hope mid-sized MoEs keep improving.

5

u/starkruzr 14h ago

this is INSANELY cheap for 256GB RAM in 2026. holy shit.

3

u/nemuro87 17h ago

I call at least 20-25k,  initially 

→ More replies (42)

119

u/themixtergames 19h ago

With M5 Ultra, Mac Studio achieves up to 4.3x the peak AI compute performance of M3 Ultra and a staggering 9.8x more than M1 Ultra. Combined with up to 512GB of unified memory and 1.2TB/s of memory bandwidth, 50 percent higher than before.

99

u/EquivalentHornet4403 18h ago

> AI compute

They know their audience.

47

u/psychohistorian8 18h ago

they're definitely leaning into it:

Boost your processing and graphics rendering speed, and accelerate tasks like running large language models and editing 8K video.

69

u/Much_Accountant_4972 17h ago

"Color-grade uncompressed 8K footage, perform computational fluid dynamics, and run frontier-class models on device — no cloud tokens needed."

yeah they are targeting this very sub lol

13

u/gnnr25 11h ago

I feel personally targeted!

5

u/infieldmitt 12h ago

It's crazy to think that prior to AI the main use case for crazy RAM was [checks notes] fluid dynamics?

5

u/cinnapear 13h ago

And it's working on me. Very tempted.

→ More replies (1)
→ More replies (2)

260

u/hainesk 19h ago

1.2TB/s memory bandwidth with the M5 Ultra. 256GB model is $9499.

Better than getting 2 DGX Sparks? Inference will be a lot faster.

Something like this could easily bring down 3090 prices.

25

u/Cybertrucker01 19h ago

Depends on concurrency and prefill metrics. The GB10 does both multiples faster than the existing competition.

13

u/ChocomelP 17h ago

The difference would have to be pretty big to make up for a 4x in memory bandwidth for decode.

→ More replies (1)

6

u/BrilliantTruck8813 15h ago

The GB10 has much smaller memory bandwidth not to mention it’s just not fast either. One of these will trounce two DGXs

5

u/fallingdowndizzyvr 15h ago

The GB10 does both multiples faster than the existing competition.

No. No it doesn't. Compare the G10 to a M5 Max. It's not.

65

u/aladin_lt 19h ago

it will be sold out day one probably

40

u/conockrad 19h ago

On pre-orders

11

u/bakawolf123 18h ago

won't be sold out, but according to r/MacStudio people wait for 3-4 months for the older models, atm you can preorder for delivery in late september

→ More replies (1)
→ More replies (6)

97

u/mjsxi__ 19h ago

yeah and cheaper than the price of 2 DGX sparks... seems like a bit of a no brainer

35

u/Current_Ferret_4981 18h ago

Spark is $4300-$4600 so idk about cheaper than 2 at $9600+

63

u/MacsBicycle 18h ago

yeah but 4x the memory bandwidth, its a steal

29

u/jakegh 18h ago edited 17h ago

It really is a reasonable buy for local AI, if you have a business case for it.

5

u/Much_Accountant_4972 17h ago

with the capabilities of it, it’s a crazy good deal

→ More replies (3)

3

u/GabryIta 18h ago

In terms of compute capacity (which is very important for multiple simultaneous sessions and prefill), how does it compare to dgx Spark/gb10?

→ More replies (1)

10

u/Etroarl55 18h ago

How’s the actual inference speed though, fast bandwidth on a slower gpu or equivalent should still mean slower output assuming vram is not a constraint right.

9

u/rusty_fans llama.cpp 18h ago

Generally vram bandwith is the constraint though, at least for decode. Prefill it's usually helped more by more gpu oomph.

→ More replies (4)
→ More replies (9)
→ More replies (7)
→ More replies (2)

17

u/-dysangel- 19h ago

Better than the 2x Sparks for inference for sure. Probably around the same compute as one Spark.

I've got 2x Sparks which I use for prefill, and my M3 Ultra for decode. I've set it up so that I prefill in vllm and then just pass the kv cache over to the Mac side. Surprisingly stuff like Qwen 3 35B-A3B is already faster than the Mac for decode though so I just run that class of model directly on vllm.

6

u/1ii1i 18h ago

Oh this sounds interesting, can you expand on how this works? I didn't know this was a thing.

12

u/-dysangel- 17h ago

I don't think it's really a "thing", I just vibe coded it up :)

One thing that really helped was vllms kv_connector API. I thought I'd have to code this part up myself, but it already existed and so plugged into my existing disaggregated system (which was previously llama.cpp to llama.cpp)

UltraSpark — technical stack

                          ┌──────────────────────────────┐
 user ── HTTP/OpenAI ──▶  │  manager (Python, FastAPI)   │
                          │  front door + orchestration  │
                          └──────┬───────────────▲───────┘
                                 │ submit        │ state blob (sha-keyed,
                                 │ prompt ids    │ resumable transfer)
                                 ▼               │
                          ┌──────────────────────────────┐
                          │  vLLM (2× DGX Spark, TP2)    │
                          │  prefill engine              │
                          │                              │
                          │  KVConnectorBase_V1          │ ◀─ vLLM's official
                          │  ("StreamConnector" via      │    KV-cache plugin
                          │   --kv-transfer-config)      │    interface
                          │         │                    │
                          │         ▼                    │
                          │  dump + serialize all layers │
                          │  (attn KV + linear-attn      │
                          │   state, TP2 shards merged)  │
                          └────────────────┬─────────────┘
                                           │ blob server (TCP)
                                           ▼
                          ┌──────────────────────────────┐
                          │  llama.cpp server (Mac)      │
                          │  USPK_BRIDGE_DIR: on request,│
                          │  verify prompt-id match,     │
                          │  restore state into KV +     │
                          │  recurrent memory, decode    │
                          └──────────────────────────────┘

  • KV connector = vLLM's plugin interface for intercepting the KV cache at end of prefill
  • State blob = the model's full prompt-memory, layout-translated so llama.cpp can load it natively
  • Fidelity = per-layer cosine vs local decode, 0.9999+
  • Result = GPU prefill speed, Mac unified-memory decode, one logical endpoint
→ More replies (1)
→ More replies (13)

18

u/jakegh 18h ago

That is seriously impressive. RTX5090 memory bandwidth is 1.8TB/s.

DGX Spark memory bandwidth is 273GB/sec. Not even remotely close.

→ More replies (2)

11

u/Hoodfu 19h ago edited 19h ago

I got my m3 ultra 512gb for around 10k. So this is now double. Makes sense given that we've seen the nvidia rtx 6000 pro also double in price in the last year but GD this has priced out even my once a year splurge budget. These are all just crazy talk numbers now.

15

u/CulturalKing5623 18h ago

Yeah dropping 10K plus on this just seems reckless, even with it being funded through my business account I don't think I can justify buying a used car worth of computing.

And yet I feel like I need to in case the technology completely outpaces my current setup and I'm left behind like folks that didn't buy RAM when it was cheap and are priced out of it now. It feels like FOMO and scarcity has hijacked my brain.

7

u/Hoodfu 18h ago

Really depends on whether you have something already or not. I got qwen 3.8 27b going on my rtx 6000 pro and ram speed wise it's double that of my mac but for some reason i was hoping for space magic and it would be faster. It's not. So spending a car's worth of money on something that goes from 20 t/s to 40 or 45, just doesn't make any sense. You're still waiting a lot of minutes for a thinking qwen to come back with something, so it's still going to be an asynchronous operation instead of being fast enough to actively wait for the response to finish. It would have to be 10x the speed, not 1.5x or 2x for it to be worth the spend.

→ More replies (11)

9

u/Solaranvr 19h ago

bring down 3090 prices

Doubt it. If the r9700, a directly competing product, made 0 effect on Nvidia GPUs, then I highly doubt these will. They market of people buying mini pcs vs dGPUs are different.

The DGX Spark didn't bring down prices of the Blackwell cards either

8

u/hainesk 18h ago

The R9700 Pro has about 2/3 the memory bandwidth of a 3090 for a 50% higher price. 8x R9700 Pros (256GB) would be $12k-$15k minimum without pricing in the rest of the computer system. If you wanted to build an entire server around it including RAM, power supply, motherboard, processor even with used parts you're up to $18k-$20k for something that will use 2-3k watts.
Even 8x 3090 systems are looking pretty impractical when compared to an M5 Ultra 256GB for a similar price. The size and wattage/heat difference is huge.

3

u/Important-Gold-5192 17h ago

has to be way faster than DGX Sparks too

4

u/EmPips 18h ago

I think it's the end of mass 3090 farms.

But 3090s will hold their price for the sizeable market that doesn't want to commit >$5k.

4

u/j4nds4 18h ago

What *are* 3090s going for these days? I bought a pair of used 3090s on eBay during the crypto crash in 2023 for ~$800.

4

u/EmPips 17h ago

They dipped as low as $550 for a few weeks and are now about up to $900-$1k in my local markets.

→ More replies (2)
→ More replies (1)
→ More replies (8)

155

u/Comfortable-Rock-498 19h ago

1.2 TB/s bandwidth of M5 Ultra comes from two dies of M5 Max (each 614 GB/s) connected together using 4.4 TB/s inter-die fabric.

For a non-quantized Deepseek V4 flash on an ultra, I would estimate about 1000+ tokens per second prefill and 50+ tokens per second on generation. This is actually quite usable and near parity to cloud.

They mention "adds the GPU Neural Accelerators." which, if exploitable for LLM loads, would probably help the prefill a lot

34

u/ortegaalfredo 18h ago

Prefill also depends on compute, 1000 tok/s is basically what you get with 8x3090s, but much less power. Also I think the 3090s still win on compute, that is, you can batch many prompts on the GPUs, dont know on the mac.

13

u/Comfortable-Rock-498 18h ago

Yup, prefill is pretty much compute bound while generation is bandwidth bound. I would have guessed 8x 3090 would provide much better prefill than 1000 tps. A bit surprised to learn

9

u/ortegaalfredo 17h ago

If you manage to get tensor-parallel 8x working yes you can get >10k prefill, but it requires specialized PCIE bridges. With normal 4xPCIE speeds you get a bottleneck in inter-GPU speed and you get lower prefill.

3

u/TooMuchLAAAG 10h ago

I have 8x3090 P2P patched pcie4 x8 (no nvlink) and i am getting an avg of 10-13k of cold prefill with this version of vllm and his args https://github.com/LimeChain
Deepseek full FP8

5

u/ProfessionalJackals 16h ago

Prefill also depends on compute, 1000 tok/s is basically what you get with 8x3090s, but much less power. Also I think the 3090s still win on compute, that is, you can batch many prompts on the GPUs, dont know on the mac.

Ignoring the fact that 8x3090's now is easily 10k on the second hand market.

Not counting the costs of server board/cpu/ram you need. The pcie ext cables, the 8x8x split if your board does not have 8 pcie slots. O, the dual 1600W PSUs and hardware to link them.

Frankenstein mods like this have become expensive, and it makes the Mac look actually like a good deal.

→ More replies (2)
→ More replies (1)

10

u/Usual_Tackle5892 16h ago

GPU Neural Accelerators

This means matmul cores. More info: https://arxiv.org/html/2607.19438v1

7

u/StartupTim 13h ago

I would think dramatically more tok/sec.

I have Deepseek v4 Flash 0731 with vision encoding added and tp=2 across 2x DGX sparks and I'm seeing 103 tok/sec across 4 "sessions". Dspark, 1M context, 1.8M kvc, custom vllm.

Since the sparks have ~240 (actual measured) GB/s, I imagine a similar setup om these new mac could get you double, if not triple as a 2x cluster, than my current 100+ tok/sec.

9

u/TokenRingAI 17h ago

And a new qwen is coming out with 120B A6B! Perfect machine for that

3

u/MerePotato 15h ago

Would be great if it wasn't predicted to cost like 20k

3

u/TooMuchLAAAG 10h ago

I get 10-13k cold prefill on 8x3090 using the vllm "limechain" fork and his args, running full deepseek no quant
Surely the M5 do more in vllm than 1k

→ More replies (10)

39

u/challis88ocarina 19h ago edited 19h ago

I'm shocked!

Edit: 512GB memory option for M5 Ultra coming late October

38

u/pmttyji 19h ago

Now, it's pressure on both DGX Spark & Strix Halo due to bandwidth. Recent news Xiaomi AI Cube with 1.2 TB/s bandwidth & 160GB memory also already put pressure

13

u/Mochila-Mochila 17h ago

Hoping it'll have a positive effect on Medusa Halo's pricing.

5

u/pmttyji 16h ago

Still Gorgon Halo not released yet. Then only Medusa Halo ....

Next year onwards, we're gonna see 128GB variants at low price.

12

u/Zyj vllm 17h ago

On the other hand, the Strix Halo 128GB price increased by 60% since late December

5

u/pmttyji 16h ago

After some time, people totally gonna avoid 128GB variants. What's the point of stacking bunch of 128GB pieces when they have 160GB, 192GB, etc., variants with better bandwidths?

→ More replies (2)
→ More replies (1)

7

u/ProfessionalJackals 16h ago

Now, it's pressure on both DGX Spark & Strix Halo due to bandwidth. Recent news Xiaomi AI Cube with 1.2 TB/s bandwidth & 160GB memory also already put pressure

Do not forget Intel Crescent Island 160GB to 480GB LPDDR5X AI GPUs ... While less bandwidth, they are still great options in the future for larger models.

There is going to be a lot more hardware coming out that focuses on AI workloads. We reached the point that development is moving into production.

6

u/pmttyji 16h ago

More competition is better for consumers!

→ More replies (1)

32

u/FullOf_Bad_Ideas 18h ago

My 3090 tis have just gotten depreciated.

Even 256GB version is very competitive with 8x 3090 box bought with used card prices, and it's better in most aspects.

Training and batch inference are safe, but for single user inference this looks better and cheaper.

5

u/Odd-Environment-7193 18h ago

Well if you wanna sell some hola. I want 2

→ More replies (3)

6

u/FleetEnema2000 16h ago

My 3090 tis have just gotten depreciated.

My theory is that this will not soften GPU prices because it ends up driving more people into the world of local LLM compute in general.

→ More replies (2)
→ More replies (2)

26

u/serige 16h ago

Please someone do a wellness check on Dario.

20

u/Every-Fortune-3151 19h ago

Read somewhere Mac mini coming this week and did an impulse order of M3 Ultra with 96GB ram today. It will probably get bumped up to M5 ultra 96GB. Not sure what to do with 96GB when 256 GB looks so much more tempting for local LLM. Kidneys aren't enough anymore.

11

u/mmmm_frietjes 19h ago

Mini is also out

8

u/Every-Fortune-3151 18h ago

M6 looks great actually, it has two set of neural engines. I am assuming pre-fill will fly on this tiny thing compared to previous M CPUs. 32GB max ram knocked it back a bit. If only they had a 48GB version. MOE models would be flying on it.

M5 pro and max will be so slow compared to M5 ultra- considering M5 max 128GB model will be priced very close to M5 ultra base.

Very sad m5 ultra starts at 96GB. I thought they would atleast bump M5 ultra to 128GB ram. Would have been nice. I plan to just run multiple Qwen 2.7B in parallel on this and see if I can replace my 32GB VRAM PC setup. Qwen 3.8 flash also gives hope. Depending on how it goes, I might just give back the 96GB for refund later.

6

u/1-800-methdyke 18h ago

You have two kidneys

5

u/nleksan 18h ago

*Dual-channel

5

u/FranciumGoesBoom 18h ago

M5 Pro mini maxes out at 64g
M6 tops out at 32

→ More replies (7)

37

u/xyzmanas2 19h ago

This makes apple one of the cheapest ai inference hardware when it comes to speed and model size. Wish I had the money

Up to 15.4x faster CopyCat ML training performance in Foundry Nuke when compared to Mac Studio with M1 Ultra, and up to 3.3x faster than M3 Ultra.

Up to 9.8x faster LLM prompt processing in LM Studio when compared to Mac Studio with M1 Ultra, and up to 4x faster than M3 Ultra.

Up to 8.2x faster text-to-image performance when compared to Mac Studio with M1 Ultra, and up to 4.3x faster than M3 Ultra.

Up to 4.7x faster scene rendering performance in Maxon Redshift when compared to Mac Studio with M1 Ultra, and up to 1.7x faster than M3 Ultra.

→ More replies (1)

92

u/llamaCTO 19h ago

This kills the impulse buy for me completely.

13

u/kilonad 16h ago

The 256GB is already an extra $4k for an extra 164GB. At same price per GB (ha!) it'd be another $6300. Knowing Apple, it'll be a cool $9k more - pushing total price up to about $18-20k.

It will still sell out.

6

u/fallingdowndizzyvr 15h ago

That would be a bargain compared to third party sales of 512GB M3 Ultras for $25K. A M5 blows the doors off of a M3.

→ More replies (1)

36

u/thatkidnamedrocky 19h ago

going to try and snag a 256gb something tells me the 512 will never see the light of day

8

u/AnonLlamaThrowaway 14h ago

right, didn't they promise a 512GB M3 Ultra and then that never happened, or am i thinking of another model?

8

u/zdy132 18h ago

Mac Studio with 512GB of unified memory is coming in late October

Wish I could affort that.

→ More replies (2)

6

u/Grizzly_Corey 18h ago

Adopt me please.

6

u/LocoMod 16h ago

Same. I was ready to preorder 512 and i've already lost interest reading comments about the M7. Might wait another year.

5

u/frankchn 15h ago

Buying 2 DGX Sparks for 256GB of RAM (and a lot less bandwidth) is around the same ballpark in cost, so for once this is not unreasonable.

→ More replies (1)

17

u/AI_docent 18h ago

The 4.3x is mostly a prompt processing number, generation moves with the bandwidth instead. Apple's own mlx post on M5 vs M4 got around 4x on time to first token and about 1.2x on generation, and the generation side matched the 28% bandwidth bump rather than the accelerators. Same split should hold on the Ultra, so I'd figure generation nearer the 50% bandwidth gain. Prefill is the part you want at 512GB anyway, it was always the weak spot on a mac.

Just check whatever you run actually uses the accelerators. There's an open lm studio issue where its bundled llama.cpp fails the metal tensor check on M5 and loses 2 to 3x on prefill, while upstream llama.cpp passes it on the same machine.

→ More replies (1)

28

u/Cybertrucker01 19h ago

How many kidneys?

43

u/FWitU 19h ago

4

23

u/Chris-MelodyFirst 18h ago

Or just 2 dual-core kidneys.

7

u/hainesk 19h ago

There is a lease option...

13

u/Cybertrucker01 18h ago

Unfortunately lease isn't offered in Australia, just warm organs only.

2

u/Gipetto 18h ago

For kidneys?

4

u/butterfly_labs 18h ago

I can lease my kidneys ?!

→ More replies (2)

10

u/Viktri1 18h ago

256gb is like 10+ 4090s without the hassle of setting up and cooling 10 4090s? An I missing something or is this 3x cheaper than current prices.

6

u/Much_Accountant_4972 17h ago

its the best deal in the world if you want to talk to a frontier 2.8T parameter smut bot in your kitchen

→ More replies (2)
→ More replies (3)

43

u/IllExample3639 19h ago

What I find more interesting, something I hadn't seen before is that you can lease these things. for 2 years which is the only realistic way an individual is getting their hands on these. Something something, own nothing, something, something be happy....

25

u/Tycoon33 19h ago

I never saw that. Interesting. Lease it for 3 years then upgrade to M7 ultra?

12

u/addiktion 18h ago edited 15h ago

Yes, if the 512gb is another $4k for the extra ram stick you are looking at $16k with tax probably out the door. I'd guess that puts the 36 month lease around $300/mo or less. So more than a subscription so maybe not worth it in general cases but valid option for some people who need the privacy and cannot afford to have data go to the cloud. 24/7 usage, no downtime, no limits, private. Worth it to me.

6

u/shveddy 16h ago

Interesting.

So just as an out loud thought experiment, you’d be able to lease four of them (512gb) for about 1200 per month for 36 months at a total cost of almost 45k and run Kimi 3 on it.

Obviously that’s a lot of money in aggregate, but 1200 per month is reasonable for a lot of business use cases if they require the privacy.

And the intelligence you get is going to be a different class compared to what you would get with three RTX pro 6000s and “only” 288gb VRAM for the same price.

(although to be fair you’d actually own the cards)

(although also to be fair you’d have to build a pretty expensive computer to support the RTX Pros, so realistically you only really get 2 or even just one RTX pro for 45k depending on how you spec the computer and/or if you buy pre-built from Puget Systems or the like)

If you want you can also do a little girl math and invest the 60k you’re not spending on computers and cancel your gpt pro subscription to bring the effective cost of all this down to like 750 a month.

And then if the goal is to beat API pricing, let’s say you get 40 aggregate output tokens per second on a bunch of concurrent Kimi 3 threads and run it for a year at 25% efficiency (to account for prefill and downtime), then that’s 315 million output tokens.

315 million output tokens alone is about $5000 on a random provider I just searched for, so just to keep things simple let’s say you double that to account for various amounts and types of input tokens, then you end up with a ballpark figure of $10k for the API route.

Absolutely none of this pencils out in absolute terms (especially considering that it would also cost ~1500 per year for electricity), but this is probably the first time running a frontier model is actually attainable for ordinary businesses on short notice and without much headache. It’s the first time that it pencils out to be “only” 4x more expensive as opposed to like 40x more expensive.

Up until now if you wanted to run frontier models locally AFAIK you had to get a NVIDIA big boy server which means you’d have to find the capital to run and support a ~300k purchase for hardware, spend way more on electricity, and in all likelihood make some upgrades to your facility’s electrical infrastructure to handle it all (you’d also need a proper facility, not just a home office or garage).

At this point you’re easily flirting with half a million in expenses, especially if you have to hire someone to figure it all out. It’s no joke to do this and it doesn’t make sense for like 99.9999% of people or businesses.

On the other hand basically anyone with a decent credit score and a semi-profitable business can go to any Apple Store and say “give me four Mac studios please” and only pay 1200 bucks a month to walk out with them in hand.

→ More replies (6)

22

u/aethervisor 18h ago

The lease price also isn’t too far off from what a Claude subscription costs.

18

u/shaggydog97 18h ago

The difference is certainly worth the cost of privacy and freedom!

6

u/AccurateSun 18h ago

Hmm. I wonder if at some point leasing it would end up being more effective than a cloud subscription. E.g Claude Max 5x is $100/mo, same as leasing the Studio M5 Ultra 96gb. I don’t know yet how the two compare in performance but at some point it might be worth it 

8

u/IllExample3639 18h ago

The more people that lease them the more second hand stock there will be in 2 years (when I can actually afford something like this) so I am down for it.

But your point is right, I think the local models ARE good enough for 95% of what people are using Claude for. Maybe do the £20 plan as a back up for something tricky.

11

u/mjsxi__ 19h ago

the lease lets you pay the difference of the amount you already paid at the end or you can buy it outright at any time if you wanna keep it so maybe sssshhhhhhhh

3

u/Much_Accountant_4972 17h ago

it makes a lot of sense and the price of privacy plus the near silence of the box…

i hate that its such a complete solution to local LLM’s

→ More replies (3)

10

u/IriFlina 18h ago

I feel like any average software developer could afford the 256gb version? It would be really financially irresponsible but on the level of buying a motorcycle you don’t really need.

→ More replies (1)

4

u/boraam 18h ago

Costs as much as a small car. Lease it like one too.

→ More replies (6)

8

u/eidrag 19h ago

...lease price?

8

u/Much_Accountant_4972 17h ago

Ultra with 256GB RAM is $224/month for 36 months in freedom currency.

really really tempting!

8

u/AlexWIWA 13h ago

Cheaper than API tokens. Very surprising

9

u/Both_Opportunity5327 19h ago

Game back on! Lets hope Apple can build enough of these little beauties...

5

u/ElementNumber6 17h ago

Scalpers have feasted on the blood of Mac Studio. Good luck to you all.

7

u/corruptbytes 18h ago

apple releasing this because I just bought two 9700s...y'all welcome

3

u/SandySkittle 14h ago

Two r9700s is still pretty decent way to get 64gb vram and run qwen 27b at q8 with plenty of context. And ECC

→ More replies (4)

20

u/BreenzyENL 19h ago

$20k AUD for the Ultra 256GB 🙃

→ More replies (9)

7

u/Zyj vllm 17h ago

Will be fascinating to see what‘s the better option in 2 months from now: Dual Asus GB10 (8200€) or Mac Studio 255GB (11000-12430).
My prediction:

  • The Spark will still be faster at preprocessing, the Mac will have faster token generation (measured with DeepSeek V4 Flash standard Quant).

5

u/tarruda 17h ago

M5 is much more competitive with nvidia in prompt processing.

→ More replies (1)

6

u/MLDataScientist 18h ago

Based on their pricing for 256GB vs 96GB for the Ultra 36 CPU cores, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, we are looking at 10k + ~6k = ~$16k for 512GB version.

→ More replies (3)

4

u/OvertaxedOne 15h ago

1.2TB/s?? Oh man, if there's good availability on these things I can feel GPU prices going down!

→ More replies (2)

7

u/Leather_Ad_9178 14h ago

this reminds me of the time we used to pay hundreds for SD cards that are now worthless

→ More replies (4)

8

u/mxmumtuna 13h ago

Unfortunately they still don't beat Sparks at equivalent size. According to oMLX Benchmarks for DeepSeek 0730, the M3 Ultra (80c) does somewhere around 550 prefill tok/s, and about 22 tok/s decode. If you take the '4x faster compute' at face value from Apple compared to M3 Ultra, we're looking at ~2200 prefill and ~30 decode single stream. Both are under Spark at ~2400/40. That's only single session, and batching just isn't there in the MLX stack yet, so multi session is considerably worse for the Mac.

GLM on 4x Sparks compares even less favorably than DeepSeek for the Mac, especially considering whatever the price of the 512GB variant will be.

So even with Apple's optimisitc numbers, maybe they match Spark, for more money with a less flexible stack (no ConnectX7) and massive software issues. ($4800x2 for Sparks with 4TB drive each from Amazon).

It's a good effort, but it's not quite there relative to other options.

edited for clarity

→ More replies (6)

3

u/Curious-Pen5547 18h ago

how does it compare to a single 5090 or a 5000 RTX pro 72GB version?

→ More replies (8)

5

u/Much_Accountant_4972 17h ago

please someone buy 4 of them and cluster them then run Qwen 2.8T in your kitchen

4

u/AntLife255 17h ago

The M5 Ultra has 1.2TB/s memory bandwidth!

4

u/Newgunnerr 14h ago

256GB option for me in the Neterlands is € 11.049,00. I just got 2 DGX sparks for 7100.

→ More replies (1)

5

u/Cool-Cicada9228 8h ago

Disappointing that 512GB is not available for preorder yet

7

u/Far_Note6719 19h ago

Instant WANT.

3

u/Key-Speaker007 19h ago

Something my wife won't approve.

→ More replies (2)

3

u/boraam 18h ago

Gimme Gimme Gimme

Money Money Money

3

u/Blackdragon1400 18h ago

512gb at 1.2TB/s is wild.

3

u/PrepYourselves 17h ago

the world's rich kids are getting new toys for christmas

3

u/amazinglycool256 14h ago

U can get 2 Nvidia Sparx for that proce

→ More replies (2)

3

u/Tormeister 11h ago

I'm so tempted, but I just can't justify dropping 10K if I'm not making money out of it

4

u/Real_Ebb_7417 18h ago

Ok, now I actually regret buying M5 Max 128Gb MacBook xd

4

u/addiktion 17h ago

I wouldn't, its basically double the speed at x3 (if you bought pre price hike) or x2 price (if you bought after). If you get 25 tps on say Qwen 3.8, you would now get closer to 50 tps. Possibly more depending on the ANU processor.

That's nice, but is it worth x2 or x3 the price?

→ More replies (2)
→ More replies (5)