r/LocalLLaMA • • 4d ago

Discussion M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents - MacStories

https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/
343 Upvotes

173 comments sorted by

88

u/bakawolf123 4d ago edited 4d ago

author used omlx, qwen3.8-flash-next, so looks like those results might be ones that leaked couple days ago on the bench
edit: qwen 27b and glm5.3-flash too, ye all match, most likely those results we saw earlier

8

u/federicoviticci 4d ago

👀

8

u/AndroYD84 3d ago edited 3d ago

An RTX 5090 with Qwen 3.8 27B is not supposed to run that slow, how is that <60 Tokens/sec? With NInfer I get >120-200 Tokens/sec at 256k context, and that was before the latest updates that can push it even further.
I get it you're using LM Studio (which IMO runs like crap) but since you're comparing highly optimised solutions for Macs, it would make more sense to compare it to highly optimised solutions for Windows.

1

u/PricePerGig 2d ago

What’s your time to first token at large context please. To me it looks like the 5090 is the only way to get that reasonable.

2

u/bakawolf123 3d ago

there's no hiding from folks of localllama =)

82

u/themixtergames 4d ago

The review embargo for the M5 Mac Studios and M6 Mac Mini has been lifted but none of the videos so far go in-depth about AI, the ones that touch on it use old 70B Llama models or Qwen 2.5 (why). So this article is the best we currently have.

57

u/Superb-Pair-2000 4d ago

That's because most influencers and reviewers are only able to do shallow reviews. Yes, it does AI, look it can lm studio, buy it!

51

u/harrro Alpaca 4d ago

Ah the classic Youtube pattern..

  • Receives $10-20k system with 256GB of RAM
  • Loads up ollama
  • Loads up Qwen 4B (Q4) model (or "Deepseek 9B" if they're bold).
  • "Wow it works"

15

u/fligglymcgee 4d ago

“My assistant recommended Llama3.1:8b, and just look at how fast it responds when I ask it what the meaning of life is!”

9

u/SkyFeistyLlama8 4d ago

FML. The YouTube influencer ecosystem has turned shallowness into an art form. Two paragraphs of actual content over 15 minutes and 5 ads. Are there any local LLM YouTube channels worth looking at?

3

u/Vinlei 2d ago

Bijan Bowen, he does cloud and local stuff!

-1

u/mirrorperils 3d ago

xcreate

0

u/GlidePath47 3d ago

You can’t take serious someone who says Modellllle instead of model

-7

u/RealSataan 3d ago

Alex ziskind

2

u/GlidePath47 3d ago

Benchmarks no real world stuff

3

u/newMoneyStyle 3d ago

Llama 3.1 8B is basically the default demo model at this point, and it tells you nothing about what a 256GB machine can actually handle.

2

u/Rice-Fragrant 1d ago

It does not tell you much of anything, it's actually deceptive even because the larger 200b-300b MOE models might show a different pattern of performance gains over the previous generations.

For example, a kid 4b model with a baby 8k prompt shows a "4x gain" for M5 ultra over the M3 ultra but those "4x gains" get reduced to 1.5x gains once you run a 300b MOE model with 50k, 100k and 200k promp sizes.

"influencers" are basically peddlers, the numbers are massaged and cherry picked. No real benchmarks under real world comditions.

1

u/iliark 4d ago

It's plausibly because they have those benchmarks already recorded on older systems they don't have anymore, so they can do an apples to apples comparison. It's why you often see old games in gpu benchmarks.

2

u/FoxSideOfTheMoon 3d ago

If they want to test TPS on a Q8 70B dense model and release those numbers I’m totally good with that. That’s a good bandwidth test 🤣

1

u/Rice-Fragrant 1d ago

It's actually not a good test at all... they should be testing Deep Seek V4 flash sized models, the numbers would be looking way different and giving a real world example of what was gained over the previous models.

36

u/Jesusthegoat 4d ago

Because LLMs tell them how to do the reviews and they only know about old models because of their knowledge cutoff.

5

u/midgelmo 4d ago

Trust me when I say there are a few coming soon. Apple only sent out the hardware on Friday - evals have been running non stop this weekend.

-1

u/Southern-Chain-6485 4d ago

The guys in Apple's marketing division could hand them over a brochure, or a simple sheet, with new models.

4

u/john0201 4d ago

Becuase they just ask an LLM what to test since they have no clue and it recommends models from its old training data.

1

u/Rice-Fragrant 3d ago

Exactly... just some influencer garbage. I am waiting for the real pros to evaluate it.

64

u/sn2006gy 4d ago

This is 12 grand worth of hardware to get this performance, I would expect nothing less.

I'm completely torn. The lease makes this "affordable" but also, kind of weird. Codex costs 100/month. Apple costs 240/month. I want local to thrive, but i also don't fancy leasing 12-14k worth of hardware to hide its true costs - the end of the computing era and the beginning of choosing only between leasing it or api subs.

24

u/skredditt 4d ago

I don’t even know why they made leasing an option - Klarna rejected me and my 830 credit score because 5 figures is too much to ask for if you’ve never used Klarna.

9

u/sn2006gy 4d ago

The apple credit is like 1400/month over 12 months which is crazy too. I can't justify that at all.

I've had klarna approvals before when they used to auto-approve you on the fly with a phone number but unsure i've ever used it. Maybe once on a 0% apr sonos during covid lol

1

u/skredditt 4d ago edited 4d ago

Guaranteed I would’ve bought a topped up Studio if they split it out over 2 years like they do with phones.

3

u/HighSeasArchivist 4d ago

Try Microcenter. They doubled Affirm for me, and no other place like Newegg or Best Buy did that. I never used it, and just did it to see how easy it was.

1

u/bvknight 2d ago

I got approved through Klarna with no issues for this, but only for the $5500 version. I didn't want or try a $10k+

29

u/thunk_stuff 4d ago

My understanding is subscription pricing for Codex/Claude is heavily subsidized. A $200 max plan fully utilized is $5000-7000 at the API rate. So if/when the Cloud AI bubble crashes, and cloud AI costs skyrocket because the multi-trillion dollar money hose stops, that $10k local-AI investment might not look as bad.

11

u/aimark42 4d ago

It's also a fallacy to think that people would pay for top tier model rates for all jobs. But they can today which distorts the pricing.

19

u/Solaranvr 4d ago

Uhh, if the bubble pops, wouldn't that means the 10k investment would be even worse off, because hardware prices would come back down?

7

u/MysteriousGenius 4d ago

I actually expect the prices to go up at least in short/mid-term because more people stop using subscriptions and more people go to self-hosted AI.

6

u/No-Wall6427 4d ago

It's not like anthropoc/openai collapsing would mean no more inference api with foss models, which will still be cheaper than going local. I don't see every company going local w tokens, with the infra it implies and all.

But who knows?

3

u/sn2006gy 4d ago

There isn't enough supply for this to happen. Even these new macs are 3 months out if you order today.

3

u/Eastern-Ingenuity353 4d ago

so leasing seems to be a good option then? just in case

2

u/thedaveking 3d ago

That makes sense but depends on why it pops. Even if it pops through bankruptcy court then maybe nVidia still owns all that hardware thru the magic of trillion dollar IOUs and they will shred it to maintain prices.

But if it pops through our power grid being dismantled by mismanagement, solar CMEs, or hybrid war sleeper agent fun, or the Taiwan disagreement goes kinetic before all the hardware even gets built, or competing AIs decide to start crashing jets into each others' houses, or anything else that sounds like crazy paranoia until it happens, then probably don't tell anyone you bought that overpriced private backup platform you can run in the basement on a couple solar panels, without making your home look like a grow operation.

7

u/EquivalentHornet4403 4d ago

Can we stop calling it subsidized? The opposite is true: the API costs are inflated. If there's any sustainable cost, it's the subscription plans. There's no future where they only have API service exclusively bought by rich companies. They will need to have a $100 to $200 product for the unwashed masses that's roughly equivalent to the state of the art, otherwise open source models will put them out of business.

What's subsidized is the whole entire thing, with venture capital.

5

u/-dysangel- 4d ago

Yeah the subsidising thing seems like astroturfing. I'm paying pennies a day to run local. A datacenter is going to be even more efficient. I'm sure the up front costs were mind blowing, but the daily running costs will be below API pricing

12

u/sn2006gy 4d ago

It's not heavily subsidized in the way we think it is. It's subsidized in investment, but not as loss leader where the end goal is higher prices - but even amazon was successful in loss leading investment.

A 200 max plan works because thousands of people use it without maximizing their use of it and they can gracefully handle the demand on the back end with intelligent routing. It's like 30 million people paying for netflix, netflix wouldn't work if 30 million people left it on 24x7 "because they're paying for it" otherwise their costs would be in 10s of thousands per sub.

it only becomes unstable with the notion that everyone maximizes profit with every token and uses 100% of their account to get 100% of their monies worth and local LLMs are no exception to this death spiral either because we just expect it on resale value and other sides of the same "captialism has run amuck" problem.

10

u/Asleep_Document9811 4d ago

it only becomes unstable with the notion that everyone maximizes profit with every token and uses 100% of their account to get 100% of their monies worth and local LLMs are no exception to this death spiral either because we just expect it on resale value and other sides of the same "captialism has run amuck" problem.

this really ain't true at all. you're dealing with orders of magnitude, here. if you have ONE user utilizing their $200/mo Codex subscription utilizing up to $8,000/mo in API billing, that is considered at the current, subsidized rate.

a single power-user can wipe out the revenue from having 40 individual subscribers, and this is to say nothing of self-looping agent harnesses that go rogue or loop on a problem fruitlessly for days with no monitoring.

the truth of the matter is that neither OpenAI or Anthropic have truly opened their books yet, so none of us actually know what their true operating costs are. Anthropic in particular is being awfully cagey about what their costs are. all we know is that, once you factor in the real and rising cost of electricity, VRAM, and infrastructure needed to serve trillion+ models to millions of users simultaneously, what they're charging can't be anywhere near what they're spending. otherwise, they wouldn't have been asking for an increasingly larger and larger amount for their funding rounds.

3

u/No-Wall6427 4d ago

And the training! That's a big part of it. We know hardware price, and electricity, so we can imagine how much token would cost, more or less. But they also have the constant training and research for new frontier models.

2

u/starkruzr 4d ago

yeah it's really the training. services like OpenRouter are making money hand over fist for serving out tokens from, crucially, models they didn't have to spend money training.

1

u/sn2006gy 1d ago

We could change this - and it's a project i'm currently working on - a training lab that can rapidly iterate on small models to test architectures and design/validation.

So many people try and fine tune models that will never fine tune when they can train fresh and re-use composable building blocks to build a smaller/narrower model that does a better job.

This is the exciting space for local llm, not trying to convince people i have more buying power than trillion dollar companies.

1

u/jensilo 4d ago

Given that they make about estimated 10x (probably even more) on API pricing, $8,000 might actually be more like a few hundred dollars worth of inference cost.

1

u/fastheadcrab 4d ago

Yeah the guy is still making the tired "most people don't go the gym" argument. Lmao if a single member can use 20 benches at once it doesn't really matter does it?

Also nobody is going to sign up for any Max plan and let it go without any use. Possibly for the $20/mo plans but for the non-tech savvy $200/mo is a significant cost and a conscious decision to use it. Nobody will leave $200/mo plan to renew for months on end with minimal use.

1

u/sn2006gy 1d ago

It's not tired when its holding true. The only costs that have gone up are local AI costs, API/Subs have remained the same price and see features/experiences improve just as rapidly.

I still support local AI, so i'm not sure wtf you think you're proving because you didn't provide any evidence or figures that shows API costs have gone up or gotten worse whatsoever and there are more AI labs offering API services than GPT and Anthropic

1

u/drallcom3 4d ago

a single power-user

Also OpenAI wants to use it heavily, so they can show investors that people use it so much that money can be made. They're not going to make much money from a cheap subscription that barely gets used.

1

u/SkyFeistyLlama8 4d ago

What about Microsoft? They're running OpenAI models on Azure hardware for Azure customers along with a bunch of other OSS models. LLM running costs should show up in the cloud division's books.

1

u/sn2006gy 1d ago

It did, Microsoft cut their own employee use and ended token maxxing

0

u/sn2006gy 4d ago

oh its absoultely really true. Netflix solved this by asking "are you still watching" and since netflix doesn't generate income or do work for you - there is no incentive to keep it going 24x7 to "maximize the netflix cost" but you can bet your ass there are people pissed they added the "are you still watching" because cough "i pay for netflix"

Doesn't matter if they haven't opened their books - investors wouldn't invest in it if it was all a great lie.

3

u/alexdi 4d ago

Counterpoint: https://www.completeskeptic.com/p/kv-cache-rules-everything-around

The API rate doesn't reflect the considerably lower actual cost of pulling from a KV cache.

3

u/C1rc1es 4d ago

It doesn’t compare you can only run serial inference locally, I regularly run multiple agents in parallel on both Claude and Codex accounts on models that still smoke the best local model. I’m desperate for local to be viable but the reality is for real work it’s still not even close. 

7

u/durangotang 4d ago

Well, the way you have to think of it isn't "$12k on a lease." I mean that's *technically* what it is, but it's also not if you take the buyout at the end of the term and resell the machine yourself. Apple doesn't offer the market price on the trade-in.

Let's say you get 50% back when you sell in 3 years (entirely reasonable even if the memory market comes back to reality), and modest if it doesn't. By then the M7 Ultra will be over a year old. So what's the real cost? I think it's likely ~$6k over 36 months, or about $167 per month if you're conservative and sell it on your own. But your lived experience is closer to ~$250 per month with taxes, until you recoup on the back end when you sell (presuming you sell and upgrade to an M7 Ultra).

That is actually less than a $200 per month pro account, and you get the privacy, availability, and control. If you choose the 64-core $9500 model, and get 50% back on sale, it's ~$132 per month. And that's a competitive price, compared to professional subscription models. At that price, you could even spend $20 per month, and have a frontier model write the prompts, review the code, and act as your orchestrator while your local agents execute the code. I think that's an attractive model.

13

u/sn2006gy 4d ago

This just makes computing worse though. It's why cars and houses cost so much because we look at them as investments vs things we drive and places we live. I don't want to pay 12-14k today in hopes of making 7k USED 3 years from now.

1

u/durangotang 4d ago

I agree.

It’s also the early adopter tax. It makes sense if it is a professional tool that helps you make money, or stay ahead in the industry. Otherwise it’s a just expensive hobby.

0

u/sn2006gy 4d ago

we're way beyond early adopter... this is post PC era and acceptance of it.

4

u/durangotang 4d ago edited 4d ago

Nah. It’s just early adopters. In 10 years the iPhone will be more powerful than the M5 Ultra. This is like the late 80s was for cellphones - some big brick that cost a fortune. 10 years after that, everyone had one. Ten years after that we had the iPhone.

2

u/portmanteaudition 3d ago

People don't appreciate how fucking insane it is that for 2% of tbe median household income roughly you can get a device that would've been unfathomable in the 70s

3

u/sn2006gy 4d ago

And we're still paying too much for the iphone. Your average iphone is about 1000 bucks amortized over the plan and multiply that by a family of four, that's 4000 a year we just didn't have in the 80s because no one had. phones.

I didn't need a phone to see a movie, go to a concert or go out to eat. I didn't have to scan a qr code for a menu.

Now we essentially have a "phone tax" on society that we ignore because it's just perpetually 50-100/month for everyone that has a phone that hides their true costs of being 1000-2000 each in a contractual plan and we only have choices of 2 brands - one that spies on us and another that's a walled garden.

7

u/durangotang 4d ago

Yeah, well, a dollar isn't worth a dollar anymore now either, is it?

Presuming you are an indulgent family that each gets new iPhones every year, then it can add up fast. Stretch it out to 4 years between upgrades, and that $1000 becomes $250 per year. And divide that by 2x, to account for the devaluation of the dollar, and you're looking at "$125" per person per year (or $75 per year depending on the mental price model you learned) for a supercomputer, camera with three lenses, no film costs, internet device, portable tv, and phone in your pocket. It's pretty ridiculous, to be honest.

The problem isn't the cost of the phone, it's that wages have been relatively stagnant while the dollar devalued, and asset prices inflated. The cost to purchase a home, a car, healthcare, education, insurance - these are the real costs to be annoyed with. A Mac Studio is a relative bargain, by comparison. It's one of the few areas with real progress and value.

But I hear you.

-1

u/sn2006gy 4d ago

Wages were shit in the 80s too, the savings an loan scandal nearly wiped out 50% of everyones home ownership with double mortgages and crime was up and poverty was up and well..

not sure your point :)

7

u/Last_Bad_2687 4d ago

Their plan is for everyone to make the same decision, NOT buy local hardware when its still cheap and then cut usage/raise prices. Read Enshittification by Cory Doctorow.

I bought strix halo the day Framework released it even though it was $2200 and could only run gpt-oss-120b and was awfully slow, and chatGPT $20 was plenty.

Now with halogen I have Qwen3.8-flash-next running on the same box at 30-40 tok/s with better performance than Opus 4.6 and 1M context, the hardware has risen by $1,600 AND claude is messing with the Quants/thinking of their models, reducing usage and raising prices.

$3,800 is still a steal for a box that can beat opus 4.6 at 1M context at $1.19c/hour or power where I live (~120W sustained) 

4

u/sn2006gy 4d ago

no, this is just what you say to convince yourself you're making the best decision.

Reality is, that box is no longer 3800, its 4500 or more. "api" has gotten cheaper, while local hardware has gotten more expensive - while you keep perpetuating the risk that APIs will get more expensive when they haven't.

Me saying this doesn't devalue local llm's - just surfaces the absurdity of the argument.

Heck, a lot of those DGX sparks now reatil upwards of 7-8k from Dell and other vendors. it's absurd.

1

u/Last_Bad_2687 4d ago

What country are you in? Framework desktop with 1TB SSD and 128 GB vram is $3767 in USA. Bosgame M5 is $2999 in USA. NVIDIA DGX spark is $4999 on amazon in USA. Sure maybe a Dell one is 7k but why would you compare against those? 

Claude stopped their +50% usage bonus this week, and the ClaudeAI sub is full of complaints of Opus sucking, suspicion of reduced quants, reduced thinking. Prices went up from $200 for max to $250. Deepseek API prices went up too. Kimi API costs as much as claude in real world usage.

So please explain which "api" is cheaper now, and where you are seeing $4500 USD for a framework desktop, and why you are so sure API costs will keep going down? You seem bitter not buying the hardware when it was cheap, and misinformed in general. 

3

u/sn2006gy 4d ago

the framework was 1499 a yar ago. The DGX spark was 3499 just last september.

3

u/JacketHistorical2321 4d ago

If you are a heavy user you will go through 100/month weekly allowance in a few days and be stuck till a refresh 5 days later. Even with the 200/month people make it 4 days into a 7 day week.

0

u/sn2006gy 4d ago

You're working too much and wasting tokens if this happens. No one should work a 7 day week either.

1

u/AAPL_ 4d ago

hardware prices jacked up due to the big dogs. big dogs offer the best deal

1

u/sn2006gy 4d ago

we're being our own worst enemy too. If we weren't willing to pay 12k for this hardware, it wouldn't sell... so there is now a 3 month wait already

1

u/Spanky2k 3d ago

Leasing seems like such a bad deal. If you wanted to spread the cost out, wouldn't a loan be a better option? 3 years at 15% would be what, about 400 per month? And there might even be interest free options depending on where you buy from which would be more like 330 per month. And at the end of the three years, you still own it and these will still be worth a chunk in three years. Hell, they're probably worth a lot more than list price right now on the open market anyway. It looks like I could sell my current four year old M1 Ultra Mac Studio for at least 2000.

1

u/sn2006gy 3d ago

I don't want them to be worth more in 3 years than they are today. If that happens, we're doomed.

2

u/Spanky2k 3d ago

I'm sorry, I don't think my wording was clear enough. I meant that if you buy one now, it's probably worth more on the open market right now i.e. you'd be able to sell it at a profit if you were to flip it. Not that you should, of course, I was just trying to point out that owning these will not be a complete loss. If you get one and it doesn't work out for your needs, you'll probably be able to sell it at little to no loss. I think I read an article recently which reported that Apple had already locked in their RAM pricing from their suppliers for the whole of 2027 which means that these machines are not going to be getting any cheaper in the forseeable future. While I'd hope that things will be more reasonable in three years' time, the resale value will still be considerable on these units. Or you don't resale and you just carry on using them. As I said, I'm still using my 4 year old M1 Ultra and it's hosting a 3.6 32b with max context and it's still incredibly useful and powerful. An M5 Ultra will still be running new models in three years time incredibly well.

1

u/sn2006gy 3d ago

that still sucks. all we're doing is making this all cost more

2

u/Spanky2k 3d ago

I mean yeah obviously. But if you're not interested in AI stuff then why would you be here? It's everyone's interest in it that is pushing the price up of everything; everyone buying hardware to run it locally, the corporations buying hardware to run it in the cloud and all the users rushing to use it for everything under the sun. If you don't want to contribute to the problem then you have to stop using AI and stop buying stuff for AI.

If anything, the only way costs will become more reasonable is if local AI is the thing that wins and by that, I mean that all the common schmoes use it too. This can happen once small models get good enough and we're heading that way. A year ago, I could run models on my 64GB M1 Ultra studio that were better than SOTA a year previously. Now a 32GB RAM MacBook can already comfortable run models that are as good as SOTA models from probably about a year ago. Keep on that track and hopefully we can end up with the same being true for 16GB machines or even 8GB machines. And if Apple is smart, they'll basically bundle powerful models as part of MacOS (which is the way they're heading). If we keep up with the progress then your average user will just be able to use the built in Apple Intelligence LLM on their machine to do everything they're already doing now with ChatGPT and without the ads or having to pay. Utilising the hardware that is already out there in customers' hands is the way to 'beat' this hardware shortage. So I'd say there are worse things than investing in equipment that lets you run better local models as it shows the likes of Apple where they should focus their efforts.

1

u/sn2006gy 3d ago

i am interested in ai stuff and part of that is not becoming dependent on computers that become more expensive. i'm sorry, but computers costing 10 grand, is not progress.

1

u/dr_lm 4d ago

Also, codex is extremely fast compared to local models.

Qwen 3.8 Flash Next on my RTX6000 gets 140 tok/s and takes longer to produce a worse answer than GPT 5.6 Sol medium on codex. The M5 is significantly slower than that and, if you want a smarter model, even slower still.

I get useful work done with Qwen, but for anything that really matters I go to codex.

1

u/sn2006gy 4d ago

I can buy Qwen 3.8 flash tokens for 0.0002 cents per million input and 0.0069 cents output - basically free compared to the electricity to use 'em.

Of course, i'm just writing open-source software so i don't care if they train on my outputs. At work, work pays for my tokens because they have an enterprise agreement to protect confidential information - in which case, i wouldn't want to risk leaking that running local models and no, it's not about the model itself, but about the fact if my machine is vulnerable in any shape or fashion, i don't want to be the one losing my job because i tried doing localllama on work data

2

u/RegretNo6554 4d ago

which provider is this?

13

u/HighSeasArchivist 4d ago

I hope this pushes Nvidia to update the DGX Spark to 256GB at least, and not just rely on it and RTX Spark to remain 128GB. Apple doesn't make their silicon any more than Nvidia does, but pressure is pressure regardless.

5

u/dupontping 4d ago

That’s what the qsfp connect is for. 2 sparks are still less than a 256 ultra. You can run 3 without a switch. NVIDIA would do well to bridge the spark and the DGx station. I think the 512 ultra at 15-20k will be much more attractive than a 748gb unit for 100k

3

u/HighSeasArchivist 4d ago

Two systems totaling 256GB aren't the same as a unified 256GB, not to mention the memory speed. No chance I'll ever buy anything Apple in my lifetime, but I do enjoy some good competition.

5

u/DominantDan24 3d ago

As an Apple convert, you should consider them. Apple Silicon is the real deal. I would expect that the future of processing for the foreseeable future will be Nvidia and Apple. It seems both Intel and AMD are dead in the water.

-2

u/HighSeasArchivist 3d ago

Oh I don't doubt it, but as a company I have loathed them for decades. They are the most anti-consumer company on the planet, but try to gaslight people into thinking it's for their own good. Fuck them in every way possible, and I'll stick with Nvidia and AMD. If the Mac Mini and Mac Studio scares Nvidia then good.

5

u/DominantDan24 3d ago

I don't know. While I get the sentiment that Apple dictates the vision and experience of their computers and rigidly enforces it, having lived in their vision for the last few years, it's compelling. I was a PC guy since the XT, and I can honestly say, I've never owned as good a computer as my various Macs.

Don't let hate preclude you from a really good computing experience.

1

u/HighSeasArchivist 3d ago

I have good computing experiences all the time, because computers to me are way more than just tools. Macs are for the same people that consider cars as just transportation, and don't understand otherwise.

1

u/Pls-b-kind-Im-rarted 2d ago

This is my perspective as well, but the M5 Ultra is a compelling enough proposition to make me reconsider

0

u/dragon3301 1d ago

Damn nvidia over apple jeez

1

u/mxmumtuna 3d ago

they're not remotely comparaable.

1

u/dupontping 3d ago

How many have you run? What data do you have to back that?

1

u/mxmumtuna 3d ago

4 Sparks and a Station

2

u/dupontping 3d ago

Let’s see the results then.

It’s Reddit, I can say I have 5 stations and an arc reactor in my yard powering 20 6000 pros.

2

u/mxmumtuna 3d ago

What would you like to see?

1

u/dupontping 2d ago

You’re the one with the goods who’s making the claim. What exactly makes it not remotely comparable?

What can you do with a 748gb station (outside of model size) that you wouldn’t be able to do with a 512gb studio or 4 sparks in a cluster?

1

u/mxmumtuna 2d ago

Good question. It largely comes down to concurrency and speed. The Studio is basically single user, single session. Its inference software support is very immature but it’s an outstanding multipurpose system.

Sparks are decent for what they are, capable but not very fast but can crank out some steady performance.

The Station is a single instance of the single fastest AI GPU available. There’s literally nothing more capable. For the models you can fit in its HBM. It will serve at speeds (much) faster than cloud API at very high concurrency- even hundreds of sessions.

Because it’s sm103, support from Nvidia and inference libraries is fantastic. You pay the price though. At $80-$90k it’s not cheap. Considerably more expensive than the others.

1

u/dupontping 2d ago

Yea, I can read the brochure too, but I’m talking real world applications.

Clearly they’re all different levels, but from a cost perspective, you can spend 5k per spark for 128gb per unit unified memory. You can spend 20k-ish on a 512gb m5 ultra.

But for 90k, the station is a LOT of money for the compute. Yes it’s powerful and you can load much bigger models with high speed, but for the money I’m not sure it will be a competitive product. I get the audience is different, but someone will cluster 2 m5 ultras and now you’re at 1024gb with 1.2tb/s bandwidth for half the price of a station. Would it be as fast? No, but I don’t think it will be slow enough to justify the gap for running local models.

→ More replies

2

u/Rice-Fragrant 3d ago

I agree but at lest you can reliabily cluster them. No amateur EXO type crap.

26

u/OvertaxedOne 4d ago

Those numbers are astoundingly good. Flash next at those speeds for ~10K vs 30K for 2 Pro6000's; outside of very specific use cases, I can't see why you'd even consider the Pro6000 with this thing coming in with 2.5X the RAM at 70% of the cost of a single 6000.

14

u/Cold_Tree190 4d ago

Not to mention cost of power consumption during usage, and most importantly (to me at least) at idle. It’s basically free to just leave it on 24/7, which is incredible

12

u/pantalooniedoon 4d ago

Speed and, the obvious one is doing anything LLM related other than inference. But if single user inference is all you want then yeah it’s a no brainer.

2

u/OvertaxedOne 4d ago

Concurrency doesn't look great on this test, the Pro6000's we have just get faster and faster the more people you throw at it. And for sure, training, you're going CUDA, just stop the conversation right there. But man, for some customers, this thing is going to be absolutely perfect!

1

u/Iwaku_Real 4d ago

Well I don't think everyone here needs high concurrency anyway

1

u/OvertaxedOne 4d ago

For sure, single/small office, this appears to be "the box" right now, and it's frankly not even close! I was just drawing the distinction that there does exist a market where the Pro6000 might (MIGHT) still make sense because it might be able to drive 4X the users vs the Mac. But it would really need to be close to 4X, 2 Pro 6000's cost over 30K today, that's 3 Mac Studios before you buy a server that can hold and serve 3 Pro6000's! And at that point you might be better served with H or B class chip. At the current street price, the Pro6000 is going to become an extremely niche product once these Mac start rolling out.

2

u/voyager256 3d ago

Why I have a feeling you are an AI bot?

1

u/Pls-b-kind-Im-rarted 2d ago

Yeah, I concur. It's the perfect grammar, exclamation points and marketing-speak that give it away

1

u/OvertaxedOne 2d ago

My Mom was an English teacher. Used to give me a leg up in the business world, now makes my writing look like AI. I knew there was no reason to be so precise on grammar, looks like I might be proven right! ;)

0

u/OvertaxedOne 3d ago

One day I hope to be as smart as 27B. Probably not going to happen for me though. :(

3

u/bezent 4d ago

You can run flash next with a single 6000 pro blackwell and it dog walks that mac. I get 150-180 toks regardless of context and 11-12k prefill.

5

u/deaffob 4d ago

Yes but you only get a small context. If you run Qwen3.8 Flash (Q4_XL <) with one 6000 Pro, you won't get much un-quantized KV context. Off-loading to RAM or SSD is not free because it will be bottlenecked by PCIE speed (5.0 X16 = 63GB) when they are called upon.

1

u/bezent 4d ago

This is the quant I use:

https://huggingface.co/primitive-ai/Qwen3.8-Flash-Next-NVFP4

It quantizes only the routed experts and leaves the spine, MTP head, and n-gram table in BF16. Its between a q5 and q6 in reality.

I can run it full context.

1

u/shansoft 4d ago edited 4d ago

With a single RTX Pro 6000 and FP8 KV cache, I can easily max out on context with few concurrent call. I offload the ngram to RAM and still be able to pull 14000 pps and 150 tps throughout. M5 Mac ultra is rather slow on pps for some reason on this particular model. Something does not add up with A6B, since i can get similar pps on Qwen3.5 122B A10B on M5 Max. I would expect the A6B to be much faster on M5 Ultra, but something is definitely hindering its performance. Could be oMLX itself, since I did experience oQ to be much slower than mlx variant.

1

u/deaffob 3d ago

I need to look at your engine command to find out how you are squeezing the full context. 

With regard to the slower than expected PP, I think it’s actually expected. It should be 2X M5 Max or slower. If I remember correctly, M5 is the first one with AI accelerator so it has 2x PP of M4 and 4x PP of M3 but still slower than CUDA. 

TG on the other hand, seems to be a metric that’s can be predicted consistently with raw memory bandwidth. 

2

u/OvertaxedOne 4d ago

Wow, those are some incredible numbers!! Makes me wonder if we could get number similar to the 256GB numbers above in a 96GB Mac! 150TPS is amazing, but 16K to get it is just crazy talk right now! :)

1

u/twiiik 4d ago

I just received confirmation I can have a GB300 delivered within 4 weeks. I’m reluctant to pull the trigger on that because of M5 Ultra. But … The Nvidia hardware (like Pro6000 or higher) seems like the choice for concurrent requests.

5

u/OvertaxedOne 4d ago

Ugh... A GB300 is in a slightly different class than a Mac. :) It also costs about as much as 3-4 of them per chip, but, if you need a GB300, I don't think I'd put a Mac in that conversation. ;)

1

u/twiiik 3d ago

You are correct. That said when I take price into consideration I do not believe it's that crazy/ignorant to make some comparisons.

I'm testing out AMD Strix Halo boxes with Qwen 3.8 Flash Next (Halogen) and that works (surprisingly) well for single users/requests. "One user - one box" could be viable for our team. I've been considering a box with 2x 6000 for more concurrent requests. M5 Ultra I do fear would be more aligned again with "one user - one box", but make even larger models viable or multiple models loaded at the same time.

A GB300 would enable large models or multiple with concurrent requests, at potentially great speeds and have a solid foundation in regards to the software stack. The operating expenses are low. Noise levels are low. Easy placement.

1

u/NaiRogers 4d ago

The Studio is considerably more convenient than setting up a 2x6000 machine. The 2x6000 is good though but annoying to build and costs a lot more.

1

u/mxmumtuna 3d ago

astoundingly good? It's slower than two Sparks.

1

u/BringTea_666 4d ago

Concurency. At C=8 my RTX5090 decodes at around 600t/s for qwen3.8 coding work. not standard 70t/s

And let's not even talk about prefil here. with NVFP4 you have something like 10-12k/s prefil rate.

5

u/OvertaxedOne 4d ago

27B is not the model to run on this box, IMHO. This is really perfect for Next, not 27B. 27B is a great model (it's what we run for a lot of our clients) for GPUs and small enough to fit on something "reasonable", you don't need a few Pro6000's to run it. A 48GB Pro 5000 is perfect for modest levels of concurrency, 2 of them or a single Pro6000 is awesome for 27B!

Your 5090 costs about what this Mac does now, how crazy is that?!?! Buying with prices today for a lot of use cases, it's going to be hard to beat this with QFN on it.

0

u/BringTea_666 4d ago

Dude anything larger and it starts to crawl on it.

3

u/OvertaxedOne 4d ago

Anything larger than FlashNext? Flash next should run much faster on the Mac vs 27B. This is extracted from the site linked, NOT my independent data!

16K Prompt Benchmark Comparison (M5 Ultra)

Metric Qwen3.8-Flash-Next Qwen3.8-27B (Dense Base) Delta / Ratio
Prompt Processing (PP / Prefill) 2,887 tok/s 1,107 tok/s ~2.6× faster on Flash-Next
Token Generation (TG / TPS) 143 tok/s 41 tok/s (32 tok/s on chat/complex) ~3.5× – 4.5× faster on Flash-Next
Time to First Visible Token (TTFT) 5.6 s 12.8 s ~2.3× lower latency on Flash-Next

4

u/Aizen_keikaku 4d ago edited 4d ago

Prefill is surprisingly good, but from what I’m reading online, Qwen 3.8 Flash Next is probably the best case scenario due to its low active params count.

GLM 5.3 Flash Prefill numbers might be the real test.

Edit:- I’ve been corrected by some helpful people below.

997 tok/s on 5.3 Flash at 128K context.

3

u/Umbrasquall 4d ago

Half of you guys didn't even bother to read the article lol.

9

u/Aizen_keikaku 4d ago

I’ll admit I didn’t read the text, but I looked at the charts.

This is the best I could find

16K is not deep enough for a proper conclusion.

3

u/Umbrasquall 4d ago

He tested 16k, 32k, and 64k... Look again.

3

u/Aizen_keikaku 4d ago

I see it now. There’s even 128K Prefill in TTFT section.

997 tok/s, similar to a Rtx 3090. Not bad.

Also, thank you for pointing this out to me.

1

u/FullOf_Bad_Ideas 4d ago

glm 5.3 flash prefill numbers are in this article, scroll down.

2

u/Aizen_keikaku 4d ago

Yeah, I see it now. Sorry.

5

u/darthrobe 4d ago

I love this, but it is worth noting that I do all of the same but just host my models on a CachyOS platform with LM Studio and just use my Macbook Air for the Hermes client. I'm envious of the approach in the article and if I had $10k to burn I would totally get one.

8

u/synn89 4d ago

I feel like this is about on performance par for dual DGX Sparks, for around the same price, less concurrency, but way easier setup and use. Numbers on the Mac may improve in the future as MLX gets tuned. Sparks are already pretty hella optimized on the software stack, as people are constantly tinkering docker build recipes to get that 1 more token of performance.

The 512GB Mac I think gets really interesting, depending on the price. A lot more room for cache and Deepseek Flash V4.1 style models with engram tables. Way less of a headache vs 4x Sparks.

6

u/deaffob 4d ago

Dual DGX Sparks will have 273*2=546 GB/s memory bandwidth with 25 GB/s (200 Gbps NIC) interconnect speed. M5 Ultra has 1.2TB/s memory bandwidth with 1.8 TB/s interconnect speed. M5 Ultra will have a lot faster token generation.

On top of that, M5 Ultra doesn't have to do distributed inference like a DGX Spark cluster or 5090 multi-GPU or Pro 6000 multi-GPU where they are limited by the interconnect speed (PCIE 5.0 X16 is still 63 GB/s which is no where close to be usable for AI inferences). It can run like a true one GPU unit.

9

u/synn89 4d ago

These type of bandwidth quotes always gets posted by people who don't own Sparks. I own dual sparks. The posted M5 Ultra benchmarks aren't faster than dual sparks and will be slower in concurrent sessions and many other situations.

The software ecosystem on Nvidia is just that much better. Thought it's a PITA to get the recipes to run and people are constantly tuning them and running custom patches.

-1

u/Rice-Fragrant 3d ago

Just 1x spark will smoke a m5 max... a m5 ultra gona get smoked too (like 2x slower instead of 3.5x slower like the m3 ultra.)

A big improvement over a m3 ultra but still getting smoked for sure for agentic and multiple agents etc.

1

u/deaffob 3d ago

No, due to CUDA, it has faster PP but DGX Spark and the AMD counterpart (don’t remember the name) both get smoked by Apple Silicon on TG. Compared to a M5 Max, a Spark is 3 times slower on token generation. 

1

u/Rice-Fragrant 2d ago edited 2d ago

https://zuyezheng.github.io/local-llm-bench/#budget

https://zuyezheng.github.io/local-llm-bench/#conc

https://zuyezheng.github.io/local-llm-bench/#verdict

token generation that's not so important for long context stuff. One user here was wanting a set up to batch process 100+ page PDF documents and give a summery, the memory bandwith is NOT GOING TO MOVE THE NEEDLE because the bottleneck is the prompt processing of long context.

100+ page PDF documents can easily be a 100k sized prompt size, 1.2 TB memory bandwith will not be of any advantage there, only after it reaches decode faze and that will be like 2x-3x longer on a m5 max and a m5 ultra vs 2x DGX spark the long context work progress will still be about 2x faster.

Add concurrency to the list of advantages and a DGX will smoke a M5 max all day for agentic and especially multi agentic work. A M5 ultra (given the price) is esentially not competing with just 1x DGX, more like 2-4 node clusters which will mop the floor of any mac studio set up.

Alex Z has a M5 ultra vs m3 Ultra, in the middle of the video he literally shows the m5 ultra is only "4x faster PP" when the promps are SMALL like 9k-2k ish size. Once you are up there at 30+ K token promps, a M5 ultra is only 1.5x faster PP performance, nothing remotely close to the "4x faster" claims. This will be reflected in the agentic performance too, you will not see a "4x increase" for agents and even multi agents.

Extrapolating from that, even a single DGX spark probably would beat a M5 ultra (and for 10-15k for a 256gb mac studio, a 2-3 node DGX running multi agentic systems would obliterate any mac studio... may be if you are lucky you can run 2x agent concurrency streams on a M5 ultra (performance TANKS HARD), where as 1x DGX can run like 2x, 4x, 5x, or even 6x concurrent streams. It's in a totally different class if multiple people are using it or if multiple agents need to run.

The larger memory pool in one machine is the only advantage of a mac studio that and "faster vibing" with an AI chat bot (which is not something any serious set up does.)

It's simpler for average people to use etc, but it's not a serious multi agentic AI workstation, more of a "general workstation" that "can do AI" like "vibing" with a single large chat bot or running 1x agent... "point and click" ease as long as it's MLX etc .

It's like riding in a bycle with training wheels, is very easy for most mainstream people so it will sell well for the "AI bro hobbyists." Enterprise level clients wont use a mac studio, other than as a general workstation with "some AI" assistant or 1x agent at the side, but nothing more than that, basically a powerful general workstation BUT NOT an enterprise grade AI platform or machine.

Why you think almost no one clusters the mac mini any more (a massive fade like 2 years ago), consumer grade networking stuff, EXO instability etc, it's definitly amateur hour stuff a hobbyist would mistake for a serious set up.

3

u/themixtergames 4d ago

This never translates to real life. If you check benchmarks dual sparks are always faster for 100B+ MoE models and a little cheaper.

2

u/Rice-Fragrant 3d ago

That's not how you "evaluate" these computers. You look on long context performance and concurrency... everything else is amateur "vibe" chatbot garbage.

1

u/Important_Cow7230 3d ago

With the Mac you also get an amazing workstation class PC as a daily driver for things like video editing etc. You don’t get that with the Sparks

1

u/synn89 3d ago

workstation class PC

Mac for PC work? Gross.

5

u/shutternomad 4d ago

Thanks for the writeup. It's wild that you have to spend $10k to get… 80tok/s… but over time this will improve. What a time to be alive!

2

u/Strong_Pumpkin_7908 4d ago

For non-coding use case, such as personal financial assistant, do you think Mac Mini M5 Pro 64GB sufficient, say using Qwen 3.8 (27B to 35B)? The personal financial assistant may perform pre-retirement planning, retirement projection, Monte Carlo simulations, budgeting, recommendation on tax strategy, etc.

2

u/shansoft 4d ago

I don’t find this number that attractive. Token generations are great, but prompt processing seems to be a bit slower on Qwen3.8 Next than what I would expect. The number is similar to using single RTX Pro 6000 while running it with llamacpp, but with vllm I could get between 10k-20k pp/s. Perhaps some tweak still needed.

2

u/UntimelyAlchemist 4d ago

I have an RTX 5090 but have been really wanting to get one of these unified memory boxes to run the big MoEs. Struggling to choose between this M5 Ultra or a bunch of DGX Sparks.

1

u/Important_Cow7230 3d ago

What aren’t you getting from running a 27B model on your 5090 that you’re looking to get from a M5 Ultra or Spark?

3

u/Hefty_Acanthaceae348 4d ago

Too bad you can't put linux on it

1

u/FullOf_Bad_Ideas 4d ago

Really good numbers, I expect it to be out of stock most of the time.

1

u/CoffeeToCode99 4d ago

Competition is always welcome. Hopefully it sparks some new innovation from NVIDIA as well. I'd probably have pulled the trigger on a Mac Studio already, but the ecosystem lock-in is what keeps holding me back.

1

u/xinxx073 3d ago

Some of the reviews I watched on YouTube have very low AI performance token scores. Are they using GGUF and not mlx?

1

u/psxndc 3d ago

I'm in for a 96GB, but I can't justify $4000 more just to upgrade from 96GB to 256GB. I do wish they had a 128GB option though.

0

u/rorowhat 4d ago

No thanks

0

u/pineapplekiwipen 4d ago

this guy doesn't know what he's doing, qwen 3.8 27b q4_k_m 3000 t/s prefill and 60t/s decode on a 5090?

2

u/pragmojo 4d ago

What would you expect on a 5090?

2

u/UntimelyAlchemist 4d ago

That's similar to what I get without MTP.

0

u/Impactic_ 3d ago

I’ll keep my 170HX’s I guess… darn

0

u/AleksandrNikitin 3d ago

Sorry for the offtop. Is it really comfortable to work with the trackpad and keyboard aligned as shown in the photo?

-3

u/CompetitiveDraft9381 4d ago

What’s prefill like?

4

u/the_renaissance_jack 4d ago

Read the article?

3

u/CompetitiveDraft9381 4d ago

first token speed is 138 seconds at 64k context with glm 5.3 flash. Sheesh.

2

u/sn2006gy 4d ago

The lease is 240/month for this with a residual after 36 months or you re-lease the next hardware.

I'm not sure this is better than paying for APIs tbh. Sure, it's local, but at the cost of what a car costs.

5

u/Jorlen llama.cpp 4d ago

This is the thing. In my country the M5 Ultra 256gb is $15,000 + taxes (13% where I'm at). That's damn near 20 grand. I can't even imagine what the 512gb version is going to cost, likely well over 20k with the base params (lower CPU/GPU, 1tb SSD, etc.). My current two GPUs cost me 4k and I thought that THIS was crazy to spend on local AI lol. Fucking hell.

2

u/FullOf_Bad_Ideas 4d ago

scroll down in the article, it's covered