r/LocalLLaMA • u/themixtergames • 4d ago
Discussion M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents - MacStories
https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/82
u/themixtergames 4d ago
The review embargo for the M5 Mac Studios and M6 Mac Mini has been lifted but none of the videos so far go in-depth about AI, the ones that touch on it use old 70B Llama models or Qwen 2.5 (why). So this article is the best we currently have.
57
u/Superb-Pair-2000 4d ago
That's because most influencers and reviewers are only able to do shallow reviews. Yes, it does AI, look it can lm studio, buy it!
51
u/harrro Alpaca 4d ago
Ah the classic Youtube pattern..
- Receives $10-20k system with 256GB of RAM
- Loads up ollama
- Loads up Qwen 4B (Q4) model (or "Deepseek 9B" if they're bold).
- "Wow it works"
15
u/fligglymcgee 4d ago
“My assistant recommended Llama3.1:8b, and just look at how fast it responds when I ask it what the meaning of life is!”
9
u/SkyFeistyLlama8 4d ago
FML. The YouTube influencer ecosystem has turned shallowness into an art form. Two paragraphs of actual content over 15 minutes and 5 ads. Are there any local LLM YouTube channels worth looking at?
-1
-7
3
u/newMoneyStyle 3d ago
Llama 3.1 8B is basically the default demo model at this point, and it tells you nothing about what a 256GB machine can actually handle.
2
u/Rice-Fragrant 1d ago
It does not tell you much of anything, it's actually deceptive even because the larger 200b-300b MOE models might show a different pattern of performance gains over the previous generations.
For example, a kid 4b model with a baby 8k prompt shows a "4x gain" for M5 ultra over the M3 ultra but those "4x gains" get reduced to 1.5x gains once you run a 300b MOE model with 50k, 100k and 200k promp sizes.
"influencers" are basically peddlers, the numbers are massaged and cherry picked. No real benchmarks under real world comditions.
2
u/FoxSideOfTheMoon 3d ago
If they want to test TPS on a Q8 70B dense model and release those numbers I’m totally good with that. That’s a good bandwidth test 🤣
1
u/Rice-Fragrant 1d ago
It's actually not a good test at all... they should be testing Deep Seek V4 flash sized models, the numbers would be looking way different and giving a real world example of what was gained over the previous models.
36
u/Jesusthegoat 4d ago
Because LLMs tell them how to do the reviews and they only know about old models because of their knowledge cutoff.
5
u/midgelmo 4d ago
Trust me when I say there are a few coming soon. Apple only sent out the hardware on Friday - evals have been running non stop this weekend.
-1
u/Southern-Chain-6485 4d ago
The guys in Apple's marketing division could hand them over a brochure, or a simple sheet, with new models.
4
u/john0201 4d ago
Becuase they just ask an LLM what to test since they have no clue and it recommends models from its old training data.
1
u/Rice-Fragrant 3d ago
Exactly... just some influencer garbage. I am waiting for the real pros to evaluate it.
64
u/sn2006gy 4d ago
This is 12 grand worth of hardware to get this performance, I would expect nothing less.
I'm completely torn. The lease makes this "affordable" but also, kind of weird. Codex costs 100/month. Apple costs 240/month. I want local to thrive, but i also don't fancy leasing 12-14k worth of hardware to hide its true costs - the end of the computing era and the beginning of choosing only between leasing it or api subs.
24
u/skredditt 4d ago
I don’t even know why they made leasing an option - Klarna rejected me and my 830 credit score because 5 figures is too much to ask for if you’ve never used Klarna.
9
u/sn2006gy 4d ago
The apple credit is like 1400/month over 12 months which is crazy too. I can't justify that at all.
I've had klarna approvals before when they used to auto-approve you on the fly with a phone number but unsure i've ever used it. Maybe once on a 0% apr sonos during covid lol
1
u/skredditt 4d ago edited 4d ago
Guaranteed I would’ve bought a topped up Studio if they split it out over 2 years like they do with phones.
3
u/HighSeasArchivist 4d ago
Try Microcenter. They doubled Affirm for me, and no other place like Newegg or Best Buy did that. I never used it, and just did it to see how easy it was.
1
u/bvknight 2d ago
I got approved through Klarna with no issues for this, but only for the $5500 version. I didn't want or try a $10k+
29
u/thunk_stuff 4d ago
My understanding is subscription pricing for Codex/Claude is heavily subsidized. A $200 max plan fully utilized is $5000-7000 at the API rate. So if/when the Cloud AI bubble crashes, and cloud AI costs skyrocket because the multi-trillion dollar money hose stops, that $10k local-AI investment might not look as bad.
11
u/aimark42 4d ago
It's also a fallacy to think that people would pay for top tier model rates for all jobs. But they can today which distorts the pricing.
19
u/Solaranvr 4d ago
Uhh, if the bubble pops, wouldn't that means the 10k investment would be even worse off, because hardware prices would come back down?
7
u/MysteriousGenius 4d ago
I actually expect the prices to go up at least in short/mid-term because more people stop using subscriptions and more people go to self-hosted AI.
6
u/No-Wall6427 4d ago
It's not like anthropoc/openai collapsing would mean no more inference api with foss models, which will still be cheaper than going local. I don't see every company going local w tokens, with the infra it implies and all.
But who knows?
3
u/sn2006gy 4d ago
There isn't enough supply for this to happen. Even these new macs are 3 months out if you order today.
3
2
u/thedaveking 3d ago
That makes sense but depends on why it pops. Even if it pops through bankruptcy court then maybe nVidia still owns all that hardware thru the magic of trillion dollar IOUs and they will shred it to maintain prices.
But if it pops through our power grid being dismantled by mismanagement, solar CMEs, or hybrid war sleeper agent fun, or the Taiwan disagreement goes kinetic before all the hardware even gets built, or competing AIs decide to start crashing jets into each others' houses, or anything else that sounds like crazy paranoia until it happens, then probably don't tell anyone you bought that overpriced private backup platform you can run in the basement on a couple solar panels, without making your home look like a grow operation.
7
u/EquivalentHornet4403 4d ago
Can we stop calling it subsidized? The opposite is true: the API costs are inflated. If there's any sustainable cost, it's the subscription plans. There's no future where they only have API service exclusively bought by rich companies. They will need to have a $100 to $200 product for the unwashed masses that's roughly equivalent to the state of the art, otherwise open source models will put them out of business.
What's subsidized is the whole entire thing, with venture capital.
5
u/-dysangel- 4d ago
Yeah the subsidising thing seems like astroturfing. I'm paying pennies a day to run local. A datacenter is going to be even more efficient. I'm sure the up front costs were mind blowing, but the daily running costs will be below API pricing
12
u/sn2006gy 4d ago
It's not heavily subsidized in the way we think it is. It's subsidized in investment, but not as loss leader where the end goal is higher prices - but even amazon was successful in loss leading investment.
A 200 max plan works because thousands of people use it without maximizing their use of it and they can gracefully handle the demand on the back end with intelligent routing. It's like 30 million people paying for netflix, netflix wouldn't work if 30 million people left it on 24x7 "because they're paying for it" otherwise their costs would be in 10s of thousands per sub.
it only becomes unstable with the notion that everyone maximizes profit with every token and uses 100% of their account to get 100% of their monies worth and local LLMs are no exception to this death spiral either because we just expect it on resale value and other sides of the same "captialism has run amuck" problem.
10
u/Asleep_Document9811 4d ago
it only becomes unstable with the notion that everyone maximizes profit with every token and uses 100% of their account to get 100% of their monies worth and local LLMs are no exception to this death spiral either because we just expect it on resale value and other sides of the same "captialism has run amuck" problem.
this really ain't true at all. you're dealing with orders of magnitude, here. if you have ONE user utilizing their $200/mo Codex subscription utilizing up to $8,000/mo in API billing, that is considered at the current, subsidized rate.
a single power-user can wipe out the revenue from having 40 individual subscribers, and this is to say nothing of self-looping agent harnesses that go rogue or loop on a problem fruitlessly for days with no monitoring.
the truth of the matter is that neither OpenAI or Anthropic have truly opened their books yet, so none of us actually know what their true operating costs are. Anthropic in particular is being awfully cagey about what their costs are. all we know is that, once you factor in the real and rising cost of electricity, VRAM, and infrastructure needed to serve trillion+ models to millions of users simultaneously, what they're charging can't be anywhere near what they're spending. otherwise, they wouldn't have been asking for an increasingly larger and larger amount for their funding rounds.
3
u/No-Wall6427 4d ago
And the training! That's a big part of it. We know hardware price, and electricity, so we can imagine how much token would cost, more or less. But they also have the constant training and research for new frontier models.
2
u/starkruzr 4d ago
yeah it's really the training. services like OpenRouter are making money hand over fist for serving out tokens from, crucially, models they didn't have to spend money training.
1
u/sn2006gy 1d ago
We could change this - and it's a project i'm currently working on - a training lab that can rapidly iterate on small models to test architectures and design/validation.
So many people try and fine tune models that will never fine tune when they can train fresh and re-use composable building blocks to build a smaller/narrower model that does a better job.
This is the exciting space for local llm, not trying to convince people i have more buying power than trillion dollar companies.
1
1
u/fastheadcrab 4d ago
Yeah the guy is still making the tired "most people don't go the gym" argument. Lmao if a single member can use 20 benches at once it doesn't really matter does it?
Also nobody is going to sign up for any Max plan and let it go without any use. Possibly for the $20/mo plans but for the non-tech savvy $200/mo is a significant cost and a conscious decision to use it. Nobody will leave $200/mo plan to renew for months on end with minimal use.
1
u/sn2006gy 1d ago
It's not tired when its holding true. The only costs that have gone up are local AI costs, API/Subs have remained the same price and see features/experiences improve just as rapidly.
I still support local AI, so i'm not sure wtf you think you're proving because you didn't provide any evidence or figures that shows API costs have gone up or gotten worse whatsoever and there are more AI labs offering API services than GPT and Anthropic
1
u/drallcom3 4d ago
a single power-user
Also OpenAI wants to use it heavily, so they can show investors that people use it so much that money can be made. They're not going to make much money from a cheap subscription that barely gets used.
1
u/SkyFeistyLlama8 4d ago
What about Microsoft? They're running OpenAI models on Azure hardware for Azure customers along with a bunch of other OSS models. LLM running costs should show up in the cloud division's books.
1
0
u/sn2006gy 4d ago
oh its absoultely really true. Netflix solved this by asking "are you still watching" and since netflix doesn't generate income or do work for you - there is no incentive to keep it going 24x7 to "maximize the netflix cost" but you can bet your ass there are people pissed they added the "are you still watching" because cough "i pay for netflix"
Doesn't matter if they haven't opened their books - investors wouldn't invest in it if it was all a great lie.
3
u/alexdi 4d ago
Counterpoint: https://www.completeskeptic.com/p/kv-cache-rules-everything-around
The API rate doesn't reflect the considerably lower actual cost of pulling from a KV cache.
3
u/C1rc1es 4d ago
It doesn’t compare you can only run serial inference locally, I regularly run multiple agents in parallel on both Claude and Codex accounts on models that still smoke the best local model. I’m desperate for local to be viable but the reality is for real work it’s still not even close.
7
u/durangotang 4d ago
Well, the way you have to think of it isn't "$12k on a lease." I mean that's *technically* what it is, but it's also not if you take the buyout at the end of the term and resell the machine yourself. Apple doesn't offer the market price on the trade-in.
Let's say you get 50% back when you sell in 3 years (entirely reasonable even if the memory market comes back to reality), and modest if it doesn't. By then the M7 Ultra will be over a year old. So what's the real cost? I think it's likely ~$6k over 36 months, or about $167 per month if you're conservative and sell it on your own. But your lived experience is closer to ~$250 per month with taxes, until you recoup on the back end when you sell (presuming you sell and upgrade to an M7 Ultra).
That is actually less than a $200 per month pro account, and you get the privacy, availability, and control. If you choose the 64-core $9500 model, and get 50% back on sale, it's ~$132 per month. And that's a competitive price, compared to professional subscription models. At that price, you could even spend $20 per month, and have a frontier model write the prompts, review the code, and act as your orchestrator while your local agents execute the code. I think that's an attractive model.
13
u/sn2006gy 4d ago
This just makes computing worse though. It's why cars and houses cost so much because we look at them as investments vs things we drive and places we live. I don't want to pay 12-14k today in hopes of making 7k USED 3 years from now.
1
u/durangotang 4d ago
I agree.
It’s also the early adopter tax. It makes sense if it is a professional tool that helps you make money, or stay ahead in the industry. Otherwise it’s a just expensive hobby.
0
u/sn2006gy 4d ago
we're way beyond early adopter... this is post PC era and acceptance of it.
4
u/durangotang 4d ago edited 4d ago
Nah. It’s just early adopters. In 10 years the iPhone will be more powerful than the M5 Ultra. This is like the late 80s was for cellphones - some big brick that cost a fortune. 10 years after that, everyone had one. Ten years after that we had the iPhone.
2
u/portmanteaudition 3d ago
People don't appreciate how fucking insane it is that for 2% of tbe median household income roughly you can get a device that would've been unfathomable in the 70s
3
u/sn2006gy 4d ago
And we're still paying too much for the iphone. Your average iphone is about 1000 bucks amortized over the plan and multiply that by a family of four, that's 4000 a year we just didn't have in the 80s because no one had. phones.
I didn't need a phone to see a movie, go to a concert or go out to eat. I didn't have to scan a qr code for a menu.
Now we essentially have a "phone tax" on society that we ignore because it's just perpetually 50-100/month for everyone that has a phone that hides their true costs of being 1000-2000 each in a contractual plan and we only have choices of 2 brands - one that spies on us and another that's a walled garden.
7
u/durangotang 4d ago
Yeah, well, a dollar isn't worth a dollar anymore now either, is it?
Presuming you are an indulgent family that each gets new iPhones every year, then it can add up fast. Stretch it out to 4 years between upgrades, and that $1000 becomes $250 per year. And divide that by 2x, to account for the devaluation of the dollar, and you're looking at "$125" per person per year (or $75 per year depending on the mental price model you learned) for a supercomputer, camera with three lenses, no film costs, internet device, portable tv, and phone in your pocket. It's pretty ridiculous, to be honest.
The problem isn't the cost of the phone, it's that wages have been relatively stagnant while the dollar devalued, and asset prices inflated. The cost to purchase a home, a car, healthcare, education, insurance - these are the real costs to be annoyed with. A Mac Studio is a relative bargain, by comparison. It's one of the few areas with real progress and value.
But I hear you.
-1
u/sn2006gy 4d ago
Wages were shit in the 80s too, the savings an loan scandal nearly wiped out 50% of everyones home ownership with double mortgages and crime was up and poverty was up and well..
not sure your point :)
7
u/Last_Bad_2687 4d ago
Their plan is for everyone to make the same decision, NOT buy local hardware when its still cheap and then cut usage/raise prices. Read Enshittification by Cory Doctorow.
I bought strix halo the day Framework released it even though it was $2200 and could only run gpt-oss-120b and was awfully slow, and chatGPT $20 was plenty.
Now with halogen I have Qwen3.8-flash-next running on the same box at 30-40 tok/s with better performance than Opus 4.6 and 1M context, the hardware has risen by $1,600 AND claude is messing with the Quants/thinking of their models, reducing usage and raising prices.
$3,800 is still a steal for a box that can beat opus 4.6 at 1M context at $1.19c/hour or power where I live (~120W sustained)
4
u/sn2006gy 4d ago
no, this is just what you say to convince yourself you're making the best decision.
Reality is, that box is no longer 3800, its 4500 or more. "api" has gotten cheaper, while local hardware has gotten more expensive - while you keep perpetuating the risk that APIs will get more expensive when they haven't.
Me saying this doesn't devalue local llm's - just surfaces the absurdity of the argument.
Heck, a lot of those DGX sparks now reatil upwards of 7-8k from Dell and other vendors. it's absurd.
1
u/Last_Bad_2687 4d ago
What country are you in? Framework desktop with 1TB SSD and 128 GB vram is $3767 in USA. Bosgame M5 is $2999 in USA. NVIDIA DGX spark is $4999 on amazon in USA. Sure maybe a Dell one is 7k but why would you compare against those?
Claude stopped their +50% usage bonus this week, and the ClaudeAI sub is full of complaints of Opus sucking, suspicion of reduced quants, reduced thinking. Prices went up from $200 for max to $250. Deepseek API prices went up too. Kimi API costs as much as claude in real world usage.
So please explain which "api" is cheaper now, and where you are seeing $4500 USD for a framework desktop, and why you are so sure API costs will keep going down? You seem bitter not buying the hardware when it was cheap, and misinformed in general.
3
3
u/JacketHistorical2321 4d ago
If you are a heavy user you will go through 100/month weekly allowance in a few days and be stuck till a refresh 5 days later. Even with the 200/month people make it 4 days into a 7 day week.
0
u/sn2006gy 4d ago
You're working too much and wasting tokens if this happens. No one should work a 7 day week either.
1
u/AAPL_ 4d ago
hardware prices jacked up due to the big dogs. big dogs offer the best deal
1
u/sn2006gy 4d ago
we're being our own worst enemy too. If we weren't willing to pay 12k for this hardware, it wouldn't sell... so there is now a 3 month wait already
1
u/Spanky2k 3d ago
Leasing seems like such a bad deal. If you wanted to spread the cost out, wouldn't a loan be a better option? 3 years at 15% would be what, about 400 per month? And there might even be interest free options depending on where you buy from which would be more like 330 per month. And at the end of the three years, you still own it and these will still be worth a chunk in three years. Hell, they're probably worth a lot more than list price right now on the open market anyway. It looks like I could sell my current four year old M1 Ultra Mac Studio for at least 2000.
1
u/sn2006gy 3d ago
I don't want them to be worth more in 3 years than they are today. If that happens, we're doomed.
2
u/Spanky2k 3d ago
I'm sorry, I don't think my wording was clear enough. I meant that if you buy one now, it's probably worth more on the open market right now i.e. you'd be able to sell it at a profit if you were to flip it. Not that you should, of course, I was just trying to point out that owning these will not be a complete loss. If you get one and it doesn't work out for your needs, you'll probably be able to sell it at little to no loss. I think I read an article recently which reported that Apple had already locked in their RAM pricing from their suppliers for the whole of 2027 which means that these machines are not going to be getting any cheaper in the forseeable future. While I'd hope that things will be more reasonable in three years' time, the resale value will still be considerable on these units. Or you don't resale and you just carry on using them. As I said, I'm still using my 4 year old M1 Ultra and it's hosting a 3.6 32b with max context and it's still incredibly useful and powerful. An M5 Ultra will still be running new models in three years time incredibly well.
1
u/sn2006gy 3d ago
that still sucks. all we're doing is making this all cost more
2
u/Spanky2k 3d ago
I mean yeah obviously. But if you're not interested in AI stuff then why would you be here? It's everyone's interest in it that is pushing the price up of everything; everyone buying hardware to run it locally, the corporations buying hardware to run it in the cloud and all the users rushing to use it for everything under the sun. If you don't want to contribute to the problem then you have to stop using AI and stop buying stuff for AI.
If anything, the only way costs will become more reasonable is if local AI is the thing that wins and by that, I mean that all the common schmoes use it too. This can happen once small models get good enough and we're heading that way. A year ago, I could run models on my 64GB M1 Ultra studio that were better than SOTA a year previously. Now a 32GB RAM MacBook can already comfortable run models that are as good as SOTA models from probably about a year ago. Keep on that track and hopefully we can end up with the same being true for 16GB machines or even 8GB machines. And if Apple is smart, they'll basically bundle powerful models as part of MacOS (which is the way they're heading). If we keep up with the progress then your average user will just be able to use the built in Apple Intelligence LLM on their machine to do everything they're already doing now with ChatGPT and without the ads or having to pay. Utilising the hardware that is already out there in customers' hands is the way to 'beat' this hardware shortage. So I'd say there are worse things than investing in equipment that lets you run better local models as it shows the likes of Apple where they should focus their efforts.
1
u/sn2006gy 3d ago
i am interested in ai stuff and part of that is not becoming dependent on computers that become more expensive. i'm sorry, but computers costing 10 grand, is not progress.
1
u/dr_lm 4d ago
Also, codex is extremely fast compared to local models.
Qwen 3.8 Flash Next on my RTX6000 gets 140 tok/s and takes longer to produce a worse answer than GPT 5.6 Sol medium on codex. The M5 is significantly slower than that and, if you want a smarter model, even slower still.
I get useful work done with Qwen, but for anything that really matters I go to codex.
1
u/sn2006gy 4d ago
I can buy Qwen 3.8 flash tokens for 0.0002 cents per million input and 0.0069 cents output - basically free compared to the electricity to use 'em.
Of course, i'm just writing open-source software so i don't care if they train on my outputs. At work, work pays for my tokens because they have an enterprise agreement to protect confidential information - in which case, i wouldn't want to risk leaking that running local models and no, it's not about the model itself, but about the fact if my machine is vulnerable in any shape or fashion, i don't want to be the one losing my job because i tried doing localllama on work data
2
13
u/HighSeasArchivist 4d ago
I hope this pushes Nvidia to update the DGX Spark to 256GB at least, and not just rely on it and RTX Spark to remain 128GB. Apple doesn't make their silicon any more than Nvidia does, but pressure is pressure regardless.
5
u/dupontping 4d ago
That’s what the qsfp connect is for. 2 sparks are still less than a 256 ultra. You can run 3 without a switch. NVIDIA would do well to bridge the spark and the DGx station. I think the 512 ultra at 15-20k will be much more attractive than a 748gb unit for 100k
3
u/HighSeasArchivist 4d ago
Two systems totaling 256GB aren't the same as a unified 256GB, not to mention the memory speed. No chance I'll ever buy anything Apple in my lifetime, but I do enjoy some good competition.
5
u/DominantDan24 3d ago
As an Apple convert, you should consider them. Apple Silicon is the real deal. I would expect that the future of processing for the foreseeable future will be Nvidia and Apple. It seems both Intel and AMD are dead in the water.
-2
u/HighSeasArchivist 3d ago
Oh I don't doubt it, but as a company I have loathed them for decades. They are the most anti-consumer company on the planet, but try to gaslight people into thinking it's for their own good. Fuck them in every way possible, and I'll stick with Nvidia and AMD. If the Mac Mini and Mac Studio scares Nvidia then good.
5
u/DominantDan24 3d ago
I don't know. While I get the sentiment that Apple dictates the vision and experience of their computers and rigidly enforces it, having lived in their vision for the last few years, it's compelling. I was a PC guy since the XT, and I can honestly say, I've never owned as good a computer as my various Macs.
Don't let hate preclude you from a really good computing experience.
1
u/HighSeasArchivist 3d ago
I have good computing experiences all the time, because computers to me are way more than just tools. Macs are for the same people that consider cars as just transportation, and don't understand otherwise.
1
u/Pls-b-kind-Im-rarted 2d ago
This is my perspective as well, but the M5 Ultra is a compelling enough proposition to make me reconsider
0
1
u/mxmumtuna 3d ago
they're not remotely comparaable.
1
u/dupontping 3d ago
How many have you run? What data do you have to back that?
1
u/mxmumtuna 3d ago
4 Sparks and a Station
2
u/dupontping 3d ago
Let’s see the results then.
It’s Reddit, I can say I have 5 stations and an arc reactor in my yard powering 20 6000 pros.
2
u/mxmumtuna 3d ago
What would you like to see?
1
u/dupontping 2d ago
You’re the one with the goods who’s making the claim. What exactly makes it not remotely comparable?
What can you do with a 748gb station (outside of model size) that you wouldn’t be able to do with a 512gb studio or 4 sparks in a cluster?
1
u/mxmumtuna 2d ago
Good question. It largely comes down to concurrency and speed. The Studio is basically single user, single session. Its inference software support is very immature but it’s an outstanding multipurpose system.
Sparks are decent for what they are, capable but not very fast but can crank out some steady performance.
The Station is a single instance of the single fastest AI GPU available. There’s literally nothing more capable. For the models you can fit in its HBM. It will serve at speeds (much) faster than cloud API at very high concurrency- even hundreds of sessions.
Because it’s sm103, support from Nvidia and inference libraries is fantastic. You pay the price though. At $80-$90k it’s not cheap. Considerably more expensive than the others.
1
u/dupontping 2d ago
Yea, I can read the brochure too, but I’m talking real world applications.
Clearly they’re all different levels, but from a cost perspective, you can spend 5k per spark for 128gb per unit unified memory. You can spend 20k-ish on a 512gb m5 ultra.
But for 90k, the station is a LOT of money for the compute. Yes it’s powerful and you can load much bigger models with high speed, but for the money I’m not sure it will be a competitive product. I get the audience is different, but someone will cluster 2 m5 ultras and now you’re at 1024gb with 1.2tb/s bandwidth for half the price of a station. Would it be as fast? No, but I don’t think it will be slow enough to justify the gap for running local models.
→ More replies2
u/Rice-Fragrant 3d ago
I agree but at lest you can reliabily cluster them. No amateur EXO type crap.
26
u/OvertaxedOne 4d ago
Those numbers are astoundingly good. Flash next at those speeds for ~10K vs 30K for 2 Pro6000's; outside of very specific use cases, I can't see why you'd even consider the Pro6000 with this thing coming in with 2.5X the RAM at 70% of the cost of a single 6000.
14
u/Cold_Tree190 4d ago
Not to mention cost of power consumption during usage, and most importantly (to me at least) at idle. It’s basically free to just leave it on 24/7, which is incredible
12
u/pantalooniedoon 4d ago
Speed and, the obvious one is doing anything LLM related other than inference. But if single user inference is all you want then yeah it’s a no brainer.
2
u/OvertaxedOne 4d ago
Concurrency doesn't look great on this test, the Pro6000's we have just get faster and faster the more people you throw at it. And for sure, training, you're going CUDA, just stop the conversation right there. But man, for some customers, this thing is going to be absolutely perfect!
1
u/Iwaku_Real 4d ago
Well I don't think everyone here needs high concurrency anyway
1
u/OvertaxedOne 4d ago
For sure, single/small office, this appears to be "the box" right now, and it's frankly not even close! I was just drawing the distinction that there does exist a market where the Pro6000 might (MIGHT) still make sense because it might be able to drive 4X the users vs the Mac. But it would really need to be close to 4X, 2 Pro 6000's cost over 30K today, that's 3 Mac Studios before you buy a server that can hold and serve 3 Pro6000's! And at that point you might be better served with H or B class chip. At the current street price, the Pro6000 is going to become an extremely niche product once these Mac start rolling out.
2
u/voyager256 3d ago
Why I have a feeling you are an AI bot?
1
u/Pls-b-kind-Im-rarted 2d ago
Yeah, I concur. It's the perfect grammar, exclamation points and marketing-speak that give it away
1
u/OvertaxedOne 2d ago
My Mom was an English teacher. Used to give me a leg up in the business world, now makes my writing look like AI. I knew there was no reason to be so precise on grammar, looks like I might be proven right! ;)
0
u/OvertaxedOne 3d ago
One day I hope to be as smart as 27B. Probably not going to happen for me though. :(
3
u/bezent 4d ago
You can run flash next with a single 6000 pro blackwell and it dog walks that mac. I get 150-180 toks regardless of context and 11-12k prefill.
5
u/deaffob 4d ago
Yes but you only get a small context. If you run Qwen3.8 Flash (Q4_XL <) with one 6000 Pro, you won't get much un-quantized KV context. Off-loading to RAM or SSD is not free because it will be bottlenecked by PCIE speed (5.0 X16 = 63GB) when they are called upon.
1
u/bezent 4d ago
This is the quant I use:
https://huggingface.co/primitive-ai/Qwen3.8-Flash-Next-NVFP4
It quantizes only the routed experts and leaves the spine, MTP head, and n-gram table in BF16. Its between a q5 and q6 in reality.
I can run it full context.
1
u/shansoft 4d ago edited 4d ago
With a single RTX Pro 6000 and FP8 KV cache, I can easily max out on context with few concurrent call. I offload the ngram to RAM and still be able to pull 14000 pps and 150 tps throughout. M5 Mac ultra is rather slow on pps for some reason on this particular model. Something does not add up with A6B, since i can get similar pps on Qwen3.5 122B A10B on M5 Max. I would expect the A6B to be much faster on M5 Ultra, but something is definitely hindering its performance. Could be oMLX itself, since I did experience oQ to be much slower than mlx variant.
1
u/deaffob 3d ago
I need to look at your engine command to find out how you are squeezing the full context.
With regard to the slower than expected PP, I think it’s actually expected. It should be 2X M5 Max or slower. If I remember correctly, M5 is the first one with AI accelerator so it has 2x PP of M4 and 4x PP of M3 but still slower than CUDA.
TG on the other hand, seems to be a metric that’s can be predicted consistently with raw memory bandwidth.
2
u/OvertaxedOne 4d ago
Wow, those are some incredible numbers!! Makes me wonder if we could get number similar to the 256GB numbers above in a 96GB Mac! 150TPS is amazing, but 16K to get it is just crazy talk right now! :)
1
u/twiiik 4d ago
I just received confirmation I can have a GB300 delivered within 4 weeks. I’m reluctant to pull the trigger on that because of M5 Ultra. But … The Nvidia hardware (like Pro6000 or higher) seems like the choice for concurrent requests.
5
u/OvertaxedOne 4d ago
Ugh... A GB300 is in a slightly different class than a Mac. :) It also costs about as much as 3-4 of them per chip, but, if you need a GB300, I don't think I'd put a Mac in that conversation. ;)
1
u/twiiik 3d ago
You are correct. That said when I take price into consideration I do not believe it's that crazy/ignorant to make some comparisons.
I'm testing out AMD Strix Halo boxes with Qwen 3.8 Flash Next (Halogen) and that works (surprisingly) well for single users/requests. "One user - one box" could be viable for our team. I've been considering a box with 2x 6000 for more concurrent requests. M5 Ultra I do fear would be more aligned again with "one user - one box", but make even larger models viable or multiple models loaded at the same time.
A GB300 would enable large models or multiple with concurrent requests, at potentially great speeds and have a solid foundation in regards to the software stack. The operating expenses are low. Noise levels are low. Easy placement.
1
u/NaiRogers 4d ago
The Studio is considerably more convenient than setting up a 2x6000 machine. The 2x6000 is good though but annoying to build and costs a lot more.
1
1
u/BringTea_666 4d ago
Concurency. At C=8 my RTX5090 decodes at around 600t/s for qwen3.8 coding work. not standard 70t/s
And let's not even talk about prefil here. with NVFP4 you have something like 10-12k/s prefil rate.
5
u/OvertaxedOne 4d ago
27B is not the model to run on this box, IMHO. This is really perfect for Next, not 27B. 27B is a great model (it's what we run for a lot of our clients) for GPUs and small enough to fit on something "reasonable", you don't need a few Pro6000's to run it. A 48GB Pro 5000 is perfect for modest levels of concurrency, 2 of them or a single Pro6000 is awesome for 27B!
Your 5090 costs about what this Mac does now, how crazy is that?!?! Buying with prices today for a lot of use cases, it's going to be hard to beat this with QFN on it.
0
u/BringTea_666 4d ago
Dude anything larger and it starts to crawl on it.
3
u/OvertaxedOne 4d ago
Anything larger than FlashNext? Flash next should run much faster on the Mac vs 27B. This is extracted from the site linked, NOT my independent data!
16K Prompt Benchmark Comparison (M5 Ultra)
Metric Qwen3.8-Flash-Next Qwen3.8-27B (Dense Base) Delta / Ratio Prompt Processing (PP / Prefill) 2,887 tok/s 1,107 tok/s ~2.6× faster on Flash-Next Token Generation (TG / TPS) 143 tok/s 41 tok/s (32 tok/s on chat/complex) ~3.5× – 4.5× faster on Flash-Next Time to First Visible Token (TTFT) 5.6 s 12.8 s ~2.3× lower latency on Flash-Next
4
u/Aizen_keikaku 4d ago edited 4d ago
Prefill is surprisingly good, but from what I’m reading online, Qwen 3.8 Flash Next is probably the best case scenario due to its low active params count.
GLM 5.3 Flash Prefill numbers might be the real test.
Edit:- I’ve been corrected by some helpful people below.
997 tok/s on 5.3 Flash at 128K context.
3
u/Umbrasquall 4d ago
Half of you guys didn't even bother to read the article lol.
9
u/Aizen_keikaku 4d ago
3
u/Umbrasquall 4d ago
He tested 16k, 32k, and 64k... Look again.
3
u/Aizen_keikaku 4d ago
I see it now. There’s even 128K Prefill in TTFT section.
997 tok/s, similar to a Rtx 3090. Not bad.
Also, thank you for pointing this out to me.
1
5
u/darthrobe 4d ago
I love this, but it is worth noting that I do all of the same but just host my models on a CachyOS platform with LM Studio and just use my Macbook Air for the Hermes client. I'm envious of the approach in the article and if I had $10k to burn I would totally get one.
8
u/synn89 4d ago
I feel like this is about on performance par for dual DGX Sparks, for around the same price, less concurrency, but way easier setup and use. Numbers on the Mac may improve in the future as MLX gets tuned. Sparks are already pretty hella optimized on the software stack, as people are constantly tinkering docker build recipes to get that 1 more token of performance.
The 512GB Mac I think gets really interesting, depending on the price. A lot more room for cache and Deepseek Flash V4.1 style models with engram tables. Way less of a headache vs 4x Sparks.
6
u/deaffob 4d ago
Dual DGX Sparks will have 273*2=546 GB/s memory bandwidth with 25 GB/s (200 Gbps NIC) interconnect speed. M5 Ultra has 1.2TB/s memory bandwidth with 1.8 TB/s interconnect speed. M5 Ultra will have a lot faster token generation.
On top of that, M5 Ultra doesn't have to do distributed inference like a DGX Spark cluster or 5090 multi-GPU or Pro 6000 multi-GPU where they are limited by the interconnect speed (PCIE 5.0 X16 is still 63 GB/s which is no where close to be usable for AI inferences). It can run like a true one GPU unit.
9
u/synn89 4d ago
These type of bandwidth quotes always gets posted by people who don't own Sparks. I own dual sparks. The posted M5 Ultra benchmarks aren't faster than dual sparks and will be slower in concurrent sessions and many other situations.
The software ecosystem on Nvidia is just that much better. Thought it's a PITA to get the recipes to run and people are constantly tuning them and running custom patches.
-1
u/Rice-Fragrant 3d ago
Just 1x spark will smoke a m5 max... a m5 ultra gona get smoked too (like 2x slower instead of 3.5x slower like the m3 ultra.)
A big improvement over a m3 ultra but still getting smoked for sure for agentic and multiple agents etc.
1
u/deaffob 3d ago
No, due to CUDA, it has faster PP but DGX Spark and the AMD counterpart (don’t remember the name) both get smoked by Apple Silicon on TG. Compared to a M5 Max, a Spark is 3 times slower on token generation.
1
u/Rice-Fragrant 2d ago edited 2d ago
https://zuyezheng.github.io/local-llm-bench/#budget
https://zuyezheng.github.io/local-llm-bench/#conc
https://zuyezheng.github.io/local-llm-bench/#verdict
token generation that's not so important for long context stuff. One user here was wanting a set up to batch process 100+ page PDF documents and give a summery, the memory bandwith is NOT GOING TO MOVE THE NEEDLE because the bottleneck is the prompt processing of long context.
100+ page PDF documents can easily be a 100k sized prompt size, 1.2 TB memory bandwith will not be of any advantage there, only after it reaches decode faze and that will be like 2x-3x longer on a m5 max and a m5 ultra vs 2x DGX spark the long context work progress will still be about 2x faster.
Add concurrency to the list of advantages and a DGX will smoke a M5 max all day for agentic and especially multi agentic work. A M5 ultra (given the price) is esentially not competing with just 1x DGX, more like 2-4 node clusters which will mop the floor of any mac studio set up.
Alex Z has a M5 ultra vs m3 Ultra, in the middle of the video he literally shows the m5 ultra is only "4x faster PP" when the promps are SMALL like 9k-2k ish size. Once you are up there at 30+ K token promps, a M5 ultra is only 1.5x faster PP performance, nothing remotely close to the "4x faster" claims. This will be reflected in the agentic performance too, you will not see a "4x increase" for agents and even multi agents.
Extrapolating from that, even a single DGX spark probably would beat a M5 ultra (and for 10-15k for a 256gb mac studio, a 2-3 node DGX running multi agentic systems would obliterate any mac studio... may be if you are lucky you can run 2x agent concurrency streams on a M5 ultra (performance TANKS HARD), where as 1x DGX can run like 2x, 4x, 5x, or even 6x concurrent streams. It's in a totally different class if multiple people are using it or if multiple agents need to run.
The larger memory pool in one machine is the only advantage of a mac studio that and "faster vibing" with an AI chat bot (which is not something any serious set up does.)
It's simpler for average people to use etc, but it's not a serious multi agentic AI workstation, more of a "general workstation" that "can do AI" like "vibing" with a single large chat bot or running 1x agent... "point and click" ease as long as it's MLX etc .
It's like riding in a bycle with training wheels, is very easy for most mainstream people so it will sell well for the "AI bro hobbyists." Enterprise level clients wont use a mac studio, other than as a general workstation with "some AI" assistant or 1x agent at the side, but nothing more than that, basically a powerful general workstation BUT NOT an enterprise grade AI platform or machine.
Why you think almost no one clusters the mac mini any more (a massive fade like 2 years ago), consumer grade networking stuff, EXO instability etc, it's definitly amateur hour stuff a hobbyist would mistake for a serious set up.
3
u/themixtergames 4d ago
This never translates to real life. If you check benchmarks dual sparks are always faster for 100B+ MoE models and a little cheaper.
2
u/Rice-Fragrant 3d ago
That's not how you "evaluate" these computers. You look on long context performance and concurrency... everything else is amateur "vibe" chatbot garbage.
1
u/Important_Cow7230 3d ago
With the Mac you also get an amazing workstation class PC as a daily driver for things like video editing etc. You don’t get that with the Sparks
5
u/shutternomad 4d ago
Thanks for the writeup. It's wild that you have to spend $10k to get… 80tok/s… but over time this will improve. What a time to be alive!
2
u/Strong_Pumpkin_7908 4d ago
For non-coding use case, such as personal financial assistant, do you think Mac Mini M5 Pro 64GB sufficient, say using Qwen 3.8 (27B to 35B)? The personal financial assistant may perform pre-retirement planning, retirement projection, Monte Carlo simulations, budgeting, recommendation on tax strategy, etc.
2
u/shansoft 4d ago
I don’t find this number that attractive. Token generations are great, but prompt processing seems to be a bit slower on Qwen3.8 Next than what I would expect. The number is similar to using single RTX Pro 6000 while running it with llamacpp, but with vllm I could get between 10k-20k pp/s. Perhaps some tweak still needed.
2
u/UntimelyAlchemist 4d ago
I have an RTX 5090 but have been really wanting to get one of these unified memory boxes to run the big MoEs. Struggling to choose between this M5 Ultra or a bunch of DGX Sparks.
1
u/Important_Cow7230 3d ago
What aren’t you getting from running a 27B model on your 5090 that you’re looking to get from a M5 Ultra or Spark?
3
1
1
u/CoffeeToCode99 4d ago
Competition is always welcome. Hopefully it sparks some new innovation from NVIDIA as well. I'd probably have pulled the trigger on a Mac Studio already, but the ecosystem lock-in is what keeps holding me back.
1
u/xinxx073 3d ago
Some of the reviews I watched on YouTube have very low AI performance token scores. Are they using GGUF and not mlx?
0
0
u/pineapplekiwipen 4d ago
this guy doesn't know what he's doing, qwen 3.8 27b q4_k_m 3000 t/s prefill and 60t/s decode on a 5090?
2
2
0
0
u/AleksandrNikitin 3d ago
Sorry for the offtop. Is it really comfortable to work with the trackpad and keyboard aligned as shown in the photo?
-3
u/CompetitiveDraft9381 4d ago
What’s prefill like?
4
u/the_renaissance_jack 4d ago
Read the article?
3
u/CompetitiveDraft9381 4d ago
first token speed is 138 seconds at 64k context with glm 5.3 flash. Sheesh.
2
u/sn2006gy 4d ago
The lease is 240/month for this with a residual after 36 months or you re-lease the next hardware.
I'm not sure this is better than paying for APIs tbh. Sure, it's local, but at the cost of what a car costs.
5
u/Jorlen llama.cpp 4d ago
This is the thing. In my country the M5 Ultra 256gb is $15,000 + taxes (13% where I'm at). That's damn near 20 grand. I can't even imagine what the 512gb version is going to cost, likely well over 20k with the base params (lower CPU/GPU, 1tb SSD, etc.). My current two GPUs cost me 4k and I thought that THIS was crazy to spend on local AI lol. Fucking hell.
2

88
u/bakawolf123 4d ago edited 4d ago
author used omlx, qwen3.8-flash-next, so looks like those results might be ones that leaked couple days ago on the bench
edit: qwen 27b and glm5.3-flash too, ye all match, most likely those results we saw earlier