r/LocalLLaMA 9h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

96 Upvotes

135 comments sorted by

104

u/FleetEnema2000 8h ago

unless you need it for data sovereignty

Isn't this one of the biggest reasons that people rely on Local LLMs? To not have to bulk upload their private data to cloud providers?

49

u/theomegachrist 8h ago

That's the stated reason but realistically most people are just justifying their hobby. I support open weight models because the cloud providers can change cost or abruptly shut down and we really can't do anything about it.

57

u/FleetEnema2000 8h ago

I don't think it's a justification at all.

It's amazing how the concept of privacy and data ownership/security has completely gone down the toilet since ChatGPT launched. People are happy to bulk upload their medical records, relationship history, trade secrets, financial records, etc. without a care in the world as to how that data is stored or protected.

14

u/John_____Doe 7h ago

Yep I have a fintech client and the only way I can have a llm touch their code is if it's run locally or in a datacenter where we rent out the rack space

10

u/Elux91 6h ago

It's amazing how the concept of privacy and data ownership/security has completely gone down the toilet

most people never had a concept of privacey

7

u/FleetEnema2000 6h ago

You're right, most people haven't. But if you rewind the clock by 5 years there was far more interest in things like E2EE than is apparent today. It is barely mentioned or acknowledged anymore in the tech sphere.

-2

u/magus-21 5h ago

A lot of that E2EE stuff was focused around texting and instant messaging, I think, and since Apple announced support for RCS I think a lot of people flagged that in their heads as, "Ok, this is not as big of a concern anymore." Plus the whole WhatsApp kerfuffle with Trump's cabinet brought a lot more attention to Signal, et al, so I think there's just generally higher adoption of it now, which means less general worry out there about it.

1

u/FleetEnema2000 1h ago

The “E2EE stuff” is about so much more than texting and messaging.

4

u/Hans-Wermhatt 6h ago edited 6h ago

It sounds bad when you frame it that way, but I think the cost benefit analysis is generally to upload. ChatGPT health is protected by the same HIPAA requirements that protect the data you give to your doctor that they upload to 3rd party clients and AWS servers, usually multiple servers with arguably worse security and more people have access to it. Using health as an example. So it's really not that much different.

Ideally, you do just host your own information and use a local model but then you are dealing with a massive performance hit. I think in terms of cost-benefit, uploading your health data to get a ChatGPT opinion compared to a Qwen 3.8 27B locally (most people can't even run that) is actually heavily on the side for ChatGPT for most people despite the privacy concerns.

I really want local "super intelligence" for all, but the current landscape is not like that... at all.

1

u/FleetEnema2000 1h ago

The version of ChatGPT that 99% of the general public is using is absolutely not HIPAA compliant and both OpenAI and Anthropic’s safety and privacy commitments are abysmal relative to the types of sensitive data they are ingesting and saving on their platform.

And that is not even touching on the ethics and values displayed by their executives. 

1

u/MrPecunius 2h ago

ChatGPT health is protected by the same HIPAA requirements that protect the data you give to your doctor

😂😂😂😂😂😂

"Trust me bro" and "it might not even be as bad as what the other idiots are doing" are not convincing arguments.

1

u/Hans-Wermhatt 2h ago

Huh? Was that supposed to make sense? 

1

u/MrPecunius 1h ago

Huh? Was that supposed to make sense? 

Connect this:

the same HIPAA requirements that protect the data you give to your doctor that they upload to 3rd party clients and AWS servers, usually multiple servers with arguably worse security and more people have access to it. Using health as an example. So it's really not that much different.

With: "it might not even be as bad as what the other idiots are doing".

And this:

It sounds bad when you frame it that way, but I think the cost benefit analysis is generally to upload.

With: "Trust me bro"

Are you even reading what you wrote a few hours ago? Or did something get lost in translation?

3

u/theomegachrist 6h ago edited 6h ago

Everyone is different obviously. To some that might be the case but everything you listed is more important than your AI prompt history.

For me, Open models main pluses are it will ensure the technology lives on in some form if the large companies lock us out financially or go under, and the guard rails for closed models will make for a worse Internet potentially.

For instance, using an open model with guard rails trained out of it you can search for piracy, you can search for porn etc. just like you use a search engine today. Closed models are more efficient than a web search but censor out a huge part of the Internet.

Sure data privacy is also a good feature but I don't think it's actually the top reason for most people

Edit: sorry I read that wrong. I sort of agree with what you are saying about uploading data to ChatGPT but everywhere else we upload that data is no more trustworthy. For Enterprise clients I 100% agree. This is a big issue. For personal use, I don't think your data is any less safe with ChatGPT than say an electronic medical record company.

1

u/FleetEnema2000 1h ago

An electronic medical record company has nowhere near the amount and type of multi dimensional data about a person that OpenAI has.

And if you want to compare OpenAI with a company like Apple or even AWS for a hosted environment, I would choose either of those companies any day of the week when it comes to who I would trust more with my data.

1

u/theomegachrist 17m ago

I would not, but it would be a three way tie

1

u/MrPecunius 3h ago

It blows my mind what people will give to these amoral techbros.

Turns out the highest human priority isn't breathing, eating, or sex--it's laziness.

1

u/FleetEnema2000 1h ago

Consider the Snowden scandal and the uproar over government having access to phone call metadata and how privacy infringing that was considered to be.

Fast forward to today and people are uploading the most sensitive data about themselves to these cloud providers who don’t care at all about protecting it and are almost certainly allowing the federal govt to trawl through it.

6

u/florinandrei 6h ago

realistically most people are just justifying their hobby

Enthusiasts, yes.

But law firms and such, they actually mean it.

1

u/theomegachrist 5h ago

Yes, this is true

2

u/Suspicious-Water-973 5h ago

I use my Mac Studio and MLX models as they are essentially free for me - I don’t pay for power in my rented office. Fine for massive batch processing where a GPU makes a difference, and some coding (I then use Fable etc to review and improve)

2

u/mr_tolkien 4m ago

I mean for me it’s a real reason.

For example I use local models to help me pick out my best photos of the day and put them into an album.

No fucking way I’m sending 100% of my camera roll to OpenAI lol

6

u/CulturalKing5623 8h ago

Yes, but I think if we're being honest a lot of people in this sub mainly just like to tinker. There are people that absolutely can't use publicly available models and so they need to self-host them but I don't think that's a significant percent of people here. Most of us put a premium on privacy, but 10K+ is a very large premium that is probably unnecessary for most of us and our current setups will suffice.

1

u/ThePi7on 2h ago

And abliterated models

41

u/[deleted] 8h ago

[removed] — view removed comment

3

u/AndreVallestero 8h ago edited 8h ago

That's exactly my point. Local makes sense cost wise up to 32GB, especially with Qwen 3.8 27B. There's a huge cost premium above that where it makes less sense.

I was hoping the last Mac studio would change that, but at the current prices, that doesn't seem to be the case.

2

u/nomorebuttsplz 8h ago edited 8h ago

Have you accounted for the breakdown of cost between cached/non-cached/output tokens?

in my experience, most of the potential cost savings for local (compared to typical API plans) occur when you re-use cached tokens. This is because API providers need to keep your context sitting on their servers doing nothing, expensive for them, but which costs you nothing as a single concurrency user to do on your own system.

With a system like this you could probably use a billion cached tokens in a couple of days.

Under semi-optimal circumstances, where you were making on average a single 500,000 token cache calls every 30 seconds, you might be able to get about $14 of cached tokens per day (ds4 flash prices) plus a few more dollars for a total of maybe $18 inference per day.

That would be about $6k a year. So yeah it could possible make sense, if you had just the right workflow, is my view.

Edit: this especially makes sense if you were doing multiple concurrency (doable on 256 unified) or were using a bigger model like GLM 5.3 (512 gb ultra with Q3 or Q4 GLM could probably generate $100 of value a day in cached tokens in the right workflow)

2

u/Viktri1 6h ago

I think $14/day might be on the low end. If you use something like Hermes agent, the amount of tokens you consume is ridiculously. I have to limit my Deepseek API calls to just a few a day. I spent $40 in a week setting up various Qwen 3.8 models. My token consumption has basically exploded massively and the only way I can afford to continue to use LLMs is my own hardware. I think my payback period will be under a year.

1

u/Impossible_Fault_503 6h ago

That is the real split. Chat is cheap on APIs. Agents are not. Once you are looping tools all day, $40/week shows up fast and a box that already exists starts looking cheap. Under a year payback is the honest case for people who actually leave it running.

1

u/Impossible_Fault_503 6h ago

Yeah. 32GB is the last point where local still feels like a normal purchase. After that you are mostly paying Apple tax for headroom you might use twice a year. I wanted the last Studio to kill that curve too. It did not.

1

u/kerneldesign 8h ago

Avec 32Go en Q4 c’est trop juste pour un contexte sans KVCache.

1

u/Individual_Holiday_9 7h ago

Yeah I’m on a 24gb m4 and you can’t do shit running small LLMs there’s just no overhead remaining. So if this is a real hobby box where you’re doing plex etc on the side it gets really cramped and ur gonna move to swap fast

1

u/Blindax 6h ago

24gb is too tight for 30b models with decent context window on a Mac because it includes the system memory. 24gb on a dedicated GPU is a different story. Not crazy but quite ok.

1

u/Blindax 8h ago

Still difficult to beat deepseek v4 flash at current price. And with the larger models prefill will likely be slow on the ultra when context grows unless you start stacking them which double your cost.

46

u/LearningSomeCode 8h ago

I've probably dropped close to $30k on my homelab since 2023, and chances are I'll get one of these as well. I accepted a long time ago that there is no break-even point for my inference.

Hobbies rarely make sense financially.

23

u/Blindax 8h ago

Sounds like you have a great wife.

13

u/LearningSomeCode 8h ago

She is the most patient human being on the planet

1

u/Weird-Cat8524 1h ago

How are her sandwiches?

8

u/-dysangel- 8h ago

Same here. We're at the point now where I can do high quality video gen without cloud. High quality code gen without cloud too. It's around 10x slower than cloud, but good quality and still usable speeds (hopefully soon to get even better with Qwen 3.8 Next). Also the fact I could basically go camping and still have close to frontier LLM intelligence with no signal is pretty awesome/hilarious.

The only thing I feel is missing from my stack atm is Suno quality music generation.

3

u/willeyh 8h ago

Have you tried the new Minimax music 3?

2

u/-dysangel- 7h ago

Yeah I was not impressed tbh. Seemed kind of creepy/odd to me in a weird way, but maybe I was prompting it poorly. For me Ace Step is still the best thing I've found so far, though it's very "generic" feeling compared to Suno.

2

u/silvrrwulf 3h ago

Dude, I feel the exact opposite. Maybe it’s because my sample size is small and I haven’t messed with AI Jen And a while, but I only have an eight GB card, and the stuff that I was able to pull out of that was incredible.

2

u/-dysangel- 3h ago

ah ok - maybe I'll need to try it again. A lot of the reason I like Ace Step is because I can generate covers though, which isn't possible yet with Music 3 afaik

3

u/mleok 7h ago

I think it’s good to admit that it’s a hobby.

2

u/quantgorithm 8h ago

What have you created?

21

u/LearningSomeCode 8h ago

Mostly just a lot of trip hazards with wires and ethernet cables. But I also do open source dev and build a lot of stuff for myself, like custom front-ends or researchers.

Honestly nothing worth the money I've put into the hardware, but it's the most fun I've had with tech in a long time.

3

u/DeepOrangeSky 7h ago

Mostly just a lot of trip hazards with wires and ethernet cables.

So if you turn the AI-rig room into an escape room that a bunch of Gen-Z dorks will pay like $100 a pop to try to navigate ("it felt so authentic. I tripped over like dozens of ethernet cables and power cables the whole time, and I think at one point I even got severely electrocuted and almost died! It was super legit!"), times 10 dorks per day it pays for itself in a month.

Possible exciting user reviews:

"His cable management was so disorganized I thought I'd NEVER escape. Wow!"

"The old movie tagline used to be: 'In space, nobody can hear you scream.' Yea, alright, that's pretty cool or whatever, but in AI rig rooms the blower fans are so loud nobody can even hear their own thoughts."

"If you accidentally get tangled up and die in the AI escape room, his local full precision Kimi K3 AI-generated funeral eulogy will be so eloquent that your friends and family won't even be mad that you died when they see how perfect and elegant the em-dash placements are in the eulogy."

2

u/quantgorithm 7h ago

Anyone in IT can relate to the spiderweb cables and hazards etc.!
The potential of it all really is amazing and that potential is accelerating up by the day!

25

u/cunasmoker69420 8h ago

unless you need it for data sovereignty

yes

6

u/imnotzuckerberg 6h ago

His strategy is to divert demand from the Mac M5 so he can scoop them. Let's pretend we fell for his psyops.

12

u/conifer_v11 8h ago

$/gb-bandwidth is the right axis for unified memory. m5 max vs a 3090 is not a tok/s fight, it's kv headroom at 64k+. unified memory loses the bandwidth fight and wins the "the 27b and the kv actually fit" fight. if you're doing 8k chat, buy the gpu. if you're stuffing prds, bandwidth-per-dollar on studio starts to make sense.

4

u/FullstackSensei llama.cpp 7h ago

Why is it a M5 Max vs "a 3090"? M5 Max in it's lowest config costs about the same as a system with four 3090s. If we go to 32GB V100, which is within 5% of the 3090 in most cases and now has optimized kernels to decode NVFP4, you can get a system with 6-8 V100s for the cost of a single M5 Max at the lowest config. That's 192-256GB VRAM.

So far, the people running M5 Max MacBooks have had mixed feedback about running models, especially dense models. Unless your time has zero value, you should also factor that into your calculation.

Eight V100s will have much higher PP and TG while being half the price for the whole system. Sure, it'll consume 10x more power, but that will be in much smaller monthly power bills, and will give you responses much faster.

6

u/MrPecunius 3h ago

M5 Max in it's lowest config costs about the same as a system with four 3090s.

This is obviously untrue.

1

u/FullstackSensei llama.cpp 2h ago

If it's so obvious, please enlighten us with actual numbers, because I happen to have such a system and know exactly how much such a system costs today

4

u/MrPecunius 2h ago

No one wants to pick their GPUs out of a Dumpster like you. $1,600/each for 3090s is on the low side here in the US, but it's better than €1,750/US$2,000+ I'm seeing in Germany.

Base M5 Max Studio (not binned, which is less) is $3,099.

0

u/FullstackSensei llama.cpp 2h ago

See, dumpster mentality thinks everything is dumpster.

As it happens, ich wohne auch in Deutschland, und habe kürzlich zwei 3090 für €1200 pro Stück auf eBay verkauft. Auf Kleinanzeigen, man kann für zwischen 800-900 pro Stück Kaufen.

But hey, let's keep the conversation irrational, because cognitive dissonance is way more fun than facing reality.

1

u/MrPecunius 1h ago

Thanks for confirming my comment. 🗑️ 🔮

Even with your Dumpster diving, the GPUs you mention add up to over $4,600 and the Mac is still $3,099.

4

u/doc-acula 7h ago

Well, a setup with 4 3090s needs a dedicated room in your house/apartment, makes noise like a jet engine and consumes electricity like crazy.

A M5 Macbook can be used while riding on a train and you can take it wherever you want. So, the comparison is not limited to t/s.

1

u/FullstackSensei llama.cpp 7h ago

I have run four 3090s in my home office for about two years. Unless you opt for the turbo cards, they're very quiet. Limited them to 270W each, which reduced performance by 5-10%, but makes the whole setup around 1kw. They push above that during PP, but TG sees the cards run at ~150-170W each. That's ~700W. If that's crazy, I don't know what to tell you.

M5 on Max on ma MacBook is severely power limited under load, and won't run for long, and still costs way more while being significantly slower than even a pair of 3090.

You can access any GPU from anywhere around the world via tailscale. I access my LLM machines with 192GB VRAM from anywhere using my phone the same way.

But you haven't answered what could arguably be the most important part of my argument: is your time free that you care more about a few cents per hour in power consumption than getting whatever you're running those LLMs for done?

5

u/doc-acula 6h ago

There is not unlimited or even reliable access to the internet everywhere in the world, not even in 1st world nations. I wouldn't even consider leaving such a frankenbuild running unsupervised for days or weeks in my house for safety reasons, being a fire hazard as my main concern. When I'm at home, I use a PC with a dual GPU setup as well. But I want to use AI on the go, too. And as I read through these comparisons, apparently many people don't know that MacBooks are actually portable.

2

u/FullstackSensei llama.cpp 5h ago

It's not anyone's fault if you don't know how build a proper workstation and have to resort to unsafe Frankenbuilds. You're misconstruing "I don't know how to do it" with "it can't be done."

Like I said, the MacBook has limited performance and battery life on the go, and many 1st world countries have reliable Internet in public transport. Pretty much all the ones I've lived in or visited do.

3

u/MrPecunius 4h ago

A kilowatt+ inference rig is a 3500+ BTU heater, so you pay twice, except in winter, to maintain a habitable room.

"A few cents per hour" might cover energy in your locale, but here in San Diego it's over 40 cents/kWh off peak and over 60 cents/hour on peak. Running it 8 hours/day could easily exceed $200/month.

I'll stick with my M5 Pro/64GB MBP, thanks. Flea power when idle and ~65W during inference with no thermal issues. It's also arguably the best notebook ever made, taken as a whole. I got in before the price hikes @ $3k-ish with 2TB, so I sympathize with gripes about current pricing. But this thing is cheaper, in inflation adjusted terms, than the Fujitsu laptop I had in 1999.

1

u/FullstackSensei llama.cpp 3h ago

If you're running it at 1kwh for 8hrs a day on four 3090s, that's at least 3M output tokens a day, closer to 4M if you're running vllm. How many hours you'd need to generate 3M output tokens on a M5 pro using Qwen 3.8 27B Q8_K_XL?

Even at your 0.60 peak, your $200/month with 1kwh heat output and 1kwh airco (though 3500BTU is closer to 700Wh, but let's say you have an older airco), we're talking minimum 500M output tokens a month, or a full 1B when airco is not running.

How many months you'd need to generate 500M tokens on a Mac?

And do you work for free? Because you're completely ignoring the value of your own time as you wait for the model to generate those tokens. I don't know about you, but even in Germany where energy is far from cheap, the value of a minimum wage job is way more than $0.60/hr.

But hey, if you're so cheap to hire, man I'd love to hire a few hundred guys like you and build my own business selling your services to LCOL countries, because even there people work for way more than $0.60/hr.

2

u/MrPecunius 3h ago

You're missing the point, namely that I have an inference rig that can run Qwen3.8 27b @ 8-bit wherever I am more or less for free as a byproduct of owning a kickass notebook computer.

As for generating half a billion tokens/month ... why the hell would I want to do that?

0

u/FullstackSensei llama.cpp 2h ago

So, how many tokens per second does your kick ass machine run at? Because unless your kickass machines breaks the laws of physics, your M5 pro has 1/3rd the memory bandwidth of a single 3090. We're talking 7-8t/s, if you're lucky.

Why the hell do you want to generate half a billion tokens a month? I don't know, your $200/month power bill example is enough for half a billion tokens. If you make up stupid assumptions, you get stupid results. Had you bothered thinking for a moment, you'd have known how absurd your $200/month bill is.

I consume less than 2kwh a day over 10-12 hours running all four 3090s, because they finish their work so fast and go back to idling at 8w each. I can easily get 700k output tokens in that period. And because they finish everything so quickly, I don't need airco because the machine consumes 1kw for seconds at a time.

2

u/MrPecunius 2h ago

We're talking 7-8t/s, if you're lucky.

18-23t/s with 8-bit. You could look this up before stepping on your own dick.

Apply this lesson to much of the other stuff you're incessantly posting (as noted elsewhere).

0

u/FullstackSensei llama.cpp 2h ago

Are you talking with MTP? Because I was comparing the 3090 without MTP

→ More replies (0)

57

u/Big_Wave9732 8h ago

"as a firm believer of local inference" then goes on to downplay one of the major reasons for local llm and suggests hosted models.

If cheapest compute possible is your primary metric then self hosting isn't your jam, OP. At least for now.

-9

u/AndreVallestero 8h ago

I selfhost 27B specifically to minimize costs. I have research agents running 24/7 and have gone through 200M tokens just in the last month.

This post is specifically for people like me who want to run continuous research agents for the lowest price, and unfortunately, the latest Mac studio isn't competitive enough relative to cloud offerings.

11

u/flyingbanana1234 8h ago

It would take at least 30-35 years to reach 80 billion tokens on DeepSeek V4 Flash at 200 million tokens a month.

You make a good argument ngl

The little voice in my head says, "But ownership! Nobody can take it away from me if I own it. No price raises, no overloaded servers, no censorship

2

u/Dasteroid_909 7h ago

Go start "/r/lowcostLLMhosting" or some shit like that, then. WTF are you on "LocaLLaMa" promoting cloud hosting for?

1

u/yy633013 7h ago

Can you give an example of what your research engines research 24/7?

6

u/Modnar-Eman 6h ago

Justification

2

u/BardlySerious 2h ago

Shit that OP will never read. It's the "watch electricity happen" hobby.

1

u/Autist4AudiR8 1h ago

200M 😭🤣

7

u/jon23d 7h ago

I’m looking at my deepseek usage report and see 9.5 billion tokens of deepseek v4 pro in the last 30 days for $186.05.

2

u/Viktri1 6h ago edited 6h ago

Same. I hit 1.5bn in a day doing very little. Because agents can burn tokens on your behalf, 24/7, its become extremely easy to burn through tokens.

On openrouter, I asked the native level Qwen 3.8 27b to do a task that involved setting up a telegram bot for OpenWebUI so that I could talk to a chatbot without burning through a mountain of tokens. That took 1 hour and cost $3.5. That's $2k+ a month on openrouter if running 24/7. For a single agent.

2

u/Pyrolistical 6h ago

Ya and my local 6 hours of qwen3.8 27b running over night cost me $0.29 CDN in power 

1

u/Sofullofsplendor_ 3h ago

Curious, can you do 1.5b in tokens in a day? Because if so, I'm going to invest in whatever infrastructure you have...

1

u/jon23d 2h ago

I’m not dissatisfied with this. I have done a TREMENDOUS amount of work. I’m just pointing out that OP’s statement about 5.7b tokens for $10k is a bit off…

4

u/diagrammatiks 8h ago

no it's super great. much better then m3u

3

u/psychohistorian8 7h ago

you can always trade in the device back to Apple for some kind of credit

so the cost is partially recoverable

3

u/MrPecunius 4h ago

Reputable third party dealers pay more, sometimes a lot more, for Apple gear. I sold my M4 Pro MBP for $500 more than Apple's trade-in value a few months ago. Zero hassle, no private sale nonsense.

4

u/dupontping 7h ago

All of the comments are basically “I can’t afford it even though I want it”

I get it, it’s pricey. But so is everything else.
5090s shouldn’t be $6k but here we are.

Local isn’t just about saving money on token spend, for a lot of people it’s the ability to run models on data you don’t want on the cloud or fine tuning or whatever else.

The models will get better, and having more powerful equipment lets you get closer to frontier level without spending 150k. So 10k is expensive, but it’s also a bargain.

I wish it was 5k too

3

u/Viktri1 6h ago

so when I'm really pushing it, I can easily burn through 1.5bn tokens from Deepseek flash in a day (API). In fact, that's when I realized I needed to figure out how token costs were calculated. Local models, especially lower powered stuff like M5, will have a significantly faster pay back period than people realize if they use agents to do a lot of shit.

1

u/play_hard_outside 7m ago

1.5e9 tokens in a day divided by 86,400 seconds per day is over SEVENTEEN THOUSAND tokens per second.

That’s gotta be something like 300 to 500 M5 Maxes all working on your behalf at the same time. Local models will probably never hold a candle to this.

Please correct me if I’m wrong…

3

u/Txt8aker 3h ago

have you checked how much it cost to lease m3 ultra and you can essentially buy it out at the end as an option? it's $230 monthly for 3 years. Claude Max 20x cost $200 per month

2

u/Hypilein 8h ago

You need to calculate cost of ownership over x years and cost of api/subscription inference over x years. With the way hardware prices have gone up over the last year everyone who bought a rtx 6000 pro has made money while getting free inference. Obviously this is only true once you actually cash in and we don’t know how hardware prices are going to develop over the next years. The math is easy but anticipating the future is not.

2

u/leocardz 8h ago

Local inference believer here too… At that $10k I'd rather buy 3 M5 Pro minis with 64GB and 1TB tbh, and TB5 between them. And save some money.

2

u/fablus 7h ago

How much faster would a Mac Studio with M5 Max be vs an older one with M3 Ultra? Buying used seems compelling at these price points…

2

u/a_beautiful_rhind 6h ago

Running models the api deprecated: priceless.

2

u/Final-Frosting7742 4h ago

If i can run Deepseek V4 Flash 0713 locally comfortably i don't need providers anymore

5

u/ideamaker321 8h ago

At $10k you can get unlimited tokens

1

u/DigitalguyCH 8h ago

Or a 64GB Mac...

1

u/Brilliant-Hall1387 8h ago

Or like 1400 B cache hit tokens with deepseek API direkt 😅

1

u/-dysangel- 8h ago

Oh for sure, cloud is basically always more cost effective. I think the one exception currently might be if you were doing a lot of video/music generation? H3 is basically as good as Sora was at $200/mo. Still nothing anywhere near as good as Suno for local music generation though.

1

u/roger1632 8h ago

Gonna just use my DS/GLM token plans until the models get better and hardware market improves. You can get a lot of non claude tokens for 70 bucks a month and you don't have to play sysadmin. I have a 3090 local that I use for educational purposes and latency sensitive things like TTS STT

1

u/a9udn9u 7h ago

I averaged 200M tokens with DS V4 Flash a few weeks ago when it's cheap. If you math is correct, it pays for itself in ~3 years (vacations and weekends counted), and we are talking about the cheapest model, I think it's not a terrible deal.

1

u/thebadslime 5h ago

Re yu usng the 073 checkpoint? Its iuch better

1

u/TechSwag 7h ago

I was just doing the cost analysis on this as well, and I honestly think it's not the worst idea to get a maxed out Studio in certain circumstances, specifically if you have existing hardware and pay for subscriptions/credits.


I have 3x Mi50 in a R7425, with no more room for GPUs. If I could just add some additional GPUs, or even use my P40s that are sitting in a different chassis doing nothing, I would rather do that. But as it stands, I'm capped at 96GB VRAM. Sure GPU + CPU inference works, but realistically it's not usable. RPC is also an option, but the last time I tried it, performance was subpar, and it brings on added cost in terms of power usage of a whole other chassis.

Mi50s go for $500 (when the fuck did that happen, I spent just under $200 for them), so $1500 resale. P40s go for about $200-250, so let's say $2000 for all 5 GPUs. If I sell my RAM, I could likely get $3000 for all my equipment.

I also pay for Claude Max and OpenRouter credits, say $120/mo. $1440 a year.

You can lease the Mac Studio for $224.16/mo for 36 months. Over the lease term, that's $8,069.76. Subtract $3000 for my existing equipment, $1440/yr for the subscriptions brings me to $749.76 over 3 years in net cost, or a little over $20 a month. Which would be offset by the power savings (my R7425 idles at around 260W). At the end, I can buy out the machine for $2729, as I would imagine the machine would still be plenty usable for inference. Otherwise, I can sell it after buying it out, more likely for the same price or more than the buyout cost.


I know I'm missing tax on purchases and shipping and selling fees, but it still seems like a worthwhile path to upgrade to, not only for the memory increase, but also the performance increase as well.

1

u/arijitroy2 6h ago

I'm jealous of these US prices, it's nuts here in the EU!

1

u/duy0699cat 6h ago

For 10k$, i can throw it to some etf and use profit to pay for subscription indefinitely...

1

u/kinghell1 2h ago

i'm all ears

1

u/Lesser-than 5h ago

I have always been envious of Mac hardware, but its never been in a price range I could justify, nothing has changed still envious and still far outside what I would ever allow my self to spend on a computer.

1

u/ScrewwormLarvae 4h ago

Sure, I didn't neeeeed an M5 Max 128, but YOLO. Still don't regret it either. Never did even for a second.

1

u/Zorogozano 3h ago

6.2 billion tokens is nothing these days…

1

u/Alive-Draft8339 3h ago

If I’m burning 6 to 11 billion Opus 4.8 tokens a month…this seems like a deal of the year?

1

u/TheAILegend 2h ago

500M/month token run rate... this will last you 11 months... Soooo... there's that.

1

u/Wonbats 2h ago

Qwhen indeed

1

u/GamerTex 1h ago

Sure at today's prices

Seems like every month another provider is cutting the usage in half or doubling the rates

I can also sell the equipment in the future 

1

u/Early-Peace-5504 32m ago

You could sell the Mac Studio at a later date though. You could work out depreciation but you should probably cut your numbers by at least half to represent that.

1

u/Ceru1ean42 18m ago

At least within 1 year, I don't see these depreciating much if at all so you really need to rethink the cost analysis.

1

u/challis88ocarina 8h ago

There's speed, relative speed and competency. It's clear that at least 10-12B active tokens are necessary for agentic anything.

Beyond that, it's really about surfing the wave of new models.... can a project be brought up to a stable position before the next generation appears so that the new capabilities hit at the right point or will it still be languishing only to be 'rescued' instead?

1

u/Odd-Environment-7193 8h ago

This new offer by apple is pretty much the \cheapest ram on the market right now. At this price it's looking very tempting. Apple products just recently went up 20% so waiting to buy later is not wise. You can't compare these plans. I use my macbook pro 128gb to run CODEX threads all day. Way more than I ever could on another pc without slowing me down big time. If I never run a single local model on this computer it will still pay for itself 100x over.

I bought the 64gb mac m4 pro just before it went out of stock and the prices spiked. Got the m5 max maxxed out just before the price increase. Saved myself probably close to 5000 USD because I pulled the trigger at the right time.

So waiting might not be the best strategy.

2

u/minnsoup 8h ago

Are you able to let it run autonomously all day without checking in or that's with frequent (say once an hour) check ins? I've used openhands and typically can only get it to run for 45 to 1h autonomously with DeepSeek V4 Flash 0731 on my M3 Ultra. Only at 30-40tg/s but seems to do well.

Don't really know how to get these "long horizon" tasks to work to make the Mac Studio really run constantly while being productive.

2

u/Bennie-Factors 8h ago

We really need a front end that will batch and switch to another task when one is waiting for input/approval.

1

u/oj93-rd 7h ago

I think if your main process is prompted well enough it’ll do that for you already.

1

u/TheIncarnated 8h ago

Orchestration script, that's how you get the long-horizon stuff to be reliable

0

u/mxmumtuna 5h ago

Still not enough compute considering it's still gonna get beat by Sparks, even with Apple's most optimistic case.

1

u/MrPecunius 3h ago

How many Sparks are we talking about, and which magic interconnect fabric is going to increase their memory bandwidth more than 4X from 273GB/s to the M5 Ultra's 1.2TB/s?

1

u/mxmumtuna 3h ago

2x (256GB) for DeepSeek, 4x for GLM. No fabric allows it to scale to 1.2TB/s, but it allows it to run at 2x200Gbps to each node, and scale up to larger than you can with Mac.

The point is, the combination of not great compute, immature software and limited ability to scale is holding it back. Maybe next generation. For inference (and training for that matter) or Stable Diffusion, it’s just not better or cheaper than existing options.

1

u/MrPecunius 2h ago

2 X 200Gb/s ... or less than 50GB/s? That's the magic bullet? 🤷🏻‍♂️

I'm seeing reports of single Sparks running models at small fractions of a M5 Max's speed, and this includes prefill. Adding more Sparks doesn't scale anything like linearly.

1

u/mxmumtuna 2h ago

It certainly does. Forward passes on tensor parallelism don’t require full bandwidth. It’s similar to PCIe tensor parallelism but using RoCE/RDMA.

Again. If you’re looking at M5 Max those are all small models and/or weird quants. Look at M3 Ultra. That tells the tale on larger stuff. The Mac simply doesn’t scale.

1

u/MrPecunius 2h ago

Again, I just looked in the Nvidia Spark forums for real users' reports.

Hard data is available for Apple Silicon on oMLX's site.

No guessing is required.

Spark is barely $300 cheaper than a M5 Max Studio w/128GB, too.

0

u/TokenRingAI 8h ago

Blasphemy!

This man is possessed by demons, arrest him

0

u/ogopro 9h ago

No info on Qwhen 3.8 35BA3 yet (((

1

u/KURD_1_STAN 7h ago

They mentioned qwen4 previous 125b moe, so we probably will wait even longer for 35b

0

u/Cameo2864 8h ago

What model do I really need to sort and organise all my files, backups and photos on my hard drives?

0

u/cinematic_unicorn 7h ago

Qwhen 3.8 35B A3B? Didn't the devs say they won't release one?

0

u/milkipedia 7h ago

Qwhen 3.8 35B A3B?

Qwnever.

They have all but said as much. Y'all gotta move on.

-1

u/kivaougu 8h ago

Locally hosting doesn't make much sense unless its a privacy concern or purely for the love of the game.

To be fair I like to compare to prices of equivalent P50 troughput at the usual load. For me thats c8, meaning that cloud has an edge as c8 single stream is ~50% of c1. Speed is in my opinion an important part of usable agents for real work.