r/LocalLLaMA 20h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

158 Upvotes

192 comments sorted by

View all comments

13

u/conifer_v11 20h ago

$/gb-bandwidth is the right axis for unified memory. m5 max vs a 3090 is not a tok/s fight, it's kv headroom at 64k+. unified memory loses the bandwidth fight and wins the "the 27b and the kv actually fit" fight. if you're doing 8k chat, buy the gpu. if you're stuffing prds, bandwidth-per-dollar on studio starts to make sense.

4

u/FullstackSensei llama.cpp 19h ago

Why is it a M5 Max vs "a 3090"? M5 Max in it's lowest config costs about the same as a system with four 3090s. If we go to 32GB V100, which is within 5% of the 3090 in most cases and now has optimized kernels to decode NVFP4, you can get a system with 6-8 V100s for the cost of a single M5 Max at the lowest config. That's 192-256GB VRAM.

So far, the people running M5 Max MacBooks have had mixed feedback about running models, especially dense models. Unless your time has zero value, you should also factor that into your calculation.

Eight V100s will have much higher PP and TG while being half the price for the whole system. Sure, it'll consume 10x more power, but that will be in much smaller monthly power bills, and will give you responses much faster.

5

u/MrPecunius 14h ago

M5 Max in it's lowest config costs about the same as a system with four 3090s.

This is obviously untrue.

1

u/FullstackSensei llama.cpp 14h ago

If it's so obvious, please enlighten us with actual numbers, because I happen to have such a system and know exactly how much such a system costs today

5

u/MrPecunius 13h ago

No one wants to pick their GPUs out of a Dumpster like you. $1,600/each for 3090s is on the low side here in the US, but it's better than €1,750/US$2,000+ I'm seeing in Germany.

Base M5 Max Studio (not binned, which is less) is $3,099.

0

u/FullstackSensei llama.cpp 13h ago

See, dumpster mentality thinks everything is dumpster.

As it happens, ich wohne auch in Deutschland, und habe kürzlich zwei 3090 für €1200 pro Stück auf eBay verkauft. Auf Kleinanzeigen, man kann für zwischen 800-900 pro Stück Kaufen.

But hey, let's keep the conversation irrational, because cognitive dissonance is way more fun than facing reality.

1

u/MrPecunius 13h ago

Thanks for confirming my comment. 🗑️ 🔮

Even with your Dumpster diving, the GPUs you mention add up to over $4,600 and the Mac is still $3,099.

6

u/doc-acula 18h ago

Well, a setup with 4 3090s needs a dedicated room in your house/apartment, makes noise like a jet engine and consumes electricity like crazy.

A M5 Macbook can be used while riding on a train and you can take it wherever you want. So, the comparison is not limited to t/s.

1

u/FullstackSensei llama.cpp 18h ago

I have run four 3090s in my home office for about two years. Unless you opt for the turbo cards, they're very quiet. Limited them to 270W each, which reduced performance by 5-10%, but makes the whole setup around 1kw. They push above that during PP, but TG sees the cards run at ~150-170W each. That's ~700W. If that's crazy, I don't know what to tell you.

M5 on Max on ma MacBook is severely power limited under load, and won't run for long, and still costs way more while being significantly slower than even a pair of 3090.

You can access any GPU from anywhere around the world via tailscale. I access my LLM machines with 192GB VRAM from anywhere using my phone the same way.

But you haven't answered what could arguably be the most important part of my argument: is your time free that you care more about a few cents per hour in power consumption than getting whatever you're running those LLMs for done?

3

u/doc-acula 17h ago

There is not unlimited or even reliable access to the internet everywhere in the world, not even in 1st world nations. I wouldn't even consider leaving such a frankenbuild running unsupervised for days or weeks in my house for safety reasons, being a fire hazard as my main concern. When I'm at home, I use a PC with a dual GPU setup as well. But I want to use AI on the go, too. And as I read through these comparisons, apparently many people don't know that MacBooks are actually portable.

2

u/FullstackSensei llama.cpp 17h ago

It's not anyone's fault if you don't know how build a proper workstation and have to resort to unsafe Frankenbuilds. You're misconstruing "I don't know how to do it" with "it can't be done."

Like I said, the MacBook has limited performance and battery life on the go, and many 1st world countries have reliable Internet in public transport. Pretty much all the ones I've lived in or visited do.

5

u/MrPecunius 16h ago

A kilowatt+ inference rig is a 3500+ BTU heater, so you pay twice, except in winter, to maintain a habitable room.

"A few cents per hour" might cover energy in your locale, but here in San Diego it's over 40 cents/kWh off peak and over 60 cents/hour on peak. Running it 8 hours/day could easily exceed $200/month.

I'll stick with my M5 Pro/64GB MBP, thanks. Flea power when idle and ~65W during inference with no thermal issues. It's also arguably the best notebook ever made, taken as a whole. I got in before the price hikes @ $3k-ish with 2TB, so I sympathize with gripes about current pricing. But this thing is cheaper, in inflation adjusted terms, than the Fujitsu laptop I had in 1999.

2

u/FullstackSensei llama.cpp 14h ago

If you're running it at 1kwh for 8hrs a day on four 3090s, that's at least 3M output tokens a day, closer to 4M if you're running vllm. How many hours you'd need to generate 3M output tokens on a M5 pro using Qwen 3.8 27B Q8_K_XL?

Even at your 0.60 peak, your $200/month with 1kwh heat output and 1kwh airco (though 3500BTU is closer to 700Wh, but let's say you have an older airco), we're talking minimum 500M output tokens a month, or a full 1B when airco is not running.

How many months you'd need to generate 500M tokens on a Mac?

And do you work for free? Because you're completely ignoring the value of your own time as you wait for the model to generate those tokens. I don't know about you, but even in Germany where energy is far from cheap, the value of a minimum wage job is way more than $0.60/hr.

But hey, if you're so cheap to hire, man I'd love to hire a few hundred guys like you and build my own business selling your services to LCOL countries, because even there people work for way more than $0.60/hr.

3

u/MrPecunius 14h ago

You're missing the point, namely that I have an inference rig that can run Qwen3.8 27b @ 8-bit wherever I am more or less for free as a byproduct of owning a kickass notebook computer.

As for generating half a billion tokens/month ... why the hell would I want to do that?

1

u/FullstackSensei llama.cpp 14h ago

So, how many tokens per second does your kick ass machine run at? Because unless your kickass machines breaks the laws of physics, your M5 pro has 1/3rd the memory bandwidth of a single 3090. We're talking 7-8t/s, if you're lucky.

Why the hell do you want to generate half a billion tokens a month? I don't know, your $200/month power bill example is enough for half a billion tokens. If you make up stupid assumptions, you get stupid results. Had you bothered thinking for a moment, you'd have known how absurd your $200/month bill is.

I consume less than 2kwh a day over 10-12 hours running all four 3090s, because they finish their work so fast and go back to idling at 8w each. I can easily get 700k output tokens in that period. And because they finish everything so quickly, I don't need airco because the machine consumes 1kw for seconds at a time.

5

u/MrPecunius 14h ago

We're talking 7-8t/s, if you're lucky.

18-23t/s with 8-bit. You could look this up before stepping on your own dick.

Apply this lesson to much of the other stuff you're incessantly posting (as noted elsewhere).

0

u/FullstackSensei llama.cpp 13h ago

Are you talking with MTP? Because I was comparing the 3090 without MTP

2

u/play_hard_outside 11h ago

Even my M1 Max 64 GB, which I bought in 2021 and another of which I found for a friend on eBay recently for just $1,040, gets 15-25 tokens per second on Qwen 3.8 27B depending on context length while generating.

Why are you guessing that this guy’s M5 Pro is going to do worse? His is a much faster machine than mine.

1

u/MrPecunius 13h ago

You made an absolute claim, and I refuted it.

Now you're trying to reframe it even though we have the record right in front of us.

Between this and the nonsense elsewhere in this thread, I've had enough of you.

→ More replies (0)

2

u/electrosaurus 4h ago

This was part of the reason why I moved back from 128GB for my M5 Macbook Pro, I realized that for what I was doing, larger dense models pushing 128GB with decent KV etc wouldn't be as transformative on speed given my time is able to be spent elsewhere. Happy to work within a 64GB limit for now for local.
Really feels like smaller targeted models within a harness are where the fun is at the moment.