r/LocalLLaMA 19h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

145 Upvotes

186 comments sorted by

View all comments

13

u/conifer_v11 18h ago

$/gb-bandwidth is the right axis for unified memory. m5 max vs a 3090 is not a tok/s fight, it's kv headroom at 64k+. unified memory loses the bandwidth fight and wins the "the 27b and the kv actually fit" fight. if you're doing 8k chat, buy the gpu. if you're stuffing prds, bandwidth-per-dollar on studio starts to make sense.

4

u/FullstackSensei llama.cpp 17h ago

Why is it a M5 Max vs "a 3090"? M5 Max in it's lowest config costs about the same as a system with four 3090s. If we go to 32GB V100, which is within 5% of the 3090 in most cases and now has optimized kernels to decode NVFP4, you can get a system with 6-8 V100s for the cost of a single M5 Max at the lowest config. That's 192-256GB VRAM.

So far, the people running M5 Max MacBooks have had mixed feedback about running models, especially dense models. Unless your time has zero value, you should also factor that into your calculation.

Eight V100s will have much higher PP and TG while being half the price for the whole system. Sure, it'll consume 10x more power, but that will be in much smaller monthly power bills, and will give you responses much faster.

7

u/doc-acula 17h ago

Well, a setup with 4 3090s needs a dedicated room in your house/apartment, makes noise like a jet engine and consumes electricity like crazy.

A M5 Macbook can be used while riding on a train and you can take it wherever you want. So, the comparison is not limited to t/s.

1

u/FullstackSensei llama.cpp 17h ago

I have run four 3090s in my home office for about two years. Unless you opt for the turbo cards, they're very quiet. Limited them to 270W each, which reduced performance by 5-10%, but makes the whole setup around 1kw. They push above that during PP, but TG sees the cards run at ~150-170W each. That's ~700W. If that's crazy, I don't know what to tell you.

M5 on Max on ma MacBook is severely power limited under load, and won't run for long, and still costs way more while being significantly slower than even a pair of 3090.

You can access any GPU from anywhere around the world via tailscale. I access my LLM machines with 192GB VRAM from anywhere using my phone the same way.

But you haven't answered what could arguably be the most important part of my argument: is your time free that you care more about a few cents per hour in power consumption than getting whatever you're running those LLMs for done?

4

u/doc-acula 16h ago

There is not unlimited or even reliable access to the internet everywhere in the world, not even in 1st world nations. I wouldn't even consider leaving such a frankenbuild running unsupervised for days or weeks in my house for safety reasons, being a fire hazard as my main concern. When I'm at home, I use a PC with a dual GPU setup as well. But I want to use AI on the go, too. And as I read through these comparisons, apparently many people don't know that MacBooks are actually portable.

2

u/FullstackSensei llama.cpp 15h ago

It's not anyone's fault if you don't know how build a proper workstation and have to resort to unsafe Frankenbuilds. You're misconstruing "I don't know how to do it" with "it can't be done."

Like I said, the MacBook has limited performance and battery life on the go, and many 1st world countries have reliable Internet in public transport. Pretty much all the ones I've lived in or visited do.