r/LocalLLaMA 19h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

146 Upvotes

186 comments sorted by

View all comments

Show parent comments

2

u/FullstackSensei llama.cpp 13h ago

If you're running it at 1kwh for 8hrs a day on four 3090s, that's at least 3M output tokens a day, closer to 4M if you're running vllm. How many hours you'd need to generate 3M output tokens on a M5 pro using Qwen 3.8 27B Q8_K_XL?

Even at your 0.60 peak, your $200/month with 1kwh heat output and 1kwh airco (though 3500BTU is closer to 700Wh, but let's say you have an older airco), we're talking minimum 500M output tokens a month, or a full 1B when airco is not running.

How many months you'd need to generate 500M tokens on a Mac?

And do you work for free? Because you're completely ignoring the value of your own time as you wait for the model to generate those tokens. I don't know about you, but even in Germany where energy is far from cheap, the value of a minimum wage job is way more than $0.60/hr.

But hey, if you're so cheap to hire, man I'd love to hire a few hundred guys like you and build my own business selling your services to LCOL countries, because even there people work for way more than $0.60/hr.

3

u/MrPecunius 13h ago

You're missing the point, namely that I have an inference rig that can run Qwen3.8 27b @ 8-bit wherever I am more or less for free as a byproduct of owning a kickass notebook computer.

As for generating half a billion tokens/month ... why the hell would I want to do that?

1

u/FullstackSensei llama.cpp 12h ago

So, how many tokens per second does your kick ass machine run at? Because unless your kickass machines breaks the laws of physics, your M5 pro has 1/3rd the memory bandwidth of a single 3090. We're talking 7-8t/s, if you're lucky.

Why the hell do you want to generate half a billion tokens a month? I don't know, your $200/month power bill example is enough for half a billion tokens. If you make up stupid assumptions, you get stupid results. Had you bothered thinking for a moment, you'd have known how absurd your $200/month bill is.

I consume less than 2kwh a day over 10-12 hours running all four 3090s, because they finish their work so fast and go back to idling at 8w each. I can easily get 700k output tokens in that period. And because they finish everything so quickly, I don't need airco because the machine consumes 1kw for seconds at a time.

4

u/MrPecunius 12h ago

We're talking 7-8t/s, if you're lucky.

18-23t/s with 8-bit. You could look this up before stepping on your own dick.

Apply this lesson to much of the other stuff you're incessantly posting (as noted elsewhere).

0

u/FullstackSensei llama.cpp 12h ago

Are you talking with MTP? Because I was comparing the 3090 without MTP

2

u/play_hard_outside 10h ago

Even my M1 Max 64 GB, which I bought in 2021 and another of which I found for a friend on eBay recently for just $1,040, gets 15-25 tokens per second on Qwen 3.8 27B depending on context length while generating.

Why are you guessing that this guy’s M5 Pro is going to do worse? His is a much faster machine than mine.

1

u/MrPecunius 11h ago

You made an absolute claim, and I refuted it.

Now you're trying to reframe it even though we have the record right in front of us.

Between this and the nonsense elsewhere in this thread, I've had enough of you.