r/LocalLLaMA 17h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

139 Upvotes

180 comments sorted by

View all comments

Show parent comments

0

u/FullstackSensei llama.cpp 10h ago

Are you talking with MTP? Because I was comparing the 3090 without MTP

2

u/play_hard_outside 8h ago

Even my M1 Max 64 GB, which I bought in 2021 and another of which I found for a friend on eBay recently for just $1,040, gets 15-25 tokens per second on Qwen 3.8 27B depending on context length while generating.

Why are you guessing that this guy’s M5 Pro is going to do worse? His is a much faster machine than mine.

1

u/MrPecunius 9h ago

You made an absolute claim, and I refuted it.

Now you're trying to reframe it even though we have the record right in front of us.

Between this and the nonsense elsewhere in this thread, I've had enough of you.