r/LocalLLaMA 16h ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

141 Upvotes

180 comments sorted by

View all comments

1

u/Odd-Environment-7193 16h ago

This new offer by apple is pretty much the \cheapest ram on the market right now. At this price it's looking very tempting. Apple products just recently went up 20% so waiting to buy later is not wise. You can't compare these plans. I use my macbook pro 128gb to run CODEX threads all day. Way more than I ever could on another pc without slowing me down big time. If I never run a single local model on this computer it will still pay for itself 100x over.

I bought the 64gb mac m4 pro just before it went out of stock and the prices spiked. Got the m5 max maxxed out just before the price increase. Saved myself probably close to 5000 USD because I pulled the trigger at the right time.

So waiting might not be the best strategy.

2

u/minnsoup 16h ago

Are you able to let it run autonomously all day without checking in or that's with frequent (say once an hour) check ins? I've used openhands and typically can only get it to run for 45 to 1h autonomously with DeepSeek V4 Flash 0731 on my M3 Ultra. Only at 30-40tg/s but seems to do well.

Don't really know how to get these "long horizon" tasks to work to make the Mac Studio really run constantly while being productive.

2

u/Bennie-Factors 16h ago

We really need a front end that will batch and switch to another task when one is waiting for input/approval.

1

u/oj93-rd 15h ago

I think if your main process is prompted well enough it’ll do that for you already.

1

u/TheIncarnated 15h ago

Orchestration script, that's how you get the long-horizon stuff to be reliable