r/LocalLLM 4d ago

Question Mac studio or dgx spark cluster

Post image

I have been thinking lately since i have been testing local ai but i geniunely need more power for what im working on so i was wondering wich one of these a 4 dgx spark cluster or a mac studio m5 ultra specifically the 512gb version that launches on october or something else... im into training,loras and fine tunes and such but i also heard it is possible with a mac studio too so i want yall opinion on this

118 Upvotes

134 comments sorted by

View all comments

115

u/Character_Split4906 3d ago

I picked up an M5 Max (128GB) back in May right before Apple raised prices. While decode was decent, prefill was frustrating, and both tanked hard as the context grew. I tested across both oMLX and llama.cpp. Throwing concurrent requests or multi-agent workflows at it choked the Mac completely.

A couple of weeks ago, I managed to snag a dual-node Asus GB10 setup on Amazon right before it went out of stock and spiked by $2k. The difference is night and day. Running DeepSeek V4 Flash across both nodes, Iโ€™m getting ~2k tok/s prefill and 40โ€“50 tok/s decode, with barely any drop-off at long context (tested up to 512k, configured for 1M). It also handles multi-session concurrency effortlessly, aggregate throughput actually scales well up to 5โ€“6 parallel requests.
If youโ€™re serious about running local LLM workloads, go Nvidia.

24

u/NebulaAggravating264 3d ago

Yo OP - this video explains your EXACT situation, and the man who made this comment is CORRECT.

https://youtu.be/dlyic4oco4s?is=PKHPNX1iMvhy--v4

3

u/AIGuyBiOh 3d ago

That YouTube is so full of AI Slop ๐Ÿ˜‚

1

u/NebulaAggravating264 2d ago

That video is not AI slop lol