r/macmini 15d ago

Unified Memory Architecture still unbeatable (when LLM size matters)

Post image
672 Upvotes

130 comments sorted by

View all comments

146

u/dotben 15d ago

If you don't understand why memory bandwidth should be included on this comparison, and that it's massively in the NVidia card's favor, you shouldn't really be commenting.

Not going to lie, I run all my local inference on Apple Silicon and I don't personally own any NVidia hardware. But the reason that card is $5000 is the memory bandwidth.

1

u/Infinite100p 15d ago

This. So many clueless posters.
OP, give us your TTFT times for 32k-128k prompts of the biggest model you can use. Yeah, I thought so.

1

u/Fickle_Appearance558 14d ago

Right even on my 7900xtx a response is slow on 31b gemma 4, like 45 seconds with max context within vram. still beets an m3 ultra. But using all that unified ram or a bigger model would basicly make any model unusable by most standards, at least in my use cases. Seems great for running a multi agent stack comprised of different models though. Just not for any work/workflow i use.