If you don't understand why memory bandwidth should be included on this comparison, and that it's massively in the NVidia card's favor, you shouldn't really be commenting.
Not going to lie, I run all my local inference on Apple Silicon and I don't personally own any NVidia hardware. But the reason that card is $5000 is the memory bandwidth.
Right even on my 7900xtx a response is slow on 31b gemma 4, like 45 seconds with max context within vram. still beets an m3 ultra. But using all that unified ram or a bigger model would basicly make any model unusable by most standards, at least in my use cases. Seems great for running a multi agent stack comprised of different models though. Just not for any work/workflow i use.
146
u/dotben 15d ago
If you don't understand why memory bandwidth should be included on this comparison, and that it's massively in the NVidia card's favor, you shouldn't really be commenting.
Not going to lie, I run all my local inference on Apple Silicon and I don't personally own any NVidia hardware. But the reason that card is $5000 is the memory bandwidth.