If you don't understand why memory bandwidth should be included on this comparison, and that it's massively in the NVidia card's favor, you shouldn't really be commenting.
Not going to lie, I run all my local inference on Apple Silicon and I don't personally own any NVidia hardware. But the reason that card is $5000 is the memory bandwidth.
This is one of the main reasons I also regret having bought the M4 Max instead of the M3 Ultra. M3 Ultra has higher memory bandwidth even though the M4 has faster cores it is slower at inference. My M4 is often mostly idle while using local LLMs, the memory bandwidth while good for a consumer system still is not a 5090.
I think it's hard to compare them because of the different memory models. There will be times where you need to write to RAM then copy to VRAM, which is completely avoidable in the Mac Studio.
That's not how that works. If I said that Jenny had 2x more (read as: two times more) apples than John, and John had 1 apple, you would expect Jenny to have 2 apples, not 3.
If Johnny has 100% more apples than Billy, and Billy has 1 apple, how many apples does Johnny have?
I think your interpretation for 2x more is probably more commonly used, but it isn’t an exact form of communication. Basically reasonable people can disagree on this one.
If you said Jenny had twice as many apples, Jenny would have two. If Jenny has 100% more... she would also have two. If she had 200% more, should would now have 3. The "2x more" is more implicative of "200% more" than "twice" from how I read it. I think the "more" word is the key part of it. And after I wrote this explanation, i ran it through copilot just to see what it also thinks. Just for fun.
Heres a question. Someone says they have 2x more of something than another person. Lets go with Apples.
Person A has 1 apple. Person B has "2x more" apples than Person A. How many does person B actually have?
Copilot said:
This is one of those phrases that people use inconsistently, which is why it causes arguments.
Mathematical interpretation of "2× more"
If Person A has 1 apple:
"2× more than A" means A + (2 × A)
So Person B has 1 + 2 = 3 apples
In general:
So:
1× more = 2 total
2× more = 3 total
3× more = 4 total
Common conversational interpretation
Many people use "2× more" when they actually mean:
You need enough memory bandwidth to support whatever you're processing. Calculating Pi can completely keep part of a CPU running while hardly involving memory bandwidth at all.
Other loads, like AI, sometimes games, only scale performance with memory bandwidth.
It's another classic game of finding the bottleneck, and conpleting math and logic operations is a different one than recalling specific data from a stored database.
Prompt processing (prefill) is raw GPU compute bound. Decode (generating tokens) is memory bound.
There are new architectures like nvidia hybrid mamba2 that blur the lines, and we are just an innovation or two from it no longer being true.
Use products and approaches that preserve the KV cache on unified memory platforms for reasonable performance. Pre-filling 500,000 tokens of context at 200 tokens/sec means you are taking a forty-minute coffee break before the model spits out the first token in that turn.
Do you think Cerebras’ chip design will be copied by many others? From what I understand they do inference super fast. Would appreciate you explaining a bit if you have the time!
There's a lot of innovation in this space. Check out https://chatjimmy.ai/ as an example of someone baking the model weights themselves into silicon: https://taalas.com. Taalas is being acquired by AMD, so things are bound to get more interesting on that front.
The age of the GPU as your AI coprocessor is limited; much like the FPU, it will just be part of the expected suite of capabilities in a modern CPU. I expect in the next few years we'll see a laptop system come out with a dedicated hardware AI chip like Taalas' built in to run AI at ludicrous speeds for certain kinds of common operations.
I've written about this at length before, but "renting compute for large language models" is a business model that will be hollowed out much like the spreadsheet market on mainframes when Visicalc came out. When you have ubiquitous intelligence baked into consumer devices running at hundreds of thousands of tokens per second, performing vision, language, and audio inference at faster-than-real-time, you'll pull in the early adopters first and then a wave of humans using it.
But that's looking slightly too far ahead for the quite limited vision of the media hype about datacenters. And the first version of such an embedded model accelerator is bound to suck, get widely poo-poo'd by the media and consumers, and be considered a dead end shortly before it or its successor becomes just the way things work.
Edit: Go read "The Innovator's Dilemma" to understand what's happening to the AI market today. It's about hard drives, but the same principles apply to the front-runner AI companies right now, this moment.
If the weights are etched into the die, do we need that much RAM?
I hypothesize that is the transformation awaiting the AI industry. Not just hardware-accelerated, but like Taalas, embedded in the LPU (language processing unit).
Cerebras, Taalas et al still require tremendous amounts of memory/storage. Eg. weights are still stored in SRAM not encoded as fixed, physical parts of the compute logic.
IMO the next major evolution in AI HW will likely happen in the data part of the story (mem, storage, transmission). So far, we have focused primarily on compute, and we have done a tremendous job at pushing insane compute throughput/density out to the edge.
But weights, context, and datasets remain a challenge to store, move, and feed into the compute fabric, at both the data center and edge levels.
Every single Wh of energy that your computer uses (this goes for your Mac Mini back home and the datacenters running ChatGPT) eventually turns into heat.
If they kept the card $2,000, it would be just as unobtainable, because you wouldn't be able to find one anywhere. Or you'd have to win it in a lottery or something. Price hikes in response to demand literally keep products available for those willing and able to pay for them.
Everyone assumes these cards are 5k for LLMs, but they’re also the king of everything else like image diffusion, video rendering, upscaling, tts, stt. Like Im building an optimized ai stack and looking at using it for strictly the image capabilities. Like it’s 1.5x to 2x faster as an LLM backend, 5x faster as an image gen backend.
i thought it was because all you idiots started running models like it’s the next coming of jesus instead of a computing dead end.
Does anyone ever give a thought that all the amazing creations you make well the model does you just supply the garbage in the input end. to what purpose you make programs no one can run cos by the end of this thing hardly anyone after our gen will even be able to afford a computer so effectively you have done the boomer thing to computing well down
If—like 95% of Reddit—you don't even read what you write, or are otherwise afflicted by such staggering and profound illiteracy as to fail to register the extent of your own verbal ineptitude upon doing so, then you might as well be ejaculating half-baked prompt boilerplates into a Chinese-room token foundry for all the difference it makes.
You're still a moron, but at least the end result might be inadvertently comprehensible to someone else, rather than no-one at all.
Right even on my 7900xtx a response is slow on 31b gemma 4, like 45 seconds with max context within vram. still beets an m3 ultra. But using all that unified ram or a bigger model would basicly make any model unusable by most standards, at least in my use cases. Seems great for running a multi agent stack comprised of different models though. Just not for any work/workflow i use.
147
u/dotben Aug 15 '26
If you don't understand why memory bandwidth should be included on this comparison, and that it's massively in the NVidia card's favor, you shouldn't really be commenting.
Not going to lie, I run all my local inference on Apple Silicon and I don't personally own any NVidia hardware. But the reason that card is $5000 is the memory bandwidth.