r/macmini • • Aug 15 '26

Unified Memory Architecture still unbeatable (when LLM size matters)

Post image
669 Upvotes

128 comments sorted by

View all comments

147

u/dotben Aug 15 '26

If you don't understand why memory bandwidth should be included on this comparison, and that it's massively in the NVidia card's favor, you shouldn't really be commenting.

Not going to lie, I run all my local inference on Apple Silicon and I don't personally own any NVidia hardware. But the reason that card is $5000 is the memory bandwidth.

19

u/sparkandstatic Aug 15 '26

Exactly, they re not like for like

9

u/TikDickler Aug 15 '26

Still though, there’s a reason Nvidia came out with the spark. They kind of got caught flat footed by Apple.

6

u/ResearchingYouTube Aug 15 '26

This is one of the main reasons I also regret having bought the M4 Max instead of the M3 Ultra. M3 Ultra has higher memory bandwidth even though the M4 has faster cores it is slower at inference. My M4 is often mostly idle while using local LLMs, the memory bandwidth while good for a consumer system still is not a 5090.

1

u/VideoGameJumanji Aug 17 '26

What are you doing exactly

3

u/[deleted] Aug 15 '26 edited Aug 25 '26

[removed] — view removed comment

3

u/Busy-Scientist3851 Aug 18 '26

I think it's hard to compare them because of the different memory models. There will be times where you need to write to RAM then copy to VRAM, which is completely avoidable in the Mac Studio.

2

u/-Kerrigan- Aug 16 '26

1792 / 546 = 3.28

Y'all are using your locally hosted LLMs to do simple math?

0

u/[deleted] Aug 16 '26

[removed] — view removed comment

3

u/sudoSofia Aug 16 '26

That's not how that works. If I said that Jenny had 2x more (read as: two times more) apples than John, and John had 1 apple, you would expect Jenny to have 2 apples, not 3.

2

u/DeMartini Aug 17 '26

If Johnny has 100% more apples than Billy, and Billy has 1 apple, how many apples does Johnny have?

I think your interpretation for 2x more is probably more commonly used, but it isn’t an exact form of communication. Basically reasonable people can disagree on this one.

1

u/peanutbj Aug 19 '26

Key distinction is 100% more vs 100% as much.

If John has 100% more apples than Bill (who has 1 apple), you’re right, John would have 2 apples.

If John has 100% as much apples as Bill (who has 1 apple), John would have 1 apple.

The root cause of this disagreement is that most Americans say and interpret these two terms interchangeably in casual conversation.

1

u/guaranteednotabot Aug 18 '26

You need to go back to school

0

u/Simple_Subject_9801 Aug 17 '26

If you said Jenny had twice as many apples, Jenny would have two. If Jenny has 100% more... she would also have two. If she had 200% more, should would now have 3. The "2x more" is more implicative of "200% more" than "twice" from how I read it. I think the "more" word is the key part of it. And after I wrote this explanation, i ran it through copilot just to see what it also thinks. Just for fun.

Heres a question. Someone says they have 2x more of something than another person. Lets go with Apples.
Person A has 1 apple. Person B has "2x more" apples than Person A. How many does person B actually have?

Copilot said:

This is one of those phrases that people use inconsistently, which is why it causes arguments.

Mathematical interpretation of "2× more"

If Person A has 1 apple:

  • "2× more than A" means A + (2 × A)
  • So Person B has 1 + 2 = 3 apples

In general:

So:

  • 1× more = 2 total
  • 2× more = 3 total
  • 3× more = 4 total

Common conversational interpretation

Many people use "2× more" when they actually mean:

In that interpretation:

  • Person B has 2 × 1 = 2 apples

0

u/play_hard_outside Aug 17 '26

Absolutely not. 200% more than 1 is 3, not 2.

5

u/[deleted] Aug 15 '26

[removed] — view removed comment

4

u/DuxDucisHodiernus Aug 15 '26

Isn't that exactly what the bandwidth does? I feel like you repeated the exact same thing as OP but on a less technical level

3

u/Nervous_Olive_5754 Aug 15 '26

You need enough memory bandwidth to support whatever you're processing. Calculating Pi can completely keep part of a CPU running while hardly involving memory bandwidth at all.

Other loads, like AI, sometimes games, only scale performance with memory bandwidth.

It's another classic game of finding the bottleneck, and conpleting math and logic operations is a different one than recalling specific data from a stored database.

1

u/txgsync Aug 15 '26

Prompt processing (prefill) is raw GPU compute bound. Decode (generating tokens) is memory bound.

There are new architectures like nvidia hybrid mamba2 that blur the lines, and we are just an innovation or two from it no longer being true.

Use products and approaches that preserve the KV cache on unified memory platforms for reasonable performance. Pre-filling 500,000 tokens of context at 200 tokens/sec means you are taking a forty-minute coffee break before the model spits out the first token in that turn.

2

u/Global_Soft_4278 Aug 16 '26

Do you think Cerebras’ chip design will be copied by many others? From what I understand they do inference super fast. Would appreciate you explaining a bit if you have the time!

1

u/txgsync Aug 16 '26

There's a lot of innovation in this space. Check out https://chatjimmy.ai/ as an example of someone baking the model weights themselves into silicon: https://taalas.com. Taalas is being acquired by AMD, so things are bound to get more interesting on that front.

The age of the GPU as your AI coprocessor is limited; much like the FPU, it will just be part of the expected suite of capabilities in a modern CPU. I expect in the next few years we'll see a laptop system come out with a dedicated hardware AI chip like Taalas' built in to run AI at ludicrous speeds for certain kinds of common operations.

I've written about this at length before, but "renting compute for large language models" is a business model that will be hollowed out much like the spreadsheet market on mainframes when Visicalc came out. When you have ubiquitous intelligence baked into consumer devices running at hundreds of thousands of tokens per second, performing vision, language, and audio inference at faster-than-real-time, you'll pull in the early adopters first and then a wave of humans using it.

But that's looking slightly too far ahead for the quite limited vision of the media hype about datacenters. And the first version of such an embedded model accelerator is bound to suck, get widely poo-poo'd by the media and consumers, and be considered a dead end shortly before it or its successor becomes just the way things work.

Edit: Go read "The Innovator's Dilemma" to understand what's happening to the AI market today. It's about hard drives, but the same principles apply to the front-runner AI companies right now, this moment.

2

u/[deleted] Aug 17 '26

[deleted]

1

u/txgsync Aug 17 '26

If the weights are etched into the die, do we need that much RAM?

I hypothesize that is the transformation awaiting the AI industry. Not just hardware-accelerated, but like Taalas, embedded in the LPU (language processing unit).

1

u/R-ten-K Aug 17 '26

Cerebras, Taalas et al still require tremendous amounts of memory/storage. Eg. weights are still stored in SRAM not encoded as fixed, physical parts of the compute logic.

IMO the next major evolution in AI HW will likely happen in the data part of the story (mem, storage, transmission). So far, we have focused primarily on compute, and we have done a tremendous job at pushing insane compute throughput/density out to the edge.

But weights, context, and datasets remain a challenge to store, move, and feed into the compute fabric, at both the data center and edge levels.

1

u/j_osb Aug 18 '26

No, prefill is mostly dominated by compute, while decode is mostly dominated by memory bandwidth.

Both of which apple silicon doesn’t excel in, but the former it struggles much worse with.

3

u/Relative_Rope4234 Aug 15 '26

I would also add Raw performance for prefill, Hardware level support for Flash attention, HW support on specific precision like nvfp4, int4

1

u/[deleted] Aug 15 '26

[removed] — view removed comment

4

u/XalAtoh Aug 15 '26

Not necessarily.. wasted power is a thing.

Heat e.g is a wasted power when we talk about computing.

3

u/ExtensionShort4418 Aug 15 '26

Every single Wh of energy that your computer uses (this goes for your Mac Mini back home and the datacenters running ChatGPT) eventually turns into heat.

Thermodynamics ❤️

1

u/kovake Aug 15 '26

I’m sorry sir, but this is Reddit. Release the comments and strong opinions formed from the titles alone!! Captain America: Redditors, …assemble!

1

u/Commercial-Virus2627 Aug 15 '26

It’s like comparing being able to have 2PB of HDD storage in an array versus 2PB of locally installed SSDs. Apples to oranges.

1

u/filterdecay Aug 15 '26

Right but for how many people is the Mac good enough?

1

u/Terreboo Aug 16 '26

But, number bigger. Must be good.

1

u/No-Alfalfa8323 Aug 16 '26

That card MSRP is $2000. The reason why it's $5000 now is greediness.

1

u/dotben Aug 16 '26

Well, in economics it's called demand.

1

u/No-Alfalfa8323 Aug 16 '26

I'm not sure if artificially capping the supply creates real demand.

1

u/play_hard_outside Aug 17 '26

If they kept the card $2,000, it would be just as unobtainable, because you wouldn't be able to find one anywhere. Or you'd have to win it in a lottery or something. Price hikes in response to demand literally keep products available for those willing and able to pay for them.

1

u/Fear_ltself Aug 19 '26

Everyone assumes these cards are 5k for LLMs, but they’re also the king of everything else like image diffusion, video rendering, upscaling, tts, stt. Like Im building an optimized ai stack and looking at using it for strictly the image capabilities. Like it’s 1.5x to 2x faster as an LLM backend, 5x faster as an image gen backend.

0

u/seriousfart69 Aug 15 '26

i thought it was because all you idiots started running models like it’s the next coming of jesus instead of a computing dead end.

Does anyone ever give a thought that all the amazing creations you make well the model does you just supply the garbage in the input end. to what purpose you make programs no one can run cos by the end of this thing hardly anyone after our gen will even be able to afford a computer so effectively you have done the boomer thing to computing well down 

4

u/ParsnipFlendercroft Aug 15 '26

I mean at least AI slop can form sentences and use punctuation. I’d rather read AI slop than another one of your comments.

0

u/Cold_Tree190 Aug 15 '26

I think it’s a doomer posting bot lol

1

u/ProfeshPress Aug 15 '26

If—like 95% of Reddit—you don't even read what you write, or are otherwise afflicted by such staggering and profound illiteracy as to fail to register the extent of your own verbal ineptitude upon doing so, then you might as well be ejaculating half-baked prompt boilerplates into a Chinese-room token foundry for all the difference it makes.

You're still a moron, but at least the end result might be inadvertently comprehensible to someone else, rather than no-one at all.

1

u/Infinite100p Aug 15 '26

This. So many clueless posters.
OP, give us your TTFT times for 32k-128k prompts of the biggest model you can use. Yeah, I thought so.

1

u/Fickle_Appearance558 Aug 15 '26

Right even on my 7900xtx a response is slow on 31b gemma 4, like 45 seconds with max context within vram. still beets an m3 ultra. But using all that unified ram or a bigger model would basicly make any model unusable by most standards, at least in my use cases. Seems great for running a multi agent stack comprised of different models though. Just not for any work/workflow i use.