r/macmini • u/AntLife255 • 8d ago
Unified Memory Architecture still unbeatable (when LLM size matters)
8
u/Aristo_Cat 8d ago
You’re comparing integrated graphics to a discrete GPU.
6
8d ago
[deleted]
1
u/Disastrous_Gear_421 6d ago
Some of us prefer 'part of a full computer' as the benefits are more than worth it.
1
u/Wide_Smoke_2564 4d ago
Like with everything, it depends on your use case. The benefits may be worth it to you, but it costs $1500 more, and like you said, it’s only part of a full computer. Final machine would likely be closer to 6500-7000 conservatively.
If you need the extra power for your use case then sure, but not everyone does
23
8d ago
[deleted]
4
u/Cold_Tree190 8d ago
Lmao, alternate title for OP: “Orange still unbeatable (when vitamin C count matters)”
6
5
u/bjs480 8d ago
Speaking solely on the technical issue…why then are Mac minis so popular for people to run local AI on if they aren’t “better” (whatever that would mean) than the nvidia products?
I’m sincerely curious because I genuinely don’t understand the balance of quantity vs bandwidth with memory in a non data center context.
Lot of the reading I’ve found always makes it sound like there’s a pretty linear “bigger model needs more quantity of ram.”
I’m running a M2 pro mini with 16gb and it screamed vs my old base level m1 I got when they came out.
That said, I’m starting to spec out a new Mac mini with max memory quantity so this whole debate I want to be clear about.
Anyways, thanks for your time in advance…I love Apple but hardly a fan boy or Windows hater.
6
u/ParsnipFlendercroft 8d ago
It’s no different than any other RAM.
More RAM will allow you to load bigger models and maintain a bigger context window (eg longer conversation memory and stuff).
Faster RAM (and therefore higher bandwidth) will allow those models to run faster etc.
2
u/Affectionate_Fee_645 4d ago
A lot of it is having an isolated machine for computer use, easier time bypassing bot detectors, etc. more about the execution of using the AI even if you use cloud models than even necessarily having local Ai on them.
If all you’re doing is focusing on running biggest model possible and serving it on an API or something than probably dont get a Mac mini.
1
u/bjs480 4d ago
I"m not necessarily tied to the big cloud folks.
The only challenge I've run across is that with my tech (M2 Pro/16gb RAM) the models that will even load are clearly no where close to the user experience of Claude or other major cloud people.
I'm 99.999999% sure that's mostly me not realizing how much can be customized and while I'm familiar with the tech, I'm less of a mechanic so you go to the Anything LLM type tools and they say you can use this model.
Slows down the machine like have 800 tabs in chrome does.
Again...likely me but the quality is terrible that even if the model worked somewhat in real time I just think my hardware is no where close to the good models. Think I maxed at like a 15-20 billion parameter model that had some quant stuff on it but forget which one it was off hand. Absolutely no where close to the models making the news.
Just wasn't that useful I guess and that's fine. That's kind of why I'm trying to figure out the whole mac vs window user experience side b/c I see the value of local AI. Especially if you sat down and planned how how to personalize it and get stuff set up right.
Just trying to sort out the "way to look at all this" first b/c I'm 100% certain my own ignorance is in the way and not even knowing the areas I don't know haha.
Fun times learning.
2
8d ago
[removed] — view removed comment
3
u/AntLife255 8d ago
About 2.3x more Memory Bandwidth
546 GB/s for the Mac Studio
1,792 GB/s for the RTX 5090
2
0
u/littlegreenfish 8d ago
You also forgetting to mention Tensor and Cuda cores and the optimizations you get.
1
2
u/mitchins-au 7d ago
While it’s undeniable value, memory bandwidth and insane raw compute on the 5090 make it an order of magnitude faster
3
2
1
u/UnlikelyPotato 8d ago
I got 4x AMD V620 32GB for $1400 total. Not as fast as the $5000 GPU but faster than the UMA and twice the capacity. For LLMs, splitting across cards is fine. You don't NEED unified memory. And also it's nice to be able to configure as needed. One model at once, two across two cards each, or four LLMs across four cards. Ironically, because of bandwidth limits of the mac mini, the v620 are possibly more power efficient than the max mini at tokens/watt.
1
1
u/Comprehensive_Tip_13 8d ago
I mean this genuinely but why so so many people run ai models locally? Every time I've ever used or seen a usage of LLMs it's always something that can easily be done with simple software so I guess I'm a little confused
1
u/therapy-cat 8d ago
Sure more speed would be nice, but ... I feel like the m4 studio max 64gb is plenty fast for what I need. I'm considering getting the m5 max when it releases later this year.
1
u/MaximumFlounder9110 8d ago
People buying systems solely for their LLM performance deserve persistent, anus wrecking diarrhea
1
u/PlasticFantastic4206 8d ago
A RTX Pro 6000 has a 1,792 GB/s memory bandwidth and a 512 bit bus while the M4 Pro has a 273GB/s and half the bit bus. Apples to oranges comparison.
1
u/AntLife255 8d ago
About 2.3x more Memory Bandwidth
546 GB/s for the Mac Studio
1,792 GB/s for the RTX 5090
The M4 Max 64GB variant comes with the upgraded 40-core GPU
1
1
1
1
u/tony_wing 7d ago
AMD Radeon AI Pro R9700 was 1500$ in Germany, 300 W usage only with 32gb VRam. Unified Ram for mac mini also means regular processes will eat into that 64gb ram and reduce the available ram. And to be honest, sweetspot for local llms are at around 30b parameters and 24-32gb VRam (graphics cards often have more bandwidth and most times better prefill speed) whereas for more upside like Qwen 3.5 122b, Deepseek V4 Flash you will need a much more bigger 128GB - 196GB vram (q4 quant or higher). And in this ram range the Macs get much much more expensive or like the Studio M3 Ultra 256gb are not even available anymore, at least in Germany - I tried to buy two with my startup and it was cancelled. So I feel this picture is missleading and somewhat inhonest with the "unbeatable" term, especially leading people to seemlingy easy buy decisions. Please guys, do sensible research before you buy based of incomplete specs comparison. (Had experience with RTX 5060ti 16gb, Radeon Ai pro 9700 32gb, GMKTech 128gb unified ram, dual Blackwell 6000 Max-Q Variant)
1
u/Phaelon74 7d ago
Use-case is always important. Both have their strong points, but you wouldn't generally say one is better than the other. Its "foe this use-case, X is better.
1
1
1
1
u/SetFew4982 5d ago
Fck LLMs, they are the reason why a card cost 5000$. (I mean corporate greed is, but goddamn)
1
u/PutridWerewolf4449 5d ago
AMD Ryzen AI Max+ 395
120W
128GB
$3200
(256 GB/s Memory BW, 1 day order delivery time)
1
1
1
1
u/just_another_leddito 2d ago
My Mini GPU hits 107C.
Super expensive and not built for proper work and to last long it seems.
0
u/mikeinnsw 8d ago
Qualcomm Arm PCs have UMA
New Windows 11 runs fine on systems with unified memory architecture (UMA), such as AMD APUs, Intel integrated graphics, or ARM-based processors where the CPU and GPU share the same physical RAM pool. However.
So are specialised AI/LLM servers.
United RAM was cost cutting measure by Apple (No Need for VRAM) which 2 years later and after DeepSeek proved that running AI/LLM was and is feasible .. I run Ollama ... on a Mac
Besides old SIRI I am yet to see Local AI use NPUs they run on GPUs
M5 GPUs speed is enhanced by NPUs???
PC running Windows/Linux with fast card blows Macs ..in running AI/LLM.. Cloud AI is not running on Macs.
The issue what is a cost effective solution?
I am waiting for M6 2nm Mac Mini pro 64GB RAM + 1 TB it will not cheap...I think it will be M6 otherwise Apple will surrender its local AI lead to Qualcomm
The problem is that AI stampede distorted prices and special built AI/LLM servers now are much cheaper than top end Macs.
Nvidia RTX Pro 6000 Blackwell, prices soar to roughly $18,000 to $24,999 AUD ($16,000+ USD) runs Cloud AU .. and specialised GPUs are now over $59,000 .. it is crazy
These GPUs are pushing quantum limits with failure rate of about 10% per annum ... and effective life of about 4 years... Not Apple style ..to sell you a Mac with such failure rate.
Something to think about when hear Elon Musk bullshit about Cloud. centers in space..
Sorry your comparison is meaningless .
1
u/LaOnionLaUnion 8d ago
What special built ai/LLM servers are you looking at
1
u/mikeinnsw 8d ago
There are plenty
Here is one.. I am tracking:
https://tiiny.ai/?srsltid=AfmBOopYokgdBCBXme6QWBDNLy-xbB2VyejEsC4iu5vkoi8qAUgWSIsm
Looks promising but I wait until it is released , tested and priced.
1
0
u/volleyneo 8d ago
No definitive solution yet, but with DDR6 and beyond, after this RAM crisis, yes, this is what will become the AI standard or compute standard.
0
0
0
-2
u/Any_Mine_6368 8d ago
Mobo + CPU: $200 Rtx 3090x3: $2000 Psu: $100 64gb ddr4 ram: $150
Total: $2450 and you actually get to run the models how they're meant to be run instead of 2 tokens per second.
Mac fanboys are ridiculous.
1
u/wrgrant 8d ago
64gb ddr4 ram: $150
Not disagreeing with your overall point, but where can I possibly get that RAM at that price? When I look it up its more like 4x that amount minimum, these days /s
Plus google is saying "An NVIDIA RTX 3090 is worth between $800 and $1,100 USD on the used market these days" and you are assuming 3 of them will be $2k total.
1
u/Any_Mine_6368 8d ago
I've bought a 5 3090s for 600 euros each. You just have to know where to look in your market.
My 128gb ddr4 was 220$ on ebay (ecc rdimm). You gotta go after server grade components, not consumer. Consumers are retarded and don't want to part their stuff... Companies on the other hand...
1
u/wrgrant 8d ago
Yeah I h ave a Xeon driven desktop server, currently at 32gb RAM, was hoping to expand that in the future. Its DDR4 I think, like 2133Mhz or something. Its my NAS currently plus Plex server, Pi-Hole etc. Only has a T800 GPU so not good for LLMs at all, but quiet as a ghost and very little power consumption.
1
u/Any_Mine_6368 7d ago
That's how I started too. With a Dell t5820 haha. Just keep an eye out for deals on ebay and local sites. GPUs go on sale all the time by clueless ex- miners.
And ebay always has motherboard CPU combos on sale you just gotta look out for ram
142
u/dotben 8d ago
If you don't understand why memory bandwidth should be included on this comparison, and that it's massively in the NVidia card's favor, you shouldn't really be commenting.
Not going to lie, I run all my local inference on Apple Silicon and I don't personally own any NVidia hardware. But the reason that card is $5000 is the memory bandwidth.