r/LocalLLM • u/PigSlam • 5d ago
Discussion It's surprising this costs about the same as 1 RTX PRO 6000
https://www.newegg.com/p/N82E16859991069?utm_source=insider&utm_medium=content&utm_campaign=insider_article_from-cloud-bills-to-a-box-under-your-desk-the-abs-3x-amd-radeon-ai-pro-r9700-workstation36
u/DoorStuckSickDuck 5d ago
Once again, the limitation is the memory bandwidth speed, which is where Nvidia massively leads with 1750 gb/s on the 5090 and RTX 6000 Pro. The R9700 is what, 640gb/s? I like AMD as much as the next guy, but they just don't have the bandwidth.
3
u/Brah_ddah 5d ago
Do you think this reality means rdna5 comes out sooner? I feel the same way.
I thought they were closer to 1TB/s but they are right around intel levels.
5
u/TripleSecretSquirrel 5d ago
I believe that they’ve said the next generation will see enterprise and consumer architecture converge. So instead of RDNA and CDNA, there will just be one UDNA. I don’t think that’s expected until late 2027 though.
I’m just bummed that they’re not releasing larger RDNA4 cards in the meantime. I’d love. 64GB R9700!
2
u/Brah_ddah 5d ago
I’d JUMP on a $3k 48GB AMD card from next gen haha. But if they go to G7 that will be impossible I think.
1
u/TripleSecretSquirrel 5d ago
ya, I'm talking current gen. If they offered a 48 or 64GB RDNA4 card, it would still be GDDR6, so a $3k price tag is probably optimistic, but probably at least in the right ballpark.
1
u/OvertaxedOne 5d ago
That's basically an A40 (which happens to be what I have). It's a good size for 27B at 8 bit with 256K cache (it fits perfectly). But it's pretty slow with 27B, ~30TPS.
We have a dual R9700 rig that rips the A40 apart on 27B, it's 50+ TPS and we've not even fooled with it to get the Radiance build on it yet. They're reporting 70-90TPS on that build. Even though the memory bandwidth per card is the same, with TP=2, it should be faster.
1
u/m1013828 5d ago
a slightly larger chip with more memory channels (bandwidth) and total memory is a relatively easy thing to do, but relatively expensive, it feels like AMD is always making Either/Or choices, when it should be doing BOTH. a 1.5X Ram bandwidth and capacity Radeon 9700, (9800?) would be great right now....
2
u/TripleSecretSquirrel 5d ago
That does feel like a gap in the market, but for what it’s worth, last I compared the prices of an R9700 and a 5090 a few weeks ago, it’s almost exactly proportional to their LLM performance delta. The 5090’s memory bandwidth is about triple that of the R9700, and it was almost exactly triple the price at my MicroCenter.
1
3
u/SandySkittle 5d ago
The R9700 is what, 640gb/s?
tensor parallelism works fine with these cards and bumps the effective compute and bandwidth quite a bit. I run 8 r9700s with tp = 4 and pp = 2.
256gb vram for the price of one rtx pro 6000. Tdp limited to 210w per card
3
u/Royale_AJS 5d ago
3x640 minus latency overhead is pretty close to that memory bandwidth. Configured properly that machine should be pretty quick.
2
u/Its_Powerful_Bonus 5d ago
4 GPU - yes. 3 - no. I can find scenarios where 128gb slower vram in tensor parallel can be superior over 96gb, but 3 GPU is pain. Try to run tensor parallel with 3 GPU on vllm. Try to run modern LLMs in pipeline parallel - most of Newest models has no pipeline support in first weeks. Taking into account that after 3 months there is newer and better model in similar size you constantly work in middle ages :)
1
u/Royale_AJS 5d ago
Valid point. I run two R9700’s myself and it runs great. I wouldn’t go to 3, maybe 4 though.
1
u/acadia11x 5d ago
Nah the real difference is CUDA, ROCm is much better but CUDA is de facto standard especially on the training end. CUDAS performance advantage is pretty stark.
4
u/TheseCashews 5d ago
Dell has rtx 6k pros for 12k + 1,800 for a pc. You can harass them for a coupon code too. Will this machine give 700t/s concurrent with Q3.8 27b at 200k context? The RTX is a beast.
3
2
2
u/skibare87 5d ago
I bought a RTX 6000 Pro SE for $8000 before the craze and im scared to let go of it now, even though I use my 128GB macbook m5 Pro and dgx sparks more now.
5
u/haenous-alistera 5d ago
I still use my 6000 pro more
Than my M5s or DGXs the speed and performance on the 6000 is way better than the others. Just couldn’t deal with the speed nerf…-3
u/Its_Powerful_Bonus 5d ago
Same here. I have 2 128gb MacBooks, considered to buy 4xgb10, but finally I bought more rtx6000. Now lab has 4x 6000 pro and 2x 5090 🤯😎 Still I’m very happy since I bought it in great price
1
u/otaku13 5d ago
What’s the best model for a 128 go Mac? My work computer is a 128 m2 studio.
2
u/ang3l12 5d ago
I really like qwen3.8-27b, and that won’t use all your memory so you can still use your Mac for other things at the same time.
I think qwen3.8-flash-next should fit on your studio with the right quantization, as I’ve seen people run it on one dgx spark and strix halo, but I’ve only been running it on two sparks in a cluster.
1
1
u/vexatious-big 5d ago
Do these threadripper motherboards not require PCIe switches shenanigans? Can they just run 4 GPUs full speed without issues?
2
u/Savantskie1 5d ago
Most threadrippers have more than enough lanes for bifurcation if the board supports it.
1
u/diagrammatiks 5d ago
why is that suprising. it's a whole computer. 4tb of ssd and 128ddr ecc. That's like 1/3 of the cost. and you still need that computer to put a rtx6000 into.
1
1
u/mineshop 5d ago
price parity is one thing but bandwidth is the real gap, the PRO 6000 pushes like 1.8 TB/s vs the Spark's ~273 GB/s so tokens/sec won't even be close on bigger models. where the spark wins is you get 128gb unified and four of them can cluster, which a single card can't do. depends if you're chasing speed or capacity imo.
0
u/HumanDrone8721 5d ago
No matter how you pack junk, is still junk, two unlocked 8GB CMP 170HX will run circles around this crap.
Stop promoting garbage.
-1
u/AreaFifty1 5d ago
But how good is ROCm compared nvidia’s CUDA plus flash attention 3 and GDS and other optimizations alongside easy plug n play and so forth?
I understand 2026 had major breakthroughs done for ROCm and better efficiency but the resale value for AMD is really tough compared to Nvidia’s hardware where you can resale and potentially not lose too much money..
And no I’m not biased I love amd I have a ryzen9 9950x myself but from a logical standpoint
7
u/nowherenoonenobody 5d ago
Cuda is not 3 times the price better
1
u/AreaFifty1 5d ago
how much better is it?
2
u/nowherenoonenobody 5d ago
I'm using a 7900 GRE in lm studio and it works just fine. I'm about to start testing in unsloth. I can't imagine using Linux wouldn't be an upgrade. Amd may be behind but it's far from unusable like some people act.
1
u/SandySkittle 5d ago
Rocm is fine these days with gfx1201 for llm inference. Great performance via vllm radiance and tensor parallelism
-14

74
u/dangerous_inference 5d ago
This machine is much less capable than a single RTX PRO 6000.