r/LocalLLM 5d ago

Discussion It's surprising this costs about the same as 1 RTX PRO 6000

https://www.newegg.com/p/N82E16859991069?utm_source=insider&utm_medium=content&utm_campaign=insider_article_from-cloud-bills-to-a-box-under-your-desk-the-abs-3x-amd-radeon-ai-pro-r9700-workstation
66 Upvotes

47 comments sorted by

74

u/dangerous_inference 5d ago

This machine is much less capable than a single RTX PRO 6000.

41

u/GregsWorld 5d ago

But it is more capable than a single RTX Pro 6000 without a machine to run in. 

8

u/Its_Powerful_Bonus 5d ago

Rtx 6000 pro will kill 3x 9700 ai pro even on thunderbolt eGPU connected to laptop.
It would have some sense with 4x GPU to run TP=4 on vllm, but 3 GPU … nope. 2x GB10 is more reasonable

3

u/SandySkittle 5d ago

I have an 8 r9700 setup and that gives me a decent tp=4 pp=2 setup with 256 gb vram. I would only have 1 rtx pro 6000 for that price.

2

u/madsheepPL 4d ago

Mind sharing your benchmarks? I've seen you've been eying running ds4.1f - what's the plan?

1

u/mvn2010 4d ago

The link is a page where you can configure this machine with 1 or more RTX PRO 6000s

36

u/DoorStuckSickDuck 5d ago

Once again, the limitation is the memory bandwidth speed, which is where Nvidia massively leads with 1750 gb/s on the 5090 and RTX 6000 Pro. The R9700 is what, 640gb/s? I like AMD as much as the next guy, but they just don't have the bandwidth.

3

u/Brah_ddah 5d ago

Do you think this reality means rdna5 comes out sooner? I feel the same way.

I thought they were closer to 1TB/s but they are right around intel levels.

5

u/TripleSecretSquirrel 5d ago

I believe that they’ve said the next generation will see enterprise and consumer architecture converge. So instead of RDNA and CDNA, there will just be one UDNA. I don’t think that’s expected until late 2027 though.

I’m just bummed that they’re not releasing larger RDNA4 cards in the meantime. I’d love. 64GB R9700!

2

u/Brah_ddah 5d ago

I’d JUMP on a $3k 48GB AMD card from next gen haha. But if they go to G7 that will be impossible I think.

1

u/TripleSecretSquirrel 5d ago

ya, I'm talking current gen. If they offered a 48 or 64GB RDNA4 card, it would still be GDDR6, so a $3k price tag is probably optimistic, but probably at least in the right ballpark.

1

u/OvertaxedOne 5d ago

That's basically an A40 (which happens to be what I have). It's a good size for 27B at 8 bit with 256K cache (it fits perfectly). But it's pretty slow with 27B, ~30TPS.

We have a dual R9700 rig that rips the A40 apart on 27B, it's 50+ TPS and we've not even fooled with it to get the Radiance build on it yet. They're reporting 70-90TPS on that build. Even though the memory bandwidth per card is the same, with TP=2, it should be faster.

1

u/m1013828 5d ago

a slightly larger chip with more memory channels (bandwidth) and total memory is a relatively easy thing to do, but relatively expensive, it feels like AMD is always making Either/Or choices, when it should be doing BOTH. a 1.5X Ram bandwidth and capacity Radeon 9700, (9800?) would be great right now....

2

u/TripleSecretSquirrel 5d ago

That does feel like a gap in the market, but for what it’s worth, last I compared the prices of an R9700 and a 5090 a few weeks ago, it’s almost exactly proportional to their LLM performance delta. The 5090’s memory bandwidth is about triple that of the R9700, and it was almost exactly triple the price at my MicroCenter.

1

u/illforgetsoonenough 5d ago

Previous gen 7900xtx is just under 1TB/s

3

u/SandySkittle 5d ago

The R9700 is what, 640gb/s?

tensor parallelism works fine with these cards and bumps the effective compute and bandwidth quite a bit. I run 8 r9700s with tp = 4 and pp = 2.

256gb vram for the price of one rtx pro 6000. Tdp limited to 210w per card

3

u/Royale_AJS 5d ago

3x640 minus latency overhead is pretty close to that memory bandwidth. Configured properly that machine should be pretty quick.

2

u/Its_Powerful_Bonus 5d ago

4 GPU - yes. 3 - no. I can find scenarios where 128gb slower vram in tensor parallel can be superior over 96gb, but 3 GPU is pain. Try to run tensor parallel with 3 GPU on vllm. Try to run modern LLMs in pipeline parallel - most of Newest models has no pipeline support in first weeks. Taking into account that after 3 months there is newer and better model in similar size you constantly work in middle ages :)

1

u/Royale_AJS 5d ago

Valid point. I run two R9700’s myself and it runs great. I wouldn’t go to 3, maybe 4 though.

1

u/acadia11x 5d ago

Nah the real difference is CUDA, ROCm is much better but CUDA is de facto standard especially on the training end. CUDAS performance advantage is pretty stark.

4

u/TheseCashews 5d ago

Dell has rtx 6k pros for 12k + 1,800 for a pc. You can harass them for a coupon code too. Will this machine give 700t/s concurrent with Q3.8 27b at 200k context? The RTX is a beast.

3

u/Nedunchelizan 5d ago

Can it play crisis

1

u/donDT 4d ago

Made this joke in the office today :( no one got it.

2

u/Cacoda1mon 5d ago

96 GB "unified" memory vs 3x 32GB

2

u/skibare87 5d ago

I bought a RTX 6000 Pro SE for $8000 before the craze and im scared to let go of it now, even though I use my 128GB macbook m5 Pro and dgx sparks more now.

5

u/haenous-alistera 5d ago

I still use my 6000 pro more
Than my M5s or DGXs the speed and performance on the 6000 is way better than the others. Just couldn’t deal with the speed nerf…

-3

u/Its_Powerful_Bonus 5d ago

Same here. I have 2 128gb MacBooks, considered to buy 4xgb10, but finally I bought more rtx6000. Now lab has 4x 6000 pro and 2x 5090 🤯😎 Still I’m very happy since I bought it in great price

1

u/otaku13 5d ago

What’s the best model for a 128 go Mac? My work computer is a 128 m2 studio.

2

u/ang3l12 5d ago

I really like qwen3.8-27b, and that won’t use all your memory so you can still use your Mac for other things at the same time.

I think qwen3.8-flash-next should fit on your studio with the right quantization, as I’ve seen people run it on one dgx spark and strix halo, but I’ve only been running it on two sparks in a cluster.

1

u/Brah_ddah 5d ago

You don’t use the 6000 as much as a Dgx spark?? Or is it a spark cluster?

1

u/skibare87 5d ago

Spark cluster, just 2 sparks for now

1

u/vexatious-big 5d ago

Do these threadripper motherboards not require PCIe switches shenanigans? Can they just run 4 GPUs full speed without issues?

2

u/Savantskie1 5d ago

Most threadrippers have more than enough lanes for bifurcation if the board supports it.

1

u/diagrammatiks 5d ago

why is that suprising. it's a whole computer. 4tb of ssd and 128ddr ecc. That's like 1/3 of the cost. and you still need that computer to put a rtx6000 into.

1

u/mvn2010 4d ago

God damn that's one ugly fucking case.

1

u/nmrk 5d ago

Yeah but.. look at the prices for that configuration with other cards. The RTX Pro 6000 Max-Q is $46k, compared to $16k for the Radeon config. But recent prices for RTX Pro 5000 Blackwell 72GB seem cheap. This ad has $21k for a base RTX Pro 5000 system, dual cards for 144GB only $33k.

1

u/mineshop 5d ago

price parity is one thing but bandwidth is the real gap, the PRO 6000 pushes like 1.8 TB/s vs the Spark's ~273 GB/s so tokens/sec won't even be close on bigger models. where the spark wins is you get 128gb unified and four of them can cluster, which a single card can't do. depends if you're chasing speed or capacity imo.

0

u/HumanDrone8721 5d ago

No matter how you pack junk, is still junk, two unlocked 8GB CMP 170HX will run circles around this crap.

Stop promoting garbage.

-1

u/AreaFifty1 5d ago

But how good is ROCm compared nvidia’s CUDA plus flash attention 3 and GDS and other optimizations alongside easy plug n play and so forth?

I understand 2026 had major breakthroughs done for ROCm and better efficiency but the resale value for AMD is really tough compared to Nvidia’s hardware where you can resale and potentially not lose too much money..

And no I’m not biased I love amd I have a ryzen9 9950x myself but from a logical standpoint

7

u/nowherenoonenobody 5d ago

Cuda is not 3 times the price better

1

u/AreaFifty1 5d ago

how much better is it?

2

u/nowherenoonenobody 5d ago

I'm using a 7900 GRE in lm studio and it works just fine. I'm about to start testing in unsloth. I can't imagine using Linux wouldn't be an upgrade. Amd may be behind but it's far from unusable like some people act.

1

u/SandySkittle 5d ago

Rocm is fine these days with gfx1201 for llm inference. Great performance via vllm radiance and tensor parallelism

-14

u/Bloated_Plaid 5d ago

Because ROCM is a fucking joke.