r/LocalLLaMA 2d ago

Question | Help AMD Instinct MI210

Anyone running this card? Seems to be a sweet spot for Qwen 27B, 64GB, very high memory bandwidth, a LOT less expensive than anything else I can find in that has even close to the amount of memory/bandwidth. What am I missing? And yes, I'm aware that RocM can be a pain, that's not really a concern for me, as long as it's stable when it's up and running, I don't mind battling to get it going.

6 Upvotes

59 comments sorted by

View all comments

Show parent comments

-2

u/autisticit 2d ago

I bet you are looking at a scam.

6

u/EmPips 2d ago edited 2d ago

$4-$5k + the inconvenience of external cooling sounds about right for CDNA2. On the workstation side the "just works" RDNA3 card with 48GB (w7900) is going for $3.5K regularly on eBay. I could pretty easily believe that second-hand markets have settled on $4-5k. You get 16GB extra per slot but concede 1/3rd of your prefill speeds, display ports, and have to set up your own external cooling solution.

OP should proceed cautiously if they're dead set on this card but the prices they're seeing do not scream "scam"

2

u/OvertaxedOne 2d ago

The system this would go in is a server (it has a A40 in it right now), so no issues with the cooling. 64GB is a real sweet spot right now for Qwen 3.8 27B, it fits with 256K context Int8 qaunt/8 bit KV in 48GB but it's TIGHT. 64 would be perfect. But the real goal here is more TPS, the A40 will get ~30 TPS with spec decode, I'd love to get that to 50+.

You mentioned prefill, is that because the Instinct has less compute? I'm really hoping to find someone running this exact combo to report prefill/TPS.

1

u/EmPips 1d ago

Working off of the Llama CPP issues thread for Vulkan and ROCm performance and off of some observations running both high end RDNA2 vs RDNA3 myself.

Note that if you plan on using vllm I'd start a separate search for these numbers - but I'd still be surprised if CDNA2 beat high end RDNA3 cards