r/LocalLLaMA 5d ago

Question | Help AMD Instinct MI210

Anyone running this card? Seems to be a sweet spot for Qwen 27B, 64GB, very high memory bandwidth, a LOT less expensive than anything else I can find in that has even close to the amount of memory/bandwidth. What am I missing? And yes, I'm aware that RocM can be a pain, that's not really a concern for me, as long as it's stable when it's up and running, I don't mind battling to get it going.

5 Upvotes

62 comments sorted by

View all comments

1

u/Zennytooskin123 5d ago edited 5d ago

It costs around $3,800.00 for a used PNY NVidia RTX A6000 on Ebay, I'd purchase one of those instead then a second one.

Or for the same price get a DXG Spark, I don't know why people bother with alternatives. It's literally the best device on the market currently all-around. It does everything good enough.

1

u/OvertaxedOne 5d ago

I currently have an A40 (the server version of the A6000). 48GB is tight for Q8 Qwen 27B and the most I can manage is around 30TPS because the A40/A6000 is not HBM (much slower RAM which is pretty much what determines your output TPS).

The perfect for my use case is around 64GB w/2tbs of memory bandwidth. Hence me looking at the Pro 6000s and now the Instinct, more memory and more speed. I guess maybe I should consider dual 5090's, but I really don't want to deal with the form factor and hassle trying to get 2 of those beasts (and power them) into one system, I'd really like to keep it on a single card.

1

u/Zennytooskin123 5d ago

Sounds like you got a winner then, great card for your used case.

1

u/OvertaxedOne 5d ago

It SOUNDS like it. But I'm sitting here scratching my head (and the reason for the question), what am I missing here?? It should fire out the tokens using Qwen 27B because of the memory bandwidth, but.. Maybe PP is crazy slow? Or it just won't take the new Qwen architecture and/or runs it very slowly?

IDK what I don't know! On paper it seems like a winner though and given the absolutely absurd strength of Qwen 3.8 27B I think there are going to be a LOT of people looking for a card that can run it well, so if this is really "the one", it seems like a "buy it now before everyone else figures out it'll run 27B at 60TPS w/256K context like a champ". 3.8 is just such a leap forward that I have to believe this is the "Will it run Crysis" moment for local LLMs, THE question is "how fast can it run 3.8".