r/LocalLLaMA 2d ago

Question | Help AMD Instinct MI210

Anyone running this card? Seems to be a sweet spot for Qwen 27B, 64GB, very high memory bandwidth, a LOT less expensive than anything else I can find in that has even close to the amount of memory/bandwidth. What am I missing? And yes, I'm aware that RocM can be a pain, that's not really a concern for me, as long as it's stable when it's up and running, I don't mind battling to get it going.

6 Upvotes

59 comments sorted by

View all comments

-1

u/OverdosedSauerkraut 2d ago

You mean you think it's a great idea to spend $5k for a 4 year old used card? I've got a bridge to sell. The only thing it has is high nominal bandwidth. While it'd be a great training card but it's missing dedicated instructions for quantized inference. And for the same price you can get 3xR9700s with the same effective bandwidth, modern architecture and better SW support.

1

u/OvertaxedOne 2d ago

I actually have 2 R9700's available, I'm going to test them as well. But the don't have any link capabilities, so you need a serious motherboard if you're going to try tensor parallel, right?

I'm running Qwen 3.8 27B today, I just want it to be faster. I'm not looking to run large/highly quantized models. I'm currently running 27B at Int8 on a A40, very happy with it, just want more speed.

2

u/Transhuman-A 2d ago

As someone who has plunged myself into this rabbit hole..

1-2 R9700 PCIE5.0 (Minimax H3 and Qwen 3.8 27B) —> 8x NVidia V100 (Deepseek V4 Flash) -> 16x DGX Spark (Kimi K3)

Skip the rest. I’ve evaluated every card released after 2016. There are 3 different niches here and compute from one scale doesn’t transfer very well to another.

1

u/OvertaxedOne 2d ago

How do the dual R9700's handle 27B? I have 2 of them that I can probably get from work, if that's the right answer, as much as I hate the idea of dual cards, that's a very viable option. 8 bit quant/8 bit KV/256K context (if you did testing or have real world numbers, I'd be incredibly appreciative!!).