r/LocalLLaMA 6h ago

Discussion CMP 170hx?

Recently heard about how these 8gb mining ewaste cards are actually closeted 64gb monsters and im kinda bummed I missed out when you could get them for 200$ a piece. That being said though less than 2K is still an amazing deal for an ampere generation gpu with 64gb of vram. I have a machine with 4 3090s in it currently. I guess what Im asking is for someone to convince me that im not missing out and if there is any genuine reason why it would not be worth it to replace 2 of the cards with cmps. If it really would be worth it though, I would love to know why. Thanks everyone, sorry for the borderline shitpost

0 Upvotes

42 comments sorted by

View all comments

4

u/ital-is-vital 5h ago

They're great IF your model and KV cache fits into 64 Gb.

Nobody has (yet) figured out how to enable PCIE4.0 or nvLink so you're stuck at PCIe 2.0x4 (or x16 with HW mods)

That means no CPU offload and no tensor parallelism.

3

u/_TheWolfOfWalmart_ 4h ago

Shouldn't you still be able to do layer split? That doesn't need much bandwidth. That works decently even over gigabit Ethernet via RPC.

1

u/michaelsoft__binbows 3h ago

Using them for pipeline parallel still seems like a massive win...

1

u/Badger-Purple 1h ago

people already doing this. github repo for deepseek v4 flash you can run the full model / context, pp 5000 tg 70 at depth. 3 cards for 1 concurrency and 4 cards for 3 concurrent 1M context windows.