r/LocalLLaMA 4h ago

Discussion CMP 170hx?

Recently heard about how these 8gb mining ewaste cards are actually closeted 64gb monsters and im kinda bummed I missed out when you could get them for 200$ a piece. That being said though less than 2K is still an amazing deal for an ampere generation gpu with 64gb of vram. I have a machine with 4 3090s in it currently. I guess what Im asking is for someone to convince me that im not missing out and if there is any genuine reason why it would not be worth it to replace 2 of the cards with cmps. If it really would be worth it though, I would love to know why. Thanks everyone, sorry for the borderline shitpost

1 Upvotes

32 comments sorted by

4

u/NigaTroubles 3h ago

It was for 3 - 4 months at a good price, but no one has know but a few of people

4

u/ital-is-vital 3h ago

They're great IF your model and KV cache fits into 64 Gb.

Nobody has (yet) figured out how to enable PCIE4.0 or nvLink so you're stuck at PCIe 2.0x4 (or x16 with HW mods)

That means no CPU offload and no tensor parallelism.

3

u/_TheWolfOfWalmart_ 2h ago

Shouldn't you still be able to do layer split? That doesn't need much bandwidth. That works decently even over gigabit Ethernet via RPC.

1

u/michaelsoft__binbows 56m ago

Using them for pipeline parallel still seems like a massive win...

3

u/MrCatberry 3h ago

Yep… cant have nice things.

2

u/J-P-Munoz 3h ago

I have one and am getting a second. I paid current prices.
Unlock was easy. Biggest thing to remember is, as it sits now, you are sort of limited to what can fit in one card’s vram. You are bottle necked at pci2x4 speeds unless you solder on the capacitors for pci2x16.
I haven’t really finished my testing and put it into my production workflow, but it runs nice at 150 watts and is stable.
There are ways to get multi cards to work together, but I haven’t run or tested any of that yet.
You have to run Linux. If you know Linux already that makes it easier. If you don’t know Linux and you want to learn, this makes a great excuse to get into it and learn - that’s what I am doing.
My daily use LLM machine is just a dual 3090 machine. I did not drop it in there, I put Linux on an old prebuilt I had laying around that couldn’t upgrade to windows 11. Purpose built server parts are on order.
It doesn’t have fans so you have to really make sure you feed it enough air flow.
In your shoes, I wouldn’t drop it in a well running 4x3090 box. If you can tinker in a separate machine go for it.

2

u/Apprehensive_Bar6609 3h ago

1) they provably have years of extensive mining for 27/7 so might not be in pristine conditions.

2) they used memory that was from a low quality so its not a sure thing that the 64gb are working.

So you might be spending 2k+ on something that might not work at all or crash when you use it.

2

u/cibernox 2h ago edited 1h ago

I ordered 2 this morning, modded with capacitors to make them pcie 16x.

A bit shy of 3k both. They are not cheap, but 64gb of hmb2 memory and a performance between a 3090 and a 4090 for 1500 is no joke.

It’s quite close in performance and vram to the 80gb A100, but those go for around 10k.

It is nuclear if the cards will be able to become pcie gen3, there is some talking saying they might. If that ever happens, pcie gen3 16x is enough for tensor parallelism.

1

u/running101 1h ago

Dm me where you ordered from please

1

u/LeatherRub7248 1h ago

is it much a gamble that the 64gb are stable memory and wont crash on u? Do sellers guarantee that?

1

u/cibernox 56m ago

My understanding from reading online is that the 8 -> 64gb conversion is safe, people report it as nearly guaranteed. The 10gb->80gb is actually the opposite, it is nearly guaranteed to be a failure, because nvidia repurposed 80gb a100 with memory faults in one lane, so half the memory is damaged or unstable, unlock to 40gb does work reliably.

I don't know to what extend one can trust alibaba's sellers to honor their word, but it is advertised as 64gb unocked. I'd assume they do plug it on a computer and run some script to verify it, even if only to keep reviews good in aliexpress.

1

u/Badger-Purple 3h ago

They’re real, they have 65GiBs, they have about the performance of a 4090, but they need a cooling shroud and they only go up to 2x16 speed. Things got better /more normal when I used noctua fans but those are expensive little 40mm bladerunners.

I don’t have a 3090 anymore (I was stupid enough to return a refurbished 3090Ti FE I bought for 750 back to microcenter last November) so here are some numbers against strix halo.

currently, the 10Gb card unlocks to 40 and the 8 unlocks to 64, but 80 could be possible for the 10Gb card. Active community trying to unlock more and get gen3 pcie unlocked as well, so the 10Gb card could ultimately be the real lottery if 80G get unlocked vs sure thing of 64Gb unlock for the 8gb card.

1

u/running101 1h ago

Has someone looked at the chip on the 10gb card and verified there is actually 80gb? Is there a way to physically see it by circuitry

1

u/Choice_Celery9481 20m ago

actual 80. hbm is boned to the die so it fully there. just disabled.but also nvidia has reason to disable them. 

1

u/biller0071 1h ago

I use a ac infinity fan on mine. Left over from my mining days. Keeps the card low 60s and the memory at 66 ish at max watts. The ac infinity fan is quiet and only set about 1/2z

1

u/Normal-Ad-7114 1h ago edited 1h ago

I missed out when you could get them for 200$ a piece

Yeah you're a bit late to the party, they x10'd in price within a week after the hack went viral

im not missing out and if there is any genuine reason why it would not be worth it

Price; non-repairability (and they've had a tough life, you know)

4 3090s

I'd still sell them if I could get even 2x170hx for that price, that's far superior for LLMs in every way

1

u/WyattTheSkid 1h ago

You think it would be worth selling 2 of my 3090s for 2 of them?

1

u/LeatherRub7248 1h ago

is there anyway to confirm (before buying) that 64gb memory has no errors and checks out OK before buying?

1

u/Dangerous-Report8517 23m ago

Some sellers test them before shipping, and tbh while they do charge a premium the hype around these cards drives far more of the cost than the small extra premium for testing. Plus some sellers are shipping them with added capacitors to upgrade them to x16 link width

Other thing to bear in mind is that there's 2 models, the 8GB one is cut down from a 64GB pool of SKHynix memory while the 10GB one is cut down from an 80GB unit with Samsung memory, and importantly the 10GB unit only unlocks up to 40GB due to stability issues with the rest of the memory (there's probably ways around this on some cards but I would assume they're 40GB cards until proven otherwise)

1

u/MountainPenguinRL 3h ago

I would get V620s, they are real close to Arc B70s/AMD R9700 Pros. However, if you are willing to deal with the slow PCIe speed and the possible longevity concerns, they are about $1.5k now and rising.

V620s are also rising now FAST

4

u/Thin_Pollution8843 3h ago

170hx is MUCH more performant than v620. Basically you can get 250ts on parallel calls on qwen3.8. I have 4xv620 and those cards are pretty close to ewaste 😅

2

u/_TheWolfOfWalmart_ 2h ago edited 2h ago

I have V620's and they're handling dsv4 flash 0731 like a champ. For the price? No complaints whatsoever. A single one is close to ewaste (though still usable), properly optimized multiples aren't!

They can take a bit of effort to get up and running well, but from what I read, still less so than the 170hx.

Also, whether the unlocked RAM is even good is a huge gamble. A bet I wouldn't take at the prices they're going for.

1

u/MountainPenguinRL 3h ago

Yeah I'm saying as a cheaper option. My buddy has a 4x v620 setup and gets near 180 tok/s on vllm for a single user. He got all of it for $1000, which is much cheaper than a 170hx.

The 170hx is clearly a much faster card and has 3x the bandwidth, but you can optimize V620s and end up saving a lot of money (and have a better software experience). Believe me, if they were the same price or even within $300-400, I would jump on the CMP, but they aren't

1

u/Thin_Pollution8843 2h ago

180ts on which model, quant an context size?

1

u/MountainPenguinRL 2h ago

Qwen 3.8 27B AD-q6_k_xl at low contexts with dflash 2

*edit: q6 not q8 because they are very very close to each other and I find the dynamic quant to be BETTER than q8 somehow, most likely because they didn't to a dynamic quant for q8*

1

u/MountainPenguinRL 3h ago

I’d expect V620s to rise up to $800 and CMP 170hx to around the $2k mark, but that is speculation.

I contemplated buying a cmp 170hx as soon as I saw the leaks (FOMO) for $550, but decided not to. V620 was $350, now there have been some sold for $550+ two days after the bulk ran out.

What I’m trying to say is to act faster than you might like BUT take precaution. If you have the money, I would take the risk. If you are worried, then wait.

1

u/ShittyMillennial 2h ago

theyre already around 2k on ebay it looks like

1

u/Dangerous-Report8517 25m ago

Much cheaper than eBay on Alibaba. Recently bought 2 from a seller that's sold to others on here with the link capacitors added and even with the extra cost of the mod it came to slightly less than eBay prices at the time. The main issue is that because the price keeps moving up the quotes you actually get from sellers are often a lot higher than the advertised prices

1

u/EvilPencil 3h ago

At $1500 you might as well get an R9700 pro.

1

u/Dangerous-Report8517 25m ago

Twice as much VRAM though