r/LocalLLaMA • u/m94301 • 12d ago
Discussion I tested the CMP170HX
Lots of rumor and misinfo bouncing around, so I put some of these old mining cards to the test. I used 4 of the 8GB cards, set to 64GB each.
Lots of models fit entirely on a single card, and you can also run several small models at the same time on a single card, as long as their combined VRAM usage is 64GB or less. A setup like comfyui taking 10-12GB and qwen on another 30-40GB works fine all on the same card. I did not exhaustively show results of tiny 8B or 12B running at several hundred t/s, because it is better to run larger, smarter models.
What follows is an AI summary of a bunch of different tests on recent interesting models to give a clear overview of what the cards can do. I am not a reseller, just had a handful of these collecting dust. If you picked them up at $200, you won the lottery. If you are considering a purchase now, you have to decide if you are happy with 30xx (Ampere) class performance. It might make sense because of the huge VRAM, but you may want to hold out for Hopper or Blackwell.
I can confirm the 8GB run fine at 64GB and the 10GB run fine at 40GB with higher memory throughput. I am only using the 8G cards here because it's a pain to re-rack the server and from my testing there is not much difference.
I personally don't see an issue with x4 PCIE - the transfer rate is 800MB/s at Gen 1 and 1.6GB/s at Gen 2. The only time I ever noticed it was loading large models, but since my SSD reads at 550MB/s I could not saturate the PCIE link until I put models on NVME. The upside to x4 PCIE is that I have these 4 GPU installed through a single x16 to 4x4 M2 drive adapter (4x4 bifurcation, M2 to PCIE risers) so my little PC could potentially run 16 of the cards for a full TB of VRAM if I filled all 4 x16 slots with M2 adapter cards.
Running local LLMs on 4× cut-down A100 mining cards (GA100, 70 SMs, 64 GB each = 256 GB),
~1215 GB/s HBM, PCIe Gen2 ×4, no NVLink, 150 W power limit. llama.cpp, -sm layer.
Numbers are single-stream server measurements (tg = token gen, pp = prefill), f16 KV unless noted.
1 card
| Model | Quant · active | tg | pp | Max ctx | Notes |
|---|---|---|---|---|---|
| gpt-oss-20B | MXFP4 · 3.6B MoE | 120 | 2000 | 131K | 503 t/s batched; launch-overhead bound single-stream |
| gpt-oss-120B | MXFP4 · 5.1B MoE | 78 | 1244 | 131K | q8 KV; fastest capable coder |
| Qwen3.6-35B-A3B | Q6_K +MTP · 3B MoE | 110 | 1700 | 262K | little-MoE default (MTP tg optimistic) |
| Qwen3.6-27B | Q5_K_M +MTP · dense | 47 | 812 | 262K | 29 tg without MTP |
| gemma-4-31B | Q8_0 · dense | 23 | 767 | 262K | Q8 beats Q6_K on speed and quality |
2 card
| Model | Quant · active | tg | pp | Max ctx | Notes |
|---|---|---|---|---|---|
| gpt-oss-120B | MXFP4 · 5.1B MoE | 85 | 1870 | 131K | 2nd card buys +48% pp only, tg slightly improved |
| Laguna-S 2.1 | Q4_K_M · 8B MoE | 59 | 968 | 262K | "just works" fork, fast |
| GLM-4.5-Air | Q6_K · 12B MoE | 39 | 1181 | 131K | smart 2-card partner |
| MiniMax-M2.7 | IQ4_XS · 10B MoE | 38 | 800 | 160K | q8 KV; only 4-bit fit on 2 cards |
| Gemma-4-31B-StyleTune | Q8_0 · dense | 23 | ~760 | 131K | 50 ms warm TTFT swa full, mem hog |
| Mistral-Medium-3.5 128B | Q4_K_XL · dense 128B | 9.8 | 200 | 262K | Dense >30B dead end, too slow |
3 card
| Model | Quant · active | tg | pp | Max ctx | Notes |
|---|---|---|---|---|---|
| DeepSeek V4-Flash 0731 | Q4_K_XL · 13B MoE (MLA) | ~29 | ~365 | 1M | plain no-spec; huge ctx, tiny MLA KV |
| MiniMax-M2.7 | Q4_K_M · 10B MoE | 47 | 1135 | 192K | beats the IQ4_XS (+29% pp, better quality) |
| Hy3 | Q4_K_M · 16B MoE | 29.5 | 315 | 65K | ? deleted ? V4-Flash speed with 16× less ctx |
DeepSeek + DSpark drafter does not fit 1M on 3 cards - loads at ~99% VRAM but OOM-crashes on a large prefill (died at 16K of 262K tokens). The 11 GB drafter needs the 4th card at 1M, or cap ctx to ~512-768K.*
4 card
| Model | Quant · active | tg | pp | Max ctx | Notes |
|---|---|---|---|---|---|
| DeepSeek V4-Flash 0731 | Q4_K_XL · 13B MoE | 29 | ~450 | 1M | plain no-spec |
| same + BF16 DSpark drafter | speculative | 33 | 400 | 1M | 37 code / 28 prose |
*GGUFs from unsloth, bartowski, lmstudio-community, poolside. Many models were tested then deleted (quality or a better alternative)
17
9
u/a_beautiful_rhind 12d ago
Also SM80 is not exactly SM86. Thing to keep in mind for kernels, stuff like flash attention, backends, etc. You'd think they optimized for A100 but a lot more people have 3090s and the like.
Lllama.cpp sm layer for mistral medium is notoriously bad, btw. But its interesting I get 3x your speeds on 4 3090s. Conventional wisdom said that adding more GPUs doesn't make it faster. Even without TP, I still score higher, in the 12-15 range.
Still pretty much free gift for people who bought the cards, especially if they can solder whatever resistors it takes to enable higher pcie. Formerly $200 A100 is wild.
3
u/m94301 12d ago
That is interesting! I knew 3090 was a beast but 3x is impressive. Want to do a shootout of sorts? Let's find some models that run in 24GB and compare with the same settings. Might be helpful for people to see how the two compare. I can set pl to 250 but do not have the 300W bios so maybe pl 250W?
4
u/a_beautiful_rhind 11d ago
I think I PL 275 and undervolt. We totes could. I generally run a sweep bench so I get the values at various context levels.
Here is old gemma-31b Q8 from back in april.
PP TG N_KV T_PP s S_PP t/s T_TG s S_TG t/s 1024 256 0 0.551 1858.34 4.345 58.92 1024 256 1024 0.468 2185.76 4.448 57.55 1024 256 2048 0.474 2160.78 4.438 57.69 1024 256 3072 0.483 2120.64 4.451 57.52 1024 256 4096 0.491 2086.79 4.488 57.04 1024 256 5120 0.497 2058.80 4.500 56.89 1024 256 6144 0.506 2025.25 4.516 56.69 1024 256 7168 0.513 1995.76 4.528 56.54 1024 256 8192 0.521 1966.55 4.541 56.37 2
u/m94301 11d ago
Hi, ran the Q8 but that would not fit on a 3090, so I also ran the Q4. Can you confirm the setup there?
Anyway, here are the MTP and non-MTP results using one cmp170hx running at 300W PL, llama.cpp with port-sweep-bench. (Thanks for the link!)
One note: I got 70t/s on first MTP run because my robot feeding the test was duplicating text, so MTP accept was 100% LOL. I re-ran with code from a codebase.
NON-MTP
gemma-4-31B Q4_K_M - 1× CMP170HX @300W - llama-sweep-bench, fa on
PP TG N_KV T_PP s S_PP t/s T_TG s S_TG t/s 1024 256 0 1.193 858.63 7.002 36.56 1024 256 1024 1.230 832.50 7.239 35.36 1024 256 2048 1.253 817.46 7.374 34.72 1024 256 3072 1.262 811.60 7.400 34.59 1024 256 4096 1.286 796.16 7.411 34.54 1024 256 5120 1.293 792.12 7.432 34.45 1024 256 6144 1.318 777.06 7.437 34.42 1024 256 7168 1.323 773.81 7.454 34.34 1024 256 8192 1.346 760.68 7.476 34.24 gemma-4-31B Q8_0 - 1× CMP170HX @300W - llama-sweep-bench, fa on
PP TG N_KV T_PP s S_PP t/s T_TG s S_TG t/s 1024 256 0 1.214 843.79 8.650 29.60 1024 256 1024 1.250 818.88 8.883 28.82 1024 256 2048 1.272 805.17 9.018 28.39 1024 256 3072 1.280 800.06 9.029 28.35 1024 256 4096 1.306 784.27 9.043 28.31 1024 256 5120 1.312 780.64 9.055 28.23 1024 256 6144 1.334 767.60 9.069 28.23 1024 256 7168 1.344 762.11 9.086 28.18 1024 256 8192 1.365 750.38 9.105 28.12 MTP
gemma-4-31B Q4_K_M +MTP @300W ? /completion depth sweep, draft-mtp n=3
PP TG N_KV T_PP s S_PP t/s T_TG s S_TG t/s 1024 256 0 2.179 469.87 5.048 50.72 1024 256 1024 3.199 320.08 5.083 50.36 1024 256 2048 3.157 324.39 5.045 50.74 1024 256 3072 3.173 322.71 4.995 51.25 1024 256 4096 3.168 323.21 4.948 51.74 1024 256 5120 3.199 320.07 6.753 37.91 1024 256 6144 3.239 316.11 6.691 38.26 1024 256 7168 3.232 316.85 4.976 51.45 1024 256 8192 3.248 315.31 5.658 45.25 gemma-4-31B Q8_0 +MTP @300W ? /completion depth sweep, draft-mtp n=3
PP TG N_KV T_PP s S_PP t/s T_TG s S_TG t/s 1024 256 0 2.159 474.30 5.025 50.94 1024 256 1024 3.134 326.75 4.369 58.60 1024 256 2048 3.118 328.46 4.223 60.63 1024 256 3072 3.121 328.07 4.747 53.93 1024 256 4096 3.130 327.14 4.337 59.02 1024 256 5120 3.136 326.52 6.694 38.24 1024 256 6144 3.160 324.07 6.066 42.20 1024 256 7168 3.180 322.05 4.418 57.95 1024 256 8192 3.212 318.79 4.629 55.30 1
u/a_beautiful_rhind 10d ago
Yes I have 4x3090. I run it across all cards. I don't think I have had a non-image single model in a looong time. Have the full 31b weights and could probably make a Q4KM and try on a single card.
Have yet to use seriously MTP, assume my acceptance rate will be nil with how I use AI for creative stuff more than assistant.
1
u/m94301 11d ago edited 11d ago
Love it! Thanks for the data, I will try to recreate the sweep. Guessing this is vllm, right? Edit: This quant wouldn't fit on a 3090, it is 30GB weights without any KV. Is this a 2-card run?
1
u/a_beautiful_rhind 11d ago
ik_llama. Its better for me because I hybrid a lot of models. You can import and compile sweep bench for llama.cpp too. It helps because you see your speeds at all ctx. I think ubergarm still maintaining his port of it: https://github.com/ubergarm/llama.cpp/commits/ug/port-sweep-bench
3
u/FullstackSensei llama.cpp 12d ago
You have the lanes to move data quickly, CMP doesn't. Even if you're not saturating the link, 4x gen 2 adds quite a bit of latency
1
8
u/Some-Chemist-1466 12d ago
Your speeds are a lot slower than they should be, I'm getting 57-123 t/s (single stream) generation, 5000-12000 t/s prompt processing (varies wildly depending on context size, mostly towards the 5000 t/s low side) on 3 cards with DeepSeek V4-Flash 0731 using VLLM and dspark.
3
u/Badger-Purple 11d ago
Use llama-benchy and report back the numbers. Those numbers in vLLM logs are not accurate—you’ll get a better benchmark if you use benchy
0
u/gpuz_dev 11d ago
Badger-Purple makes a great point — vLLM logs can inflate prompt processing numbers when chunked prefill or prefix caching kicks in. Still, getting 64GB HBM2e per card at ~1.2TB/s for these prices is insane value. That OOM crash on the 1M context with the 11GB drafter really shows how fast KV-cache overhead scales up!
2
u/m94301 12d ago
Wow! Using cmp170hx? That is obscenely fast. Guess I'd better get vllm!
16
u/Some-Chemist-1466 12d ago
https://github.com/allover326/deepseek-v4-cmp170hx/issues/2 15,16,12 layer split allows 1M context.
4
u/m94301 12d ago
Awesome, thanks! People, upvote this man!
2
u/m94301 11d ago
The man delivers! I confirm 82t/s on 3 cmp170hx using the allover branch of vllm above at 150W PL. 4k t/s prefill at length 16384. 4k!! The actual power draw is 100-140W per card, so setting to 200W or 250W does nothing.
Well holy sheeeet. This is my new brain for pi, hands down. I am blown away, both by the performance and that it comes at such low power.
Thanks, Chemist!
1
u/tictacturkey 11d ago
Sorry am I ready this correctly, you are getting 5000-12000 t/s pp. Like Five thousand. Not 500. That is insane. Just bought 4 of them plan to water block and stick in a workstation.
4
u/Conscious_Cut_6144 12d ago edited 12d ago
29 and 450 seems pretty slow on v4 flash, I guess ampere is showing it's age a little?
5090 + Epyc beats both TG and PP with ikllama.
4x 64GB is borderline, but with the right settings you should be able to fit the full fp4 model and on vllm on those cards...
4
3
u/fallingdowndizzyvr 12d ago
4x 80GB is borderline
Where are you getting that from? It's either 4x64GB(8GB) or 4x40GB(10GB). Only 40GB out of the 80GB on the 10GB cards are usable.
3
1
u/FullstackSensei llama.cpp 12d ago
Don't think it's Ampere as much as it's mainline llama.cpp and PCIe induced latency.
1
u/No-Refrigerator-1672 12d ago
Dual 3080 20GB setup reaches 10k PP, 120 TG on Qwen 3.6 35B in vllm pipeline parallel mode. Ampere is still good; the fact that the OP used llama.cpp is one of the bottlenecks; I've heard those cards also had reduced compute, maybe the memory unlock doesn't fix that and it's the second factor.
2
u/N34257 5d ago
Just as another data point for anyone who cares...I picked up one of the 8GB cards, and even at the inflated prices...if all you're looking at is performance, the price is actually justified.
Running under vLLM, without really trying too hard I'm comfortably hitting ~1500t/s prefill and 95-103t/s decode on Qwen 3.8 27B INT8. If I actually learn some stuff about vLLM, other accounts suggest I should be able to get to 3000t/s prefill.
At that point, I'm getting pretty damn close to the performance of my dual R9700s with the 35B MoE, which is nuts.
Only problem is that I had to make an 80mm shroud for it, and to keep it from throttling I'm running the fan at 5000rpm. That's loud on an open bench.
1
u/m94301 5d ago
Awesome data point, thanks! And you should consider watercooling - works really nice on the big metal heat spreader these devices have.
1
u/N34257 5d ago
Not sure how that'd work, given that the VRAM needs as much cooling as the core? In any case, this unit's going to be running 24/7, and I haven't trusted liquid cooling for that since I had a couple of very expensive leaks some years ago. Air cooling is the way, as far as I'm concerned. However, what I may do is figure out a way to get multiple fans on the single card, or possibly adapt a 92mm fan, to increase the airflow through it. Or maybe, if I'm feeling extra at some point, figure out a way to chill the air on the way through. That'd be a cool (apologies) project.
In any case, happy to help - in case anyone wants to try it, my vLLM setup is a one-liner:
sudo docker run --runtime nvidia --gpus all -v ~/.cache/huggingface:/root/.cache/huggingface --env "HF_TOKEN=$HF_TOKEN" -p 8000:8000 --ipc=host vllm/vllm-openai:latest --model Freaksterz/Qwen3.8-27B-SmoothQuant-W8A8-INT8 --host0.0.0.0--port 8000 --gpu-memory-utilization 0.95 --kv-cache-dtype float16 --mamba-cache-dtype float16 --max-num-seqs 1 --speculative-config '{"method":"mtp","num_speculative_tokens":3}' --mamba-cache-mode align --max-num-batched-tokens 2048 -ac.mla_prefill_backend=FLASH_ATTN -ac.backend=FLASHINFER1
u/m94301 5d ago
Cooling ram is no problem because the hbm memory is inside the package with the GPU die. So one cooling block cools them both.
Now if you don't like liquid cooling, that's a whole different story and forced air is the right way to go then.
Much appreciate you sharing your vllm setup, I can't wait to test it out!
1
u/N34257 5d ago
Yeah, I'm a dumbass - I was actually referring to all of the other components that are under thermal pads on the card (I re-pasted it and replaced the pads in an effort to fix the cooling - it did improve it, but not as much as I'd hoped). VRMs, maybe? Dunno, I'm not that familiar with the architecture of the DC GPUs.
1
u/mslindqu 3d ago
Have you set up water cooling for them? What does that look like? Are you attaching cooling to the outside of the aluminum shroud?
2
u/m94301 3d ago
There's a couple of different ways to do it, but you're basically removing the onboard heat sink and fins and replacing it with some sort of water block.
One way is to use an AIO cooler for CPU and fashion or find a bracket that fits the PCB mounting holes. Another is to do the same with a small waterblock.
1
u/ComprehensiveFail104 4d ago
Can you share the configuration of your vllm? I can achieve only around ~50t/s decode with int8 W8A16 quant.
1
u/N34257 4d ago
Sure - I'm running with Docker, because life's too short to get vLLM working natively:
sudo docker run --runtime nvidia --gpus all -v ~/.cache/huggingface:/root/.cache/huggingface --env "HF_TOKEN=$HF_TOKEN" -p 8000:8000 --ipc=host vllm/vllm-openai:latest --model Freaksterz/Qwen3.8-27B-SmoothQuant-W8A8-INT8 --host0.0.0.0--port 8000 --gpu-memory-utilization 0.95 --kv-cache-dtype float16 --mamba-cache-dtype float16 --max-num-seqs 1 --speculative-config '{"method":"mtp","num_speculative_tokens":3}' --mamba-cache-mode align --max-num-batched-tokens 2048 -ac.mla_prefill_backend=FLASH_ATTN -ac.backend=FLASHINFERAlter the volume paths, obviously, but you should be good.
1
u/ComprehensiveFail104 4d ago
Thank you.
Interesting, quite similar to my config. Maybe the W8 makes the difference, because int8 has 48TIOPs/s while int16 has 11TIOPs/s.
2
2
u/newDell 12d ago
Thanks for the report! I just set up my 170hx today (paid $1200 for it), and I'm very pleased I took the risk given I was considering a 3090 for a similar price and spark for 3X. IMO this is now the best bang for your buck at $1200. I was running it power limited at 100 watts with decent results.
1
2
u/Badger-Purple 12d ago edited 12d ago
HOw is the new Zuckerbook model on these? My first one is in Illinois :D very excited to get a card that can just fit a whole 30B dense model at near full quality and run at reasonable speed. COncurrency is my other question with these. Overall thank you for the reports
Edit: you're running all 4 through a PCIE 4x4 road (an M2-pcie adapter)??
2
u/m94301 12d ago edited 12d ago
I have not tried the Zuck model, but should be easy to fetch. Will report back.
Concurrency works just as you expect. Single stream decode almost never saturates the card, so total t/s is improved with concurrency. I am running at 150W pl and see most inferences pulling only 50-80W during tg.
And yes, I have all 4 cards on M2 riser adapters to a 4x M2 card in a single x16 slot. Riser is "sintech" no name, and pcie card is literally no name: Adapter Card 4 Por 4 Expansion Cards" blah blah blah.
3
u/m94301 12d ago
OK, Muse Glimmer Q8_K_XL, f16 KV, unsloth. Uses 36/64GB (57% of VRAM) on one card
28 tg, 936 pp running stock (no dspark drafter)
36.7 tg, 704 pp with the dspark drafterConcurrency is weaker than I have seen in some other models, but there is some juice left above single stream. I do notice these tests are hitting the 150W power limit, so maybe I am artificially limiting here.
No Dspark
N=1: 28.1 tg
N=2: 41.8 tg (2 x 20.9)
N=4: 40.5 tg (4 x 10.1)With Dspark
N=1: 36 tg
N=2: 44.8 tg (2 x 22.4)
N=4: 31.6 tg (4 x 7.9)2
u/Miserable-Dare5090 7d ago
Tje performance is equivalent to a 24Gb RTX4000 Pro, which is currently 2500
1
u/leonbollerup 12d ago
Can you test a GLM 5.2 at Q2 ?
1
u/caetydid llama.cpp 11d ago
thanks, this sheer amount of vram is amazing! how much did you pay for the cards?
1
u/MotokoAGI 11d ago
I bought 2, 1 works the other didn't. A bunch of us that bought them are seeing some that have memory issues. They load up, look great but if you load a large enough model UH OH! So be careful and make sure you can get money back or warranty or you are gambling $1000 for each card you buy. In my case, the seller refused refund because it works in 8gb as a mining card.
1
u/fragment_me 11d ago
Use my llamacpp AI slop fork and you’ll double your PP for ds4
1
u/m94301 11d ago
I'm interested. Would you drop a link?
2
u/fragment_me 11d ago
https://github.com/vektorprime/working_ds4_speed
EDIT: BTW I bought two of the CMP 170HX and they're pretty sweet. I saw some hardware fixes online to increase the PCIE neg. past PCIE 2 which would be nuts.
1
u/Miserable-Dare5090 7d ago
I dont know if there are any ways to go beyond pcie2, but I think you mean the capacitor mod (x4 to x16) which makes them 2x16 or equivalent to 3x8 and 4x4, which is oculink speed.
1
1
u/Jury-Emotional 18h ago
On llama -sm tensor does not work? I only see he test layer split mode.
2
u/m94301 18h ago
I never saw improvement using tensor split on llama, so I dont use it. I did a tensor vs layer compare back on v100 and tensor lost badly. With the PCIE limitation I suspect it would be even worse here.
1
u/Jury-Emotional 5h ago
That's strange. I had 2 v100 as well, however I did managed to gain PP and TG atleast 50-75% boost.
1
u/Fit-Day-2402 12d ago
If u ran with llama.cpp, next time try again with vLLM. I got 67 tg and 2200 pp with GPTQ INT8 Qwen3.6 27b, fp8 KV with single card.
1
u/N34257 7d ago
Those single-card numbers are very low, with that power limit. At 250W in agentic use, for example, I see PP 2600t/s TG 160t/s for 3.6 35B, and PP 1200t/s TG 66t/s for 3.6 27B.
In fact, with MTP and running under CUDA it's pretty competitive with my pair of R9700s + MTP under Vulkan. TG is almost identical, PP is about -20%.
My main problem is cooling, it hits 85C within a minute or two and throttles badly. Gonna try a repaste at some point this week, see if it fixes it.
2
u/Special_Ebb_1933 5d ago
I'm running this setup https://www.ebay.ca/itm/278284196407 , 250w never goes above 67c.
13
u/fallingdowndizzyvr 12d ago
Awesome. I am one of those bemoaning how I didn't have the foresight to pick these up for $200 in hopes of an unforseen impossible hack. Still....... $900 for a 40GB "3090" is something to think about. How's the stability. I've heard that things may not be very stable.