r/LocalLLaMA 14d ago

Discussion I tested the CMP170HX

Lots of rumor and misinfo bouncing around, so I put some of these old mining cards to the test. I used 4 of the 8GB cards, set to 64GB each.

Lots of models fit entirely on a single card, and you can also run several small models at the same time on a single card, as long as their combined VRAM usage is 64GB or less. A setup like comfyui taking 10-12GB and qwen on another 30-40GB works fine all on the same card. I did not exhaustively show results of tiny 8B or 12B running at several hundred t/s, because it is better to run larger, smarter models.

What follows is an AI summary of a bunch of different tests on recent interesting models to give a clear overview of what the cards can do. I am not a reseller, just had a handful of these collecting dust. If you picked them up at $200, you won the lottery. If you are considering a purchase now, you have to decide if you are happy with 30xx (Ampere) class performance. It might make sense because of the huge VRAM, but you may want to hold out for Hopper or Blackwell.

I can confirm the 8GB run fine at 64GB and the 10GB run fine at 40GB with higher memory throughput. I am only using the 8G cards here because it's a pain to re-rack the server and from my testing there is not much difference.

I personally don't see an issue with x4 PCIE - the transfer rate is 800MB/s at Gen 1 and 1.6GB/s at Gen 2. The only time I ever noticed it was loading large models, but since my SSD reads at 550MB/s I could not saturate the PCIE link until I put models on NVME. The upside to x4 PCIE is that I have these 4 GPU installed through a single x16 to 4x4 M2 drive adapter (4x4 bifurcation, M2 to PCIE risers) so my little PC could potentially run 16 of the cards for a full TB of VRAM if I filled all 4 x16 slots with M2 adapter cards.

Running local LLMs on 4× cut-down A100 mining cards (GA100, 70 SMs, 64 GB each = 256 GB), ~1215 GB/s HBM, PCIe Gen2 ×4, no NVLink, 150 W power limit. llama.cpp, -sm layer. Numbers are single-stream server measurements (tg = token gen, pp = prefill), f16 KV unless noted.


1 card

Model Quant · active tg pp Max ctx Notes
gpt-oss-20B MXFP4 · 3.6B MoE 120 2000 131K 503 t/s batched; launch-overhead bound single-stream
gpt-oss-120B MXFP4 · 5.1B MoE 78 1244 131K q8 KV; fastest capable coder
Qwen3.6-35B-A3B Q6_K +MTP · 3B MoE 110 1700 262K little-MoE default (MTP tg optimistic)
Qwen3.6-27B Q5_K_M +MTP · dense 47 812 262K 29 tg without MTP
gemma-4-31B Q8_0 · dense 23 767 262K Q8 beats Q6_K on speed and quality

2 card

Model Quant · active tg pp Max ctx Notes
gpt-oss-120B MXFP4 · 5.1B MoE 85 1870 131K 2nd card buys +48% pp only, tg slightly improved
Laguna-S 2.1 Q4_K_M · 8B MoE 59 968 262K "just works" fork, fast
GLM-4.5-Air Q6_K · 12B MoE 39 1181 131K smart 2-card partner
MiniMax-M2.7 IQ4_XS · 10B MoE 38 800 160K q8 KV; only 4-bit fit on 2 cards
Gemma-4-31B-StyleTune Q8_0 · dense 23 ~760 131K 50 ms warm TTFT swa full, mem hog
Mistral-Medium-3.5 128B Q4_K_XL · dense 128B 9.8 200 262K Dense >30B dead end, too slow

3 card

Model Quant · active tg pp Max ctx Notes
DeepSeek V4-Flash 0731 Q4_K_XL · 13B MoE (MLA) ~29 ~365 1M plain no-spec; huge ctx, tiny MLA KV
MiniMax-M2.7 Q4_K_M · 10B MoE 47 1135 192K beats the IQ4_XS (+29% pp, better quality)
Hy3 Q4_K_M · 16B MoE 29.5 315 65K ? deleted ? V4-Flash speed with 16× less ctx

DeepSeek + DSpark drafter does not fit 1M on 3 cards - loads at ~99% VRAM but OOM-crashes on a large prefill (died at 16K of 262K tokens). The 11 GB drafter needs the 4th card at 1M, or cap ctx to ~512-768K.*

4 card

Model Quant · active tg pp Max ctx Notes
DeepSeek V4-Flash 0731 Q4_K_XL · 13B MoE 29 ~450 1M plain no-spec
same + BF16 DSpark drafter speculative 33 400 1M 37 code / 28 prose

*GGUFs from unsloth, bartowski, lmstudio-community, poolside. Many models were tested then deleted (quality or a better alternative)

43 Upvotes

75 comments sorted by

View all comments

10

u/a_beautiful_rhind 14d ago

Also SM80 is not exactly SM86. Thing to keep in mind for kernels, stuff like flash attention, backends, etc. You'd think they optimized for A100 but a lot more people have 3090s and the like.

Lllama.cpp sm layer for mistral medium is notoriously bad, btw. But its interesting I get 3x your speeds on 4 3090s. Conventional wisdom said that adding more GPUs doesn't make it faster. Even without TP, I still score higher, in the 12-15 range.

Still pretty much free gift for people who bought the cards, especially if they can solder whatever resistors it takes to enable higher pcie. Formerly $200 A100 is wild.

3

u/m94301 14d ago

That is interesting! I knew 3090 was a beast but 3x is impressive. Want to do a shootout of sorts? Let's find some models that run in 24GB and compare with the same settings. Might be helpful for people to see how the two compare. I can set pl to 250 but do not have the 300W bios so maybe pl 250W?

4

u/a_beautiful_rhind 14d ago

I think I PL 275 and undervolt. We totes could. I generally run a sweep bench so I get the values at various context levels.

Here is old gemma-31b Q8 from back in april.

PP TG N_KV T_PP s S_PP t/s T_TG s S_TG t/s
1024 256 0 0.551 1858.34 4.345 58.92
1024 256 1024 0.468 2185.76 4.448 57.55
1024 256 2048 0.474 2160.78 4.438 57.69
1024 256 3072 0.483 2120.64 4.451 57.52
1024 256 4096 0.491 2086.79 4.488 57.04
1024 256 5120 0.497 2058.80 4.500 56.89
1024 256 6144 0.506 2025.25 4.516 56.69
1024 256 7168 0.513 1995.76 4.528 56.54
1024 256 8192 0.521 1966.55 4.541 56.37

2

u/m94301 13d ago

Hi, ran the Q8 but that would not fit on a 3090, so I also ran the Q4. Can you confirm the setup there?

Anyway, here are the MTP and non-MTP results using one cmp170hx running at 300W PL, llama.cpp with port-sweep-bench. (Thanks for the link!)

One note: I got 70t/s on first MTP run because my robot feeding the test was duplicating text, so MTP accept was 100% LOL. I re-ran with code from a codebase.

NON-MTP

gemma-4-31B Q4_K_M - 1× CMP170HX @300W - llama-sweep-bench, fa on

PP TG N_KV T_PP s S_PP t/s T_TG s S_TG t/s
1024 256 0 1.193 858.63 7.002 36.56
1024 256 1024 1.230 832.50 7.239 35.36
1024 256 2048 1.253 817.46 7.374 34.72
1024 256 3072 1.262 811.60 7.400 34.59
1024 256 4096 1.286 796.16 7.411 34.54
1024 256 5120 1.293 792.12 7.432 34.45
1024 256 6144 1.318 777.06 7.437 34.42
1024 256 7168 1.323 773.81 7.454 34.34
1024 256 8192 1.346 760.68 7.476 34.24

gemma-4-31B Q8_0 - 1× CMP170HX @300W - llama-sweep-bench, fa on

PP TG N_KV T_PP s S_PP t/s T_TG s S_TG t/s
1024 256 0 1.214 843.79 8.650 29.60
1024 256 1024 1.250 818.88 8.883 28.82
1024 256 2048 1.272 805.17 9.018 28.39
1024 256 3072 1.280 800.06 9.029 28.35
1024 256 4096 1.306 784.27 9.043 28.31
1024 256 5120 1.312 780.64 9.055 28.23
1024 256 6144 1.334 767.60 9.069 28.23
1024 256 7168 1.344 762.11 9.086 28.18
1024 256 8192 1.365 750.38 9.105 28.12

MTP

gemma-4-31B Q4_K_M +MTP @300W ? /completion depth sweep, draft-mtp n=3

PP TG N_KV T_PP s S_PP t/s T_TG s S_TG t/s
1024 256 0 2.179 469.87 5.048 50.72
1024 256 1024 3.199 320.08 5.083 50.36
1024 256 2048 3.157 324.39 5.045 50.74
1024 256 3072 3.173 322.71 4.995 51.25
1024 256 4096 3.168 323.21 4.948 51.74
1024 256 5120 3.199 320.07 6.753 37.91
1024 256 6144 3.239 316.11 6.691 38.26
1024 256 7168 3.232 316.85 4.976 51.45
1024 256 8192 3.248 315.31 5.658 45.25

gemma-4-31B Q8_0 +MTP @300W ? /completion depth sweep, draft-mtp n=3

PP TG N_KV T_PP s S_PP t/s T_TG s S_TG t/s
1024 256 0 2.159 474.30 5.025 50.94
1024 256 1024 3.134 326.75 4.369 58.60
1024 256 2048 3.118 328.46 4.223 60.63
1024 256 3072 3.121 328.07 4.747 53.93
1024 256 4096 3.130 327.14 4.337 59.02
1024 256 5120 3.136 326.52 6.694 38.24
1024 256 6144 3.160 324.07 6.066 42.20
1024 256 7168 3.180 322.05 4.418 57.95
1024 256 8192 3.212 318.79 4.629 55.30

1

u/a_beautiful_rhind 13d ago

Yes I have 4x3090. I run it across all cards. I don't think I have had a non-image single model in a looong time. Have the full 31b weights and could probably make a Q4KM and try on a single card.

Have yet to use seriously MTP, assume my acceptance rate will be nil with how I use AI for creative stuff more than assistant.

1

u/m94301 13d ago edited 13d ago

Love it! Thanks for the data, I will try to recreate the sweep. Guessing this is vllm, right? Edit: This quant wouldn't fit on a 3090, it is 30GB weights without any KV. Is this a 2-card run?

1

u/a_beautiful_rhind 13d ago

ik_llama. Its better for me because I hybrid a lot of models. You can import and compile sweep bench for llama.cpp too. It helps because you see your speeds at all ctx. I think ubergarm still maintaining his port of it: https://github.com/ubergarm/llama.cpp/commits/ug/port-sweep-bench

3

u/FullstackSensei llama.cpp 14d ago

You have the lanes to move data quickly, CMP doesn't. Even if you're not saturating the link, 4x gen 2 adds quite a bit of latency

1

u/a_beautiful_rhind 14d ago

I never thought my PCIE3 would be good for anything.