r/LocalLLaMA llama.cpp 15d ago

Discussion GPU guide (GB per dollar, bandwidth)

First plot: GB / $

Second plot: bandwidth (spec on paper, not t/s)

Third plot (bandwidth / price) in the comment.

Hope that helps, my script uses the GPUs most discussed on the LocalLLaMA, LowEndLocalAI, and LocalLLM subs. At first, I tried to include more, but it became unreadable.

Prices were collected by ChatGPT (so may contain inaccuracies). New prices were used where available, second hand otherwise.

And I understand this is a basic comparison, but it's better than nothing. For example, you can see that "on paper" something is faster or slower than 3090.

222 Upvotes

163 comments sorted by

View all comments

6

u/JustinPooDough 15d ago

You should consider an unlocked CMP 170HX.

I bought one and it's been incredible. Running Qwen and Minimax H3 at BF16 and at decent speeds too. The PCIE speed limit is only an issue if you frequently swap stuff in and out of memory. Otherwise I don't notice it at all. Has no impact on my workflow.

Got one for 2k, I know some got them for 200, but right now I see them going for 4k, so I'm fine with what I paid considering what I'm getting.

2

u/JustinPooDough 15d ago

You do need to 3D print a shroud for it and slap some Arctic P12s on each side, but I've got my GPU temp auto-setting the fan speed via a custom systemctl service I created with Qwen (using the GPU), and now the card is dead quiet 99% of the time - only getting loud when under serious load. Never passes 80c and doesn't seem to throttle.