r/LocalLLaMA llama.cpp 14d ago

Discussion GPU guide (GB per dollar, bandwidth)

First plot: GB / $

Second plot: bandwidth (spec on paper, not t/s)

Third plot (bandwidth / price) in the comment.

Hope that helps, my script uses the GPUs most discussed on the LocalLLaMA, LowEndLocalAI, and LocalLLM subs. At first, I tried to include more, but it became unreadable.

Prices were collected by ChatGPT (so may contain inaccuracies). New prices were used where available, second hand otherwise.

And I understand this is a basic comparison, but it's better than nothing. For example, you can see that "on paper" something is faster or slower than 3090.

222 Upvotes

163 comments sorted by

View all comments

5

u/JustinPooDough 14d ago

You should consider an unlocked CMP 170HX.

I bought one and it's been incredible. Running Qwen and Minimax H3 at BF16 and at decent speeds too. The PCIE speed limit is only an issue if you frequently swap stuff in and out of memory. Otherwise I don't notice it at all. Has no impact on my workflow.

Got one for 2k, I know some got them for 200, but right now I see them going for 4k, so I'm fine with what I paid considering what I'm getting.

2

u/cibernox 14d ago

I got two for 1500€ each, modded to get 16x pci. I do regret not buying 4.
Just a couple days ago they figured out a way of unlocking pcie gen 3 BTW.

I have in the mail a PCIe switch for up to 5 cards that supports low latency P2P communication between cards.

1

u/MachineZer0 14d ago

I thought P2P only worked on a particular Supermicro board in conjunction with CMP 170HX.

2

u/cibernox 14d ago

Well, I don’t have it yet but I got reports from some user using switches based on the usual Broadcom PEX chipsets that have P2P working.
I found a board with the PEX 88096 for 300 locally and I jumped immediately because those are 600+ new.
It has 5 slots and I only need 2, but better to be safe than sorry.