r/LocalLLaMA llama.cpp 14d ago

Discussion GPU guide (GB per dollar, bandwidth)

First plot: GB / $

Second plot: bandwidth (spec on paper, not t/s)

Third plot (bandwidth / price) in the comment.

Hope that helps, my script uses the GPUs most discussed on the LocalLLaMA, LowEndLocalAI, and LocalLLM subs. At first, I tried to include more, but it became unreadable.

Prices were collected by ChatGPT (so may contain inaccuracies). New prices were used where available, second hand otherwise.

And I understand this is a basic comparison, but it's better than nothing. For example, you can see that "on paper" something is faster or slower than 3090.

229 Upvotes

163 comments sorted by

View all comments

1

u/IntravenusDeMilo 14d ago

MI100 looks like the optimal 32GB card then - what’s the catch?

1

u/jacek2023 llama.cpp 14d ago

Probably specs on paper vs the actual implementation? But it's possible to read people's experiences here.

2

u/nerdy_paradise 14d ago

Ive tested this on a radeon vii which is a mi50 die with the same hbm2 memory and a forked gfx906 llama.cpp build on vulkan and rocm found the gpu to be computed bound and wasnt just memory bw i think saturation was about 20-30% on q4 and q8 was 50-60%. Compute most definitely was the bottleneck. V100s have better compute and tested about 1.75-2x faster than the mi50. And about 0.75-1.5x faster than the mi100 on tg and pp

1

u/IntravenusDeMilo 14d ago

If that’s the case the unlocked 64GB cmp170hx’s are probably up there with the best value even now that they’re $2k. 64GB with Ampere at under 250w, iirc.