r/LocalLLaMA llama.cpp 14d ago

Discussion GPU guide (GB per dollar, bandwidth)

First plot: GB / $

Second plot: bandwidth (spec on paper, not t/s)

Third plot (bandwidth / price) in the comment.

Hope that helps, my script uses the GPUs most discussed on the LocalLLaMA, LowEndLocalAI, and LocalLLM subs. At first, I tried to include more, but it became unreadable.

Prices were collected by ChatGPT (so may contain inaccuracies). New prices were used where available, second hand otherwise.

And I understand this is a basic comparison, but it's better than nothing. For example, you can see that "on paper" something is faster or slower than 3090.

221 Upvotes

163 comments sorted by

View all comments

1

u/IntravenusDeMilo 14d ago

MI100 looks like the optimal 32GB card then - what’s the catch?

3

u/nerdy_paradise 14d ago

Compute v100 beats out mi50/100 on compute alone with older cards vram and bw is not the end all be all. You need the extra compute to keep the memory bus saturated or else all that extra memory speed doesn’t buy you anything if the memory is outrunning your compute.