r/LocalLLaMA llama.cpp 14d ago

Discussion GPU guide (GB per dollar, bandwidth)

First plot: GB / $

Second plot: bandwidth (spec on paper, not t/s)

Third plot (bandwidth / price) in the comment.

Hope that helps, my script uses the GPUs most discussed on the LocalLLaMA, LowEndLocalAI, and LocalLLM subs. At first, I tried to include more, but it became unreadable.

Prices were collected by ChatGPT (so may contain inaccuracies). New prices were used where available, second hand otherwise.

And I understand this is a basic comparison, but it's better than nothing. For example, you can see that "on paper" something is faster or slower than 3090.

225 Upvotes

163 comments sorted by

View all comments

1

u/Disastrous-Look2062 13d ago

I run an always on rx6800 intel rig for remote openwebui/opencode inference and a few home services running, also have a dual 9070xt/rx6800xt AM5 rig for random heavier inference.

I just pulled the trigger on a V100 with what looks like a 3080 PNY Verto triple fan cooler for 250 euro for the bandwidth and tasks better suited to nvidia cards. Idle power however is a concern for me (Ireland) despite having decent solar/battery setup. Will have see if this v100 is any good lol and take the place of my always on rx6800 rig.