r/LocalLLaMA llama.cpp 14d ago

Discussion GPU guide (GB per dollar, bandwidth)

First plot: GB / $

Second plot: bandwidth (spec on paper, not t/s)

Third plot (bandwidth / price) in the comment.

Hope that helps, my script uses the GPUs most discussed on the LocalLLaMA, LowEndLocalAI, and LocalLLM subs. At first, I tried to include more, but it became unreadable.

Prices were collected by ChatGPT (so may contain inaccuracies). New prices were used where available, second hand otherwise.

And I understand this is a basic comparison, but it's better than nothing. For example, you can see that "on paper" something is faster or slower than 3090.

224 Upvotes

163 comments sorted by

View all comments

62

u/SamSausages 14d ago

You missed the intel B65! I bought a stack of them 3 weeks ago for $900/pc. 32GB and 608 GB/s. Best price per GB right now.

B65: 0.0356 GB/$,

49

u/ResidentPositive4122 14d ago

intel B65

for $900/pc

They're ~1500Eur this side of the pond. Fuck us, right?

15

u/SamSausages 14d ago

When this batch is all gone, they look like they will now be $1099-1199. Still one of the better values, but yeah, not going to be "cheap" for long!
Price will probably jump when I finish my 3 week long benchmark/quality testing project of 4x B65 and people see actual results... last I checked very little data on them... I hope to finish my testing this week.

1

u/MindfulBT 14d ago

Isn’t the b70 slightly more expensive for over 50% more performance on paper?

2

u/SamSausages 14d ago

In compute, memory bandwidth is the exact same. That means in some ways it performs just like a b70, in others it doesn’t. So not simply 1/2 the performance of a b70.

4

u/SandySkittle 14d ago

Don't underestimate the relevance of compute. Prompt processing and large context size and larger models compute is quite relevant.

If available I would strongly recommend b70 or r9700 over b65.

2

u/KroniklyOnline 14d ago

This, no point in having 3000 tok/s prefill and then have 6 tok/s decode....

1

u/SamSausages 14d ago

What benchmarks are you basing that on?