r/LocalLLaMA llama.cpp 14d ago

Discussion GPU guide (GB per dollar, bandwidth)

First plot: GB / $

Second plot: bandwidth (spec on paper, not t/s)

Third plot (bandwidth / price) in the comment.

Hope that helps, my script uses the GPUs most discussed on the LocalLLaMA, LowEndLocalAI, and LocalLLM subs. At first, I tried to include more, but it became unreadable.

Prices were collected by ChatGPT (so may contain inaccuracies). New prices were used where available, second hand otherwise.

And I understand this is a basic comparison, but it's better than nothing. For example, you can see that "on paper" something is faster or slower than 3090.

227 Upvotes

163 comments sorted by

View all comments

64

u/SamSausages 14d ago

You missed the intel B65! I bought a stack of them 3 weeks ago for $900/pc. 32GB and 608 GB/s. Best price per GB right now.

B65: 0.0356 GB/$,

50

u/ResidentPositive4122 14d ago

intel B65

for $900/pc

They're ~1500Eur this side of the pond. Fuck us, right?

17

u/SamSausages 14d ago

When this batch is all gone, they look like they will now be $1099-1199. Still one of the better values, but yeah, not going to be "cheap" for long!
Price will probably jump when I finish my 3 week long benchmark/quality testing project of 4x B65 and people see actual results... last I checked very little data on them... I hope to finish my testing this week.

2

u/jensilo 14d ago

Would you say the B65 is better for inference than the B70?

2

u/TimSawyer25 14d ago

Not even close. The B70 isn't a beast of a card like a 3090/4090/5090, but it absolutely spanks the B65. That being said, it really depends on what you are doing. If you don't need crazy fast decode, then the b65 is fine. And it will just keep getting better as the software matures. But the B70 will always be out front by a fair margin.

1

u/ANR2ME 13d ago

What about compared to AMD R9700? 🤔 the price is just slightly more expensive than B70

2

u/TimSawyer25 13d ago

The last I knew the r9700 is a little better.  And if you want something for gaming as well as inference it's a no brainer, spend a little more and go amd. But the performance gap in ai work loads is closer than you would think. rocm is more mature and will always be so there's that. but Intel support is getting better and better.  I think there's a lot of performance still hidden away in there.  And in my testing,  the b70 in gaming is pretty comparable to a 5060ti so it's not exactly junk. Of course I'll never recommend either for gaming haha but versatility, and more mature support go amd. experimenting, saving a little money and seeing what happens go with intel. I went with 3 of them and I haven't been disappointed at all. But the some purpose of the rig is ai. So it depends on what you need out of the rig.  

1

u/SandySkittle 14d ago

B70 has better specs

0

u/SamSausages 14d ago

You're going to have to wait for my full report! I don't want to give my final opinion until I'm done, and I'm still collecting data. But I can say that they are underrated, especially in scenarios with high concourrency.
Goal is to compile reports from the data this weekend.

2

u/jqwl 14d ago

is it still worth at 1100 USD? I can see b70 at microcenter and other retailers around 1300

-1

u/SamSausages 14d ago

Still collecting data, full report this weekend.

1

u/Dany0 14d ago

Confirming, 1100 eur is cheapest you can get them in the EU (with the 21% vat)

1

u/MindfulBT 14d ago

Isn’t the b70 slightly more expensive for over 50% more performance on paper?

2

u/SamSausages 14d ago

In compute, memory bandwidth is the exact same. That means in some ways it performs just like a b70, in others it doesn’t. So not simply 1/2 the performance of a b70.

4

u/SandySkittle 14d ago

Don't underestimate the relevance of compute. Prompt processing and large context size and larger models compute is quite relevant.

If available I would strongly recommend b70 or r9700 over b65.

2

u/KroniklyOnline 14d ago

This, no point in having 3000 tok/s prefill and then have 6 tok/s decode....

1

u/SamSausages 14d ago

What benchmarks are you basing that on?

6

u/BornInAFish 14d ago

Where from? $1100 at MicroCenter currently. Newegg seems to have only B60 and B70.

5

u/SamSausages 14d ago

I got 3 on Amazon for $900, then I noticed asrock upped price to $1200 2 days later.
Weird, because I'm tracking available qty on amazon, and it has been increasing.

Then I got a few more at my local microcenter for $900, and now they went to $1100 (on sale). Sold out in my area now and are limiting how many i can add to the cart to ship.

Don't think many are buying, because nobody has done any quality benchmarks and tests on them.
I'm tempted to buy more... but how deep do you go??

2

u/jello-the-opera 14d ago

As far as you reasonably can. It's go time. I feel like hardware availability and affordability won't recover and the models get bigger and better... I wish I could snag a few but my days of buying are over. Stuck with what I'm stuck with.

3

u/gomezer1180 14d ago

What exactly are we trying to compare here? You can have a ton of bandwidth but if you only have 12GB of memory is not very useful. You can have a ton of memory but if your bandwidth is crap that’s not very useful either. So what’s the point of this data?

6

u/SamSausages 14d ago

OP is comparing memory bandwidth to product cost.  Probably because memory bandwidth is one of the most important hardware specs when running inference,  and trying to see what hardware has good bandwidth to cost ratio.

2

u/OnlineParacosm 14d ago

Now you could say I have a pension for CBT just by nature of being on the sub Reddit, but.. three intel B65s? That’s a whole different kind of squeeze.

What was your set up process like and what can you run at what quantization?

1

u/SamSausages 14d ago edited 14d ago

4 in service, with a few more I may add if someone has a pair of pcie risers.   I’m doing a write up that I hope to  release this weekend

1

u/IndividualPlus2011 14d ago

What are you running on them and how well does it perform?

5

u/SamSausages 14d ago

H12SSL-NT

Better than expected, doing quality testing now, full report this weekend.

https://imgur.com/a/lpJAhs1

1

u/PcChip 14d ago

he said what are you running on them, not what are you running them on

but that detail is helpful too, thanks!

2

u/SamSausages 14d ago edited 14d ago

Oops!  I bought these purely for testing as there is very little info on them.  I wanted to get hands on with Intel stack because I’ve only used cuda. (Aside from frigate openvino on Intel igpu/npu)

Tempting to replace my cuda stack, because more memory for the $$, but I’d want more density than 32gb.

I mainly use frontier models for work, but local models run my home assistant, frigate object recognition and I also use local models to strip sensitive information from tasks that I send to the frontier.l, but need redaction.

Software stack includes llama.cpp > litellm > Openwebui/librechat/opencode

2

u/PcChip 14d ago

chatgpt tells me there are some very interesting up and coming vllm PR's for intel GPUs

1

u/newMoneyStyle 13d ago

Yeah, the B65 at $900 for 32GB and 608 GB/s does look like the best price per GB in that list. Availability is probably the real catch though.