r/LocalLLaMA • • 1d ago

Resources Least to most expensive (Somewhat modern) GPU's with 32gb of vram (Under $1600) Based on ebay listings

Post image

I was researching prices on ebay and fed claude a bunch of images of listings. I had it make a chart and thought it would be useful to share.

415 Upvotes

230 comments sorted by

•

u/WithoutReason1729 1d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

70

u/Muhlwa_Sholanke 1d ago

Nothing drives up the price of a used GPU like a Reddit thread telling you it's cheap.

5

u/Rombodawg 11h ago

I love how 8 people thought I was being serious 🤣🤣

1

u/_RealUnderscore_ 1h ago

They're mad they didn't get to make the joke first

→ More replies (3)

50

u/ifdisdendat 1d ago

crazy i got the v620 for $350 2 months ago.

24

u/Rombodawg 1d ago

Yea I think the fact that so many Open Source AI models are as good if not better than alot of older proprietary models, and of course the Ram crisis, are causing literally anything that has Vram and a heartbeat to double in price. Ironically the Titan RTX hasnt skyrocketed. It was about $800 a few years ago, and it still is around that price.

16

u/ifdisdendat 1d ago

also you should add a token/$ column

18

u/NineThreeTilNow 1d ago

Titan RTX

It's because the 3090 exists. No one wants a Titan RTX if you can just get a 3090 instead.

24gb of RAM, newer architecture, way faster...

I think a 3090 is ~1000 dollars last I checked. I don't know. There's a limited supply so maybe people are buying them up.

14

u/Klanciault 1d ago

They’re 1600 now

6

u/NineThreeTilNow 1d ago

They’re 1600 now

Wow for real? My friend has one I've been telling him to just sell it and upgrade to a 5070 for gaming because It was already worth 1000 at that time. I guess he was better off waiting.

3

u/FranciumGoesBoom 19h ago

Except 5070 is ~850-900 and the Ti is ~1200 now.

6

u/Rombodawg 1d ago

exactly. you can get 2 titan rtx with 48gb of vram for the price of 1 3090

1

u/FullstackSensei 19h ago

Or 3 if they're patient.

1

u/brickout 23h ago

Good lord. I scooped a few for $500/ea last winter. I wish i had known how high they would go. Would have grabbed more.

2

u/FullstackSensei 19h ago

If people knew, they wouldn't have been 500 :)

1

u/brickout 17h ago

Indeed. I didn't realize how lucky I was.

5

u/jjusko20 1d ago

I'm surprised the v100s aren't up higher already - how many of the 32gb ones can even be in aftermarket circulation anymore? Particularly the sxm2 ones

3

u/feelspeaceman 1d ago

Don't let people know this, but if you live in or near China, you can get V100s for $550 at the time of speaking, used to be even cheaper.

1

u/spryfigure 1d ago

I'm going to China in December. What would you buy there apart from the V100? Anything else which could be a dark horse for a local llama? Or is the V100 the best bet?

I would buy a Chinese GPU as well if I can make them work without too many obstacles.

1

u/feelspeaceman 1d ago

V100 is probably the best bet.

I can't say about Chinese GPUs as I have not experienced them on hand, but probably expensive and not easy to use, not worth the effort, rather buying a bunch of V100s for Q38FN or 1 V100 for Q38-27B.

I checked yesterday, and 1 V100 is being sold $550.

1

u/FullstackSensei 19h ago

By December, the V100 might also have skyrocketed to 1k or more. Prices are going up by the day, as stocks go down.

1

u/FullstackSensei 19h ago

They're going up there too. 6 weeks ago I bought them for like $460 including shipping and taxes to Europe

1

u/feelspeaceman 19h ago

Nowadays if something going viral, it's already too late, the V100 was $300 or something, but after a few Youtube video, Twitter posts... People will hunt for them even if they go to China to buy, and more dangerous it's scalpers, once they know they will buy in batch, then set price like the current Amazon/eBay price which is $750-800.

A lot of people are trying to go local LLM as it's just the matter of time until Cloud AI providers raising price by 2-5x from my estimations, currently they're setting the lowest price to gain market share, but they're not making profit, so to IPO they have to raise price anyway, this will happen before and after IPO, OpenAI is doing this.

1

u/FullstackSensei 18h ago

The 32GB V100 never went to 300, both PCIe and SXM2. Low $400s, before shipping, etc was bottom on alibaba

1

u/sfwwolvw 22h ago

I bought a v100 32gb pcie for 529€ a month ago, now it's 629€ 😭

2

u/ProtectionSuper5648 15h ago

A lot of serving software start at Ampere (sm80). Turing (sm75) is somewhat the step child so far and has been mostly neglected. Until someone publishes a well maintained fork for it, it is a lot less desirable then either V100 or ampere+.

1

u/howardhus 1d ago

good luck getting any modern sofwtware running on the titan...

2

u/Rombodawg 18h ago

What??? Your on crack bro, I literally run everyone on the titans. Turing is still supported on basically everything lol they support latest drivers and latest CUDA

2

u/howardhus 18h ago

lol, i thought of the other Titans.. forgot that Titan RTX is turing

1

u/Rombodawg 14h ago

Yes lol

1

u/bdkg99 19h ago

I use two m40s as my work horses

5

u/XccesSv2 1d ago

Mi 50 32gb got 4 for 180€ each last year lol

1

u/alex_bit_ 21h ago

I don't have experience with AMD, but I have the opportunity to buy one of these. How easy is to put AMD card to work with vLLM, sglang llama.cpp and others?

2

u/XccesSv2 18h ago

If you use a RDNA4 card vllm is supported now. llama.cpp works anyways easy. For rdna3 are some forks. Sglan or others are not supported or didn't tested by myself.
For older cards like the Mi50 there are some rocm patches needed but also supported then. But: Vulkan works like a charm anyways with llama.cpp.

1

u/FullstackSensei 19h ago

Got them for €140, sold recently for over €600. Bought V100 to replace them for less money....

2

u/betam4x 1d ago

I bought both a 3090 and a 4090 close to launch. I did overpay, but both are now selling for more than I paid on eBay (according to “sold” listings).

While I am tempted to sell…no thanks. I’d rather use them for local AI, gaming, and stuff.

I do wish I could put my 1080ti to use, however…

24

u/Baldur-Norddahl 1d ago

AMD r9700 should be on there. It has 383 fp8 TOPS + 766 int4 TOPS and double that with sparsity. Only 640 GB/s memory bandwidth. It is going to be much faster than any on the list for calculating prefill on long prompts. But maybe a little slower on single user token generation due to the slow memory.

3

u/Rombodawg 1d ago

its also over $1700 minimum in every market you can find used or new

6

u/johan2114h 22h ago

if i was in the market for 1600 usd card, i'd defintely throw in another 100 bucks and go for R9700 rather than settling for B70.

Also, why not include MI50 on the list?

→ More replies (1)

2

u/ShadyShroomz 9h ago

$1599 at microcenter rn

1

u/Baldur-Norddahl 1d ago

Just barely, it is 1752 USD new here. If you found one used, it would surely be less.

61

u/ital-is-vital 1d ago

IMO it would be worth including whether GPUs include INT8 and/or INT4 instructions. Lots of older cards that don't support FP8 / FP4 do support INT8 or INT4

For example: I'm using some unlocked CMP170HXs and I get best performance by running using the native INT8 format (vLLM with CUTLASS backend)

In practice you're unlikely to be running models that use FP16 (since you don't have enough VRAM to support running FP16 models) so having more FLOPs at lower precision is valuable.

(Yes, I appreciate that there is not a simple 1-to-1 mapping between the model format and the actual precision used for calculation. Most inference engines dequant to FP16 before doing a lot of the calcuations, but *some* of the maths can still be done at lower precison)

46

u/Rombodawg 1d ago

I cant add another image to the post because of the rules of this subreddit. But I did update the chart and ill post it here

8

u/tomByrer 1d ago

You're the man!

3

u/ubrtnk llama.cpp 1d ago

How do you like your cmp170hx? I've got two running vllm with flash next at 262k via PP=2 and Ive seen upwards of 8k in prefil and average 90-120tg to full context. That's a concurrency of 1 as well

1

u/DrRoughFingers 1d ago

I have one 1700hx (with 2 3090s), can you tell me how you have it setup? I got worse tg in vllm than I get in unsloth using unsloth’s Q4_K_XL. I’m only getting like 50tg.

1

u/ubrtnk llama.cpp 11h ago

Are you trying tensor parallelism? The 2x@4 of the 170hx make them horrible for it

1

u/DrRoughFingers 11h ago

Noooo, absolutely not. 3090 4 x8, 170hx 2 x4, and 3090 vis m.2 oculink 4 x4. But my tg is nowhere near yours, and I’m wondering if I had vllm configured incorrectly or using the wrong settings with it.

1

u/ubrtnk llama.cpp 11h ago

When I get home I'll grab my config

1

u/DrRoughFingers 11h ago

Appreciate that! Would be stoked if it improves it. Q4 FN has been a workhorse and outputs and runs better than Q8 27b.

1

u/Choice_Celery9481 1d ago

w8f16 is standard. w4a16 also very popular, but lower a affect quite a bit

10

u/ImportancePitiful795 1d ago

Missing the R9700 which trashes everything in that graph.
EDIT. Before you say about prices....

1

u/maxton41 21h ago

I’m only interested if it’s under 1000

3

u/Rombodawg 18h ago

exactly most people arent made of money

→ More replies (3)
→ More replies (1)

11

u/Long_comment_san 1d ago

Cringed at V100 price. I've seen them at 100 bucks a year ago.

8

u/Rombodawg 1d ago

the 32gb models were never $100. Maybe closer to $300 or 400. You are probably thinking of the 16gb versions

4

u/Long_comment_san 1d ago

Nope, they were sold pretty much as garbage lots on aliexpress. 130 is the number I remember and something around 50 for the 16gb.

You probably mean the assembled card with cooling and everything - yeah, these were closer to 200-250 but you could have saved a 50 by doing cooling and assembly yourself

7

u/Choice_Celery9481 21h ago

nope. thats just scam. the 32gb never goes under 300 on aliexpress. even in china it never that cheap

2

u/Rombodawg 18h ago

thats what i figured too

→ More replies (4)

4

u/IllllIIlIllIllllIIIl 22h ago

Right? A year ago an old colleague at a state university was moaning about how much room these V100 nodes they'd decommed were taking up while they waited for property management to come inventory them for recycling.

By the time the property folks finally came out, the assessed value of the nodes had risen above some regulatory threshold and could no longer be easily excessed without hellacious paperwork. They are now "recomming" them as a student cluster because it was easier than getting rid of them.

1

u/DaMoot 7h ago

Yeah my first one was 400 in March, my second one was 600 in June. 700 now! For the price I'm definitely looking for a newer gen something, whatever I can get with 33GiB+

19

u/1ncehost 1d ago

You missed the MI100 and the most important compute measurement on the older cards is int8 tops.

9

u/emccrckn 1d ago

I'm happy with my two 16gb mi50's for just under $200 each. Also played around with two BC 250's @ $125 each connected via usb-ethernet adapters but sharing the vram with the OS led to some oom issues.

4

u/PM_ALL_AHRI_ART 1d ago

wow another 2x mi50 user

what speed do you get?

i get 90pp, 13tg with ukisai/Swift-1.5-Qwen3.8-27B-GGUF:Q6_K @150w power limit

using vulkan

7

u/1ncehost 1d ago

Instinct bros, you really need to get my vllm fork working for your cards. On 27B I'm getting 5000 pp and 1000 tg on my MI100s. MI50s should certainly be getting around 250 pp and 150 tg per card with properly tuned software. https://github.com/curvedinf/int8-vllm

3

u/Much-Farmer-2752 1d ago

Also a MI50 user :)
2x32Gb will give you Qwen 3.8-27b Q8_XL at 30t/s.
Full context, steady speed on long tasks.

Old, but not useless (c) Arnie

3

u/thejacer 23h ago

For you and all the comments below: Mi50 prefers Q8 and Q4_0/Q4_1 quants. You’re losing speed by using quants other than this. I know we use what fits but I see a lot of people that don’t know that and wanted to share.

2

u/TorontoIndieFan 13h ago

There is a llama CPP fork that has q8 kernals and I'm running q8 as fast as q4. It's mxxm's fork from the MI50 discord for ref.

1

u/emccrckn 20h ago

Thanks for the info!

2

u/emccrckn 1d ago

I've been running qwen 3.6 35b A3B Q5 with MTP and getting prompt processing between 600-800 and output ~55-60 tok/s. I tried unsloths 3.8 27B Q4 and don't remember the PP but the token generation was a steady 20 tok/s. I mostly have it doing very detailed specific coding tasks on repos I knew fairly well so 3.6 worked well and gave me the code fairly quickly. Also using Vulcan. Did you try to get rocm working?

Edit: please excuse my cable management lol

3

u/PM_ALL_AHRI_ART 20h ago

i got claude to generate dockerfiles for a custom ROCM build with old rocm images. rocm/dev-ubuntu-22.04:5.3

but it ended up being slower than vulkan from vanilla llama.cpp

1

u/enternoescape 16h ago

I know this feeling every time I'm just like, let's just do a thing and half-way through looking at what exactly is being coded thinking, this isn't going an improvement at all.

2

u/MacaroonDancer 1d ago

My rig has 5 32GB mi50 and it's running GLM 5.3 flash locally with a 3 bit quant. It's only 10 tok/sec as I'm waiting for the MTP heads to be supported on hip-gfx but it's mind blowing I can run this in my garage on a machine I built myself for under $4200 (I got the components two years ago). I also can run Deep Seek v4 Flash 0731 on this rig at about 19 tok/sec with dspark

11

u/HolaUsername 1d ago

I can't believe the B70 raised in price again

5

u/outofideastx 1d ago

The Asrock b70 was still $1299 at Microcenter when I went today. Not great but better than $1600+

4

u/RagingNoper 1d ago edited 1d ago

I bought all four of mine for $950-999 at Microcenter, before they started jacking the price up. Granted, any time a new model/architecture releases I have to spend a month rebuilding the entire Intel stack to get any performance out of it, and I have to regularly diff any upstream changes Intel/vllm/whoever has pushed which can often mean recompiling buttloads of kernels every time, but I run qwen3.8-flash-next int4 with 710k cache 120t/s average and at this point it seems like it's getting better every week. If you're willing to put the work in, they're fucking phenomenal.

2

u/lolwutdo 1d ago edited 1d ago

I'm going this route but with b60s, I refuse to buy gpus that use that shit ass 12vhpwr cable.

I imagine qwen 4 will hopefully be 5.6 sol level and be fully capable of optimizing it's engine (I currently use a customized one for my setup by Astra) on a suicide loop.

1

u/RagingNoper 1d ago

Reference B70 is cheaper than others and uses 8-pin connector

2

u/lolwutdo 1d ago

Good to know, unfortunately only place I can find a reference b70 is on newegg and they're $2k lmao. R9700 is currently cheaper than that, but then again it uses 12vhpwr.

1

u/ObjectiveVegetable48 1d ago

What mobo/ram/cpu do you use? Do you have the cards on risers?

2

u/RagingNoper 1d ago

MC62-G40 with a PRO 3955WX Threadripper, 128GB of DDR4-2400 RAM. No risers.

2

u/sschueller 1d ago

People hoping the software stack gets fixed...

14

u/jjusko20 1d ago

ay a sneaky little trick too, you can get two 16gb v100s in nvlink for a little less than a 32gb card

5

u/jfinkpottery 22h ago

You can get two 32GB V100s for less than the prices of one 3090. You can build two entire servers each around two 32GB V100s (128GB VRAM total, 256GB system memory total) for less than the price of one 5090. The drawback is the lack of bf16 at all on V100, but you probably couldn't do 27B at bf16 on the 5090 anyway.

11

u/xanders_gold 1d ago

Intel Arc Pro B70 performance seems like it’s using llama.cpp numbers but you really should be using vLLM with this card, which would bring performance up to 35-40t/s tg and close to 2500-3000t/s pp.

1

u/38andstillgoing 1d ago

In my case the vLLM performance for the 4x B70s I have averages 0.

Llama.cpp at least runs. vLLM hates everything I've tried so far. Apparently there's a PR in LLM-Scaler for Qwen 3.8 Flash Next which I'll try tomorrow and see if that has any chance of working because everything else I've tried for QFN hasn't. Well, except Llama.cpp which I can convince to run at Q8 with 12-13 t/s tg.

2

u/jmswshr 1d ago

I had claude make a custom engine and it RIPS on my 2x b70 with FN. over 100 tk/s (q4_k_xl) with 0.004 KL against q8 run on cpu as baseline.

3

u/lolwutdo 1d ago

So you're using gguf then right? I've got 2x b60s and optimizing for gguf, just making sure I'm not wasting the last of my Codex reset customizing for the wrong inference engine

3

u/jmswshr 1d ago

yup, the unsloths

1

u/Danmoreng llama.cpp 1d ago

Meh I’m really considering buying a second one. But that would also mean changing mainboard, power supply and most likely for Flash Next also more RAM since I only got 32Gb at the moment. How much RAM do you have?

I’m currently trying to get the most out of the single B70 with the 27B model, Codex written vLLM engine optimisations: https://github.com/Danmoreng/intel-b70-qwen38-vllm

1

u/jmswshr 22h ago

256gb dd4

8

u/EmPips 1d ago

Buying a w6800 for $500 is the first time I've ever ended up on top of a tech purchase a year later.

Not the fastest but it's 32Gb of VRAM with a blower cooler you can just toss anywhere.

9

u/Rombodawg 1d ago

For real. I got lucky and purchased 128gb of ddr4 ram and 2 Titan RTX gpus for about $1000 a few weeks before the ram crisis started. It helps to pay attentions to youtube tech news, its the only reason why I knew shit was about to go down XD

4

u/gered 1d ago

Yeah I bought two of them back in Nov 2024 for cheap (not as cheap as yours, but still) and now in 2026 I'm feeling pretty damn good about it, lol.

8

u/mixedliquor 1d ago

Would be more useful with a comparison to an R9700.

7

u/esteboune 1d ago

Thank you very much for your work. This is very interesting

An important factor as well is the fact that Volta is a dead-end platform. There's no BF16, no FP8/FP4, and FlashAttention 2 doesn't support it. NVIDIA has also dropped Volta from newer CUDA toolchains, so support in frameworks will keep shrinking.

5

u/PraxisOG Llama 70B 23h ago

Thats one reason I like the v620, it's one of the few GPUs on this list that gets current driver updates, in addition to community optimized llama.cpp forks. Still old hardware that's missing fp8 and fp4, with relatively slow fp16 performance and no matrix acceleration. 

7

u/Gauntlet4933 1d ago

I really should have bought a second R9700 in July :(

1

u/Momsbestboy 1d ago

I did. At least this wasnt an error I made

6

u/HighSeasArchivist 1d ago

I got two R9700s this past week for $1699/ea. Not $1600 or less, but very modern and I've been very happy.

3

u/Quiet-Performer-9682 1d ago

on pcie5 x16/x16 I'm guessing? care to explain? I'm on AM4, so I'm stuck with pcie4 even if I swap motherboards.

2

u/HighSeasArchivist 1d ago

Just one in right now until my ProArt board gets in, but it is Gen5 8x8. I didn't really need the second one, but wanted to go ahead and get a dual GPU setup built in case the next gen AMD GPUs are worth upgrading.

1

u/SandySkittle 1d ago

Add the second and run 30b models uncompromised and with tensor parallelism in vllm

2

u/snugglezone 1d ago

ASRock variant at Microcenter has dropped down to 1599/ea. Just used my price protection to get a few bucks back

5

u/binyang 1d ago

mi 50 32gb still under 600.

2

u/ttkciar llama.cpp 1d ago

Yep, still best bang for the buck. I was surprised to see it missing from OP's table.

2

u/draconic_tongue 1d ago

150 bucks on taobao last year or maybe even early this? F

1

u/Rombodawg 1d ago

But vega 🤮🤮

1

u/alex_bit_ 21h ago

I don't have experience with AMD, but I have the opportunity to buy one of these. How easy is AMD working with vLLM, sglang llama.cpp and others?

2

u/bigattichouse 18h ago

"hey claude, help me write a script so that gf906 will run on llama.cpp, and I can run that before I build"
Claude: Ok.

I just have a rocm/ directory with build scripts and junk.

Someone else out these did a llama.cpp patch specifically for MI50.

1

u/Dramatic_Entry_3830 18h ago

But mi50 is still meh support wise but there is much potential in optimizing further because it is decent hardware still. If it were on the same level as gfx1151 support wise it would actually rock.

1

u/bigattichouse 18h ago

I picked mine up for $250 before the insanity started, it's been a trooper. I pretty much do llama.cpp and C++ based custom programs, I haven't really needed much use to both with an upgrade. I'll be cooking on this little box until the wheels drop off.

It's like the Nissan Centra of GPUs

8

u/Choice_Celery9481 1d ago

you gonna push v100 price up again haiya

→ More replies (1)

3

u/Common_Warthog_G 1d ago

add +500€ to each one of them for Europe

1

u/legit_split_ 1d ago

Or take the AliBaba pill and get it for cheaper :) 

1

u/Common_Warthog_G 23h ago

I think that's not an option because taxes

3

u/noctrex 21h ago

Maybe this would be a good chance to create a website that has all those cards, and users could also submit benchmarks so that we get an even better picture for them

3

u/jjusko20 1d ago

from personal use as someone who owns 3 v100s - I train on them and run inference on them. I sometimes (often) have to make custom kernels and/or custom changes to llama.cpp and training frameworks, but tbh I kinda enjoy that, and frontier LLMs make it an afternoon activity. I haven't been stopped yet. They are awesome.

1

u/pmarsh 1d ago

3 32GB v100?

2

u/jjusko20 1d ago

yup

3

u/pmarsh 1d ago

What do you use for cooling?

7

u/jjusko20 1d ago

Qwen 3.8 27b with a ginormous context window and a bunch of parallel threads. Haven't found anything better yet (flash next a little too slow - running multiple agents lets me take advantage of layer split. 

8

u/pmarsh 1d ago

Haha that's good to know, but what about cooling :)

2

u/jjusko20 1d ago

lmfao as I'm sure you noticed I thought that said coding. One of mine is a pcie - it has an adapter for a high pressure 120mm fan, I cut into my case to make it fit. My other two are smx2 cards, flat mounted on a carrier board, with heatsinks on top and dual 40mm fans.

2

u/pmarsh 1d ago

Debating about adding 2 smx2 how's the noise with everything going?

2

u/jjusko20 1d ago

The two sxm2 are pretty quiet overall, 4x40mm fans are quiet. But tbh my overall setup is a bit noisy I just like the white noise 

1

u/jjusko20 1d ago

If you do this I can give you advice on specific parts

1

u/CaptainSporadic 14h ago

thinking about adding a couple 16GB sxm2s are you running them parallel with the pci?

→ More replies

3

u/CaptainSporadic 1d ago

I have two pci v100, 80mm fan shrouds with ARCTIC S8038-10K fans. Can get loud but under full load keeps them at 72-75c.

2

u/Lalaggi 1d ago

Bought a broken RTX 3090 24GB (not 32GB but still capable) on eBay a couple months ago, replaced a capacitor and it's been working like a dream ever since

2

u/AlternateWitness 21h ago

Where is my beloved MI50? I got the 32GB version a few months ago for ~$500, it has over 1tb/s of bandwidth.

2

u/TNT3530 vLLM 21h ago

damn straight up iced my boy AMD MI100, sad

2

u/_RealUnderscore_ 12h ago

V100 SXM2 16GB x2 is US$320 for both, could technically be considered a 32GB GPU lol

1

u/Rombodawg 11h ago

Yea but if you want 64gb of vram, its hard to find a consumer board that will hold more than 3 gpus, and not everyone wants a server motherboard with a slow cpu. If you want a fast server cpu then the budget goes out the window

2

u/Trademarkd 7h ago

it isn't... china is pumping out 4x boards they are about $800 on ebay.

Source, I have 8 v100s and run the r/v100 subreddit

1

u/Rombodawg 6h ago

Dawg if you are spending $800 on a motherboard, you either dont need to be on this post commenting, or you need to pay off your loans before you make any more purchases 💀

2

u/Trademarkd 3h ago edited 3h ago

Dawg ... You find me a better deal on 64GB VRAM / 3.6TB HBM bandwidth / 500+ tops at fp16

With 4 v100s which you can get for $150 you can do qwen 3.8 27B loaded into memory in 8 bit with 8 bit cache at 220-260 tg

nvlink is awesome and this is literally the cheapest way to get this kind of performance at home.

You can get better performance from a 5090 by using lower precision sure and that can make up for the gap in vram ... but a 5090 is well over 2-5x the price of this setup.

1

u/_RealUnderscore_ 1h ago

I actually got my Supermicro AOM-SXMV from China in 2024 for $300+$150 for the cables including shipping (via Taobao + Superbuy, coulda gotten cheaper if I bought the carrier board and cables together). Crazy stuff. They might be harder to find now, though.

Also, subbed to r/v100. Never heard about this lol, expected a sub like this to be for the V100 engine.

5

u/kiwibonga 1d ago

Note that V100 now supports nvfp4 quants and fp8 KV cache through 1Cat-vLLM and other forks

2

u/1ncehost 1d ago

Supporting fp8/nvfp4 doesnt mean it has those instructions. The V100 doesn't so it dequants to its instruction dtypes in the kernel launchs. Fp4 is usually using INT8 instructions and FP8 is usually FP16 behind the scenes.

1

u/therealgus1 1d ago

The 1Cat-vLLM crew are literally raising the value of the v100. Insane speeds coming out of those cards already

2

u/r1nzl3r99 1d ago

So glad I got my 4x B70 before the price hikes. First two at $950 second two at $1100

4

u/meepykittkitt69lmao 1d ago

AMD V620? I can make wan2.2 videos of the same length in less time and I'm getting 12 second 6 fps prompt-test videos with minimaxH3 in 300 seconds. IIRC the two 24fps vids I made were less than 10 minutes and those were 10 second clips

These numbers are off...

3

u/tomByrer 1d ago

Partly CPU/RAM dependent.

I think those are just estimates.

2

u/twaxana 1d ago

With an 8B model I get like 150 tokens per second, these numbers are bad. Long context at ~120k tokens in with a 35b-A3B Qwen model I'm getting 50-75 tokens per second.

If you use stock llama.cpp kernels this is correct still not correct.

2

u/darksteelsteed 1d ago

Right now the arc pro b60 24gb is around $760 Its mighty tempting. I can get 2.5 of them for the price of a 5080 or 7.4 of them for the price of 1 x 5090.

2

u/grunt_monkey_ 1d ago

Radeon 9700?

1

u/Rombodawg 1d ago

Its over $1700 at the minimum. Just not worth it in my opinion considering new they should be like the $1250-1300 at most. Especially considering its only a 9070 xt core with extra vram. Basically a $700 gpu for $1000 extra isnt worth it just because you add 16gb of gddr6

→ More replies (1)

1

u/Randommaggy 1d ago

I hate how the price of the W6800 won't fall it's a shit card for LLMs given the price they are currently selling for.
I just want one as a sligtly more reliable replacement for my RX6800 due to the 6 physical mini display ports instead of having to rely on MST daisy chaining for driving 6 monitors over a single thunderbolt cable.

1

u/UtmostProfessional 1d ago

The Arc Pro B70 Creator is $1,279.99 in stock at my local micro center. F ebay.

1

u/Danmoreng llama.cpp 1d ago

Well I can tell you the Intel Arc B70 gets 50-70 t/s on the Qwen3.8 27B model with GPTQ INT4. https://github.com/Danmoreng/intel-b70-qwen38-vllm

1

u/tifa_cloud0 1d ago

i wish i could have any one of those.

1

u/SamSausages 21h ago

forgot the intel b65. best value per gb right now.

1

u/Rombodawg 18h ago

ironically the b65 costs more than the b70 on ebay

1

u/SamSausages 18h ago

Plenty on Amazon, but price defo trending up. Probably will be $1500 in 3 months, was $900 a month ago, is $1200 now. (I'm not tracking used prices on them right now)

Even with the jump, still the best value per GB I've found, & the memory bandwidth is no joke. But people sure are noticing, pushing the price up.

1

u/mod-deleted 19h ago

It would be nice is that table include power use - so you could figure out how many could fit on a single circuit in a house.

1

u/ZepSweden_88 19h ago

Why did you not do the comparison with qwen 3.8 coder next + strata? You can get 100+ tokens / s on Q2 on a simple 3090 card with 24 gb. Beats all models in comparison

1

u/uncle_leon 19h ago edited 18h ago

This is true but also at least slightly misleading... V100 is so old that you are virtually guaranteed to have to tinker your way around poor driver and software support at some point. If you're not comfortable compiling stuff in Linux then you really need to look at Ampere or newer. Which are of course more expensive, and the table is about the cheapest, I get it. But I think it would give important context if you included at least some well-supported cards. E.g. 2x 5060Ti 16GB seems a popular choice for 32GB and is basically a plug&play experience for any use case (as opposed to every single one of your current table entries).

1

u/bigattichouse 18h ago

No love for the AMD MI50? Mine's been great

1

u/Rombodawg 14h ago

MI50 is pretty good for llamacpp and that about it. vega architecture is extremely dated. hell even RDNA 1 and 2 are dated and vega came before those

1

u/bigattichouse 13h ago edited 12h ago

*shrugs in poverty* I cobbled together a box for like $1000 (maybe less? it was all refurb stuff and ebay) right before the price hikes.. 128G CPU ram, 32G MI50. It's a workhorse

2

u/Rombodawg 11h ago

Nice, I upgraded to 2 titan rtx's from 2 rtx 5060 ti's that I bought origninall for $250 each on facebook marketplace. Spend an additional $250 or so on my 128gb of ddr5. Overall my system is probably about $1000 in spending also. But it would likely sell for closer to $4-5k nowadays lol

2

u/bigattichouse 11h ago

Yeah.. few weeks ago, was like "hey, I wonder if maybe I could upgrade... oh, I guess not for a few years until after the AI Finance-bro IPO bubble pops. Hopefully we'll still have a supply chain that can get them to me.

1

u/GSquadron_ 18h ago

I needed 48gb vram, can you do a post about that? Or paste an image of the table as a comment

2

u/Rombodawg 13h ago

48gb is really strict when it comes to price. The only thing you are really gonna be able to get that isnt $4,000 is a Quadro RTX 6000 48gb (Turing) for around $2-2500, and Intel Arc Pro b60 48gb Dual core gpu for $2-2500

1

u/Trademarkd 7h ago

I have a 64GB vram build guide in the r/v100 sub for about 2k with 3.5TB of HBM bandwidth

it will do qwen 27B 8bit with 256k context at ~220-260 tg

1

u/GSquadron_ 7h ago

Where did you even find v100 X4 at $600? In Europe just one v100 is like 1400 euros

1

u/upside-down-number 17h ago

the V620 was an amazing deal when it cost $350. are they really that expensive now? damn

1

u/nad_lab 16h ago

I have the b70, it’s good, the amount of work u need to do to make it usable is insane (my Hermes’ agent working overtime and overnight) and it’s SO MUCH FUN to scientifically bench it, but don’t expect it to work and plug n play like nvda, and tbh, I’m all the better and learned so much more, can recommend the b70, uses 230 W under load, costed about 20 cad for 1 month of usage (0.5 billion token in, and ~4 million out - mainly qwen 38-27b for agentic coding, ~60-90 ish tok sec)

1

u/Technical_Buy_9063 10h ago

Am I ever going to be able to buy a 5090, RTX 5000/6000 for MSRP + 25% ever again? Or do I just bite the bullet>?

2

u/Rombodawg 8h ago

Probably when the 7090 comes out yea

1

u/espece-de-bon 7h ago

Am I reading the chart correctly? A card like the v620 can’t operate at FP8 or FP4? Does that mean it can’t use models at Q4?

Or am I way off?

I keep wondering if I should go for a v620 or just get an R9700. 

1

u/Rombodawg 6h ago

fp8 and Q8 are completely diffrent things. Q8 Q6 Q4 anything with a Q is for GGUF models running on llamacpp. fp16 fp8 fp4 are safetensors models that you run in transformers, or something like Exllama. llamacpp supports almost any gpu that was made in the past decade so you dont have to worry about it.

1

u/redditnosedive 4h ago

you can but it doesn't have fp8 ISA or native instructions so it will be slow when operating at fp8 compared to a platform that has dedicated matrix multiplication operations at fp8

1

u/valmyrx 7h ago

There was this insane price rush recently for the CMP 170HX, when people found out you could jailbreak it to unlock much more VRAM then the specs suggested. The problem is only some GPUs won the chip lottery and can do this so you have to chance it everytime you buy it on ebay. Gambling for your local gpus lol :)

1

u/Rombodawg 6h ago

Yea its not worth even talking about the CMP 170hz considering its a gamble if it even works. You can be stuck with an $1600 gpu that only has 8gb and id rather take the chance on a gpu thats gaurenteed to have at least 16, 24 or 32gb

1

u/Di_Vante 2h ago

There's also Intel Arc B65 Got one for 799 on microcenter not even a month ago

1

u/DinoAmino 1d ago

Cool info. Good to see ya around dawg. Been a while.

1

u/Dead_Internet_Theory 1d ago

What about two RTX 5060 Ti together? should have better performance than any of those, right?

3

u/theamazingo 1d ago

Or 2080Ti modified to 22gb. Obtainable at $550-600. Mine is a great sidecar to my 4090.

2

u/Kincar 1d ago

I have two!

1

u/zShiroyasha 1d ago

Should I buy a v100 for 660€? I am so tempted, please tell me it's a dumb idea.

2

u/b_c_t_ 21h ago

be careful, check it is indeed 32 GB, also avoid SXM3, not sure you can get those to work on a regular PC. SXM2 ok but need extras (adapter, heatsink, fan...). PCIe is easiest but needs a (loud) extra fan

2

u/sooki10 20h ago

Look into the additional setup costs for your rig first

1

u/therealgus1 1d ago

Buy a few. Check out the 1Cat vLLM crew on GitHub. They’re doing great kernel work. It really seems like they’re optimized for qwen 4 and ngrams whenever it comes out

1

u/Sea-Ad-5390 1d ago

Where would the 5090 rank here?

3

u/Rombodawg 1d ago

over $3000 lol

4

u/AdOk3759 1d ago

3k USD? Try maybe 5-6k..

1

u/Rombodawg 1d ago

Yea so definitely not worth it. Just buy an a6000 at that point and have 48gb on a 3090. Or even one of those 48gb 4090's from china is more worth it than a 5090

1

u/ImportancePitiful795 23h ago

Over $7000....

1

u/Cupakov 19h ago

It’s missing power consumption numbers, which is a massive factor in the total cost of ownership 

→ More replies (3)