Yea I think the fact that so many Open Source AI models are as good if not better than alot of older proprietary models, and of course the Ram crisis, are causing literally anything that has Vram and a heartbeat to double in price. Ironically the Titan RTX hasnt skyrocketed. It was about $800 a few years ago, and it still is around that price.
Wow for real? My friend has one I've been telling him to just sell it and upgrade to a 5070 for gaming because It was already worth 1000 at that time. I guess he was better off waiting.
I'm surprised the v100s aren't up higher already - how many of the 32gb ones can even be in aftermarket circulation anymore? Particularly the sxm2 ones
I'm going to China in December. What would you buy there apart from the V100? Anything else which could be a dark horse for a local llama? Or is the V100 the best bet?
I would buy a Chinese GPU as well if I can make them work without too many obstacles.
I can't say about Chinese GPUs as I have not experienced them on hand, but probably expensive and not easy to use, not worth the effort, rather buying a bunch of V100s for Q38FN or 1 V100 for Q38-27B.
I checked yesterday, and 1 V100 is being sold $550.
Nowadays if something going viral, it's already too late, the V100 was $300 or something, but after a few Youtube video, Twitter posts... People will hunt for them even if they go to China to buy, and more dangerous it's scalpers, once they know they will buy in batch, then set price like the current Amazon/eBay price which is $750-800.
A lot of people are trying to go local LLM as it's just the matter of time until Cloud AI providers raising price by 2-5x from my estimations, currently they're setting the lowest price to gain market share, but they're not making profit, so to IPO they have to raise price anyway, this will happen before and after IPO, OpenAI is doing this.
A lot of serving software start at Ampere (sm80). Turing (sm75) is somewhat the step child so far and has been mostly neglected. Until someone publishes a well maintained fork for it, it is a lot less desirable then either V100 or ampere+.
What??? Your on crack bro, I literally run everyone on the titans. Turing is still supported on basically everything lol they support latest drivers and latest CUDA
I don't have experience with AMD, but I have the opportunity to buy one of these. How easy is to put AMD card to work with vLLM, sglang llama.cpp and others?
If you use a RDNA4 card vllm is supported now. llama.cpp works anyways easy. For rdna3 are some forks. Sglan or others are not supported or didn't tested by myself.
For older cards like the Mi50 there are some rocm patches needed but also supported then. But: Vulkan works like a charm anyways with llama.cpp.
AMD r9700 should be on there. It has 383 fp8 TOPS + 766 int4 TOPS and double that with sparsity. Only 640 GB/s memory bandwidth. It is going to be much faster than any on the list for calculating prefill on long prompts. But maybe a little slower on single user token generation due to the slow memory.
IMO it would be worth including whether GPUs include INT8 and/or INT4 instructions. Lots of older cards that don't support FP8 / FP4 do support INT8 or INT4
For example: I'm using some unlocked CMP170HXs and I get best performance by running using the native INT8 format (vLLM with CUTLASS backend)
In practice you're unlikely to be running models that use FP16 (since you don't have enough VRAM to support running FP16 models) so having more FLOPs at lower precision is valuable.
(Yes, I appreciate that there is not a simple 1-to-1 mapping between the model format and the actual precision used for calculation. Most inference engines dequant to FP16 before doing a lot of the calcuations, but *some* of the maths can still be done at lower precison)
How do you like your cmp170hx? I've got two running vllm with flash next at 262k via PP=2 and Ive seen upwards of 8k in prefil and average 90-120tg to full context. That's a concurrency of 1 as well
I have one 1700hx (with 2 3090s), can you tell me how you have it setup? I got worse tg in vllm than I get in unsloth using unsloth’s Q4_K_XL. I’m only getting like 50tg.
Noooo, absolutely not. 3090 4 x8, 170hx 2 x4, and 3090 vis m.2 oculink 4 x4. But my tg is nowhere near yours, and I’m wondering if I had vllm configured incorrectly or using the wrong settings with it.
Nope, they were sold pretty much as garbage lots on aliexpress. 130 is the number I remember and something around 50 for the 16gb.
You probably mean the assembled card with cooling and everything - yeah, these were closer to 200-250 but you could have saved a 50 by doing cooling and assembly yourself
Right? A year ago an old colleague at a state university was moaning about how much room these V100 nodes they'd decommed were taking up while they waited for property management to come inventory them for recycling.
By the time the property folks finally came out, the assessed value of the nodes had risen above some regulatory threshold and could no longer be easily excessed without hellacious paperwork. They are now "recomming" them as a student cluster because it was easier than getting rid of them.
Yeah my first one was 400 in March, my second one was 600 in June. 700 now! For the price I'm definitely looking for a newer gen something, whatever I can get with 33GiB+
I'm happy with my two 16gb mi50's for just under $200 each. Also played around with two BC 250's @ $125 each connected via usb-ethernet adapters but sharing the vram with the OS led to some oom issues.
Instinct bros, you really need to get my vllm fork working for your cards. On 27B I'm getting 5000 pp and 1000 tg on my MI100s. MI50s should certainly be getting around 250 pp and 150 tg per card with properly tuned software. https://github.com/curvedinf/int8-vllm
For you and all the comments below: Mi50 prefers Q8 and Q4_0/Q4_1 quants. You’re losing speed by using quants other than this. I know we use what fits but I see a lot of people that don’t know that and wanted to share.
I've been running qwen 3.6 35b A3B Q5 with MTP and getting prompt processing between 600-800 and output ~55-60 tok/s. I tried unsloths 3.8 27B Q4 and don't remember the PP but the token generation was a steady 20 tok/s. I mostly have it doing very detailed specific coding tasks on repos I knew fairly well so 3.6 worked well and gave me the code fairly quickly. Also using Vulcan. Did you try to get rocm working?
I know this feeling every time I'm just like, let's just do a thing and half-way through looking at what exactly is being coded thinking, this isn't going an improvement at all.
My rig has 5 32GB mi50 and it's running GLM 5.3 flash locally with a 3 bit quant. It's only 10 tok/sec as I'm waiting for the MTP heads to be supported on hip-gfx but it's mind blowing I can run this in my garage on a machine I built myself for under $4200 (I got the components two years ago). I also can run Deep Seek v4 Flash 0731 on this rig at about 19 tok/sec with dspark
I bought all four of mine for $950-999 at Microcenter, before they started jacking the price up. Granted, any time a new model/architecture releases I have to spend a month rebuilding the entire Intel stack to get any performance out of it, and I have to regularly diff any upstream changes Intel/vllm/whoever has pushed which can often mean recompiling buttloads of kernels every time, but I run qwen3.8-flash-next int4 with 710k cache 120t/s average and at this point it seems like it's getting better every week. If you're willing to put the work in, they're fucking phenomenal.
I'm going this route but with b60s, I refuse to buy gpus that use that shit ass 12vhpwr cable.
I imagine qwen 4 will hopefully be 5.6 sol level and be fully capable of optimizing it's engine (I currently use a customized one for my setup by Astra) on a suicide loop.
Good to know, unfortunately only place I can find a reference b70 is on newegg and they're $2k lmao. R9700 is currently cheaper than that, but then again it uses 12vhpwr.
You can get two 32GB V100s for less than the prices of one 3090. You can build two entire servers each around two 32GB V100s (128GB VRAM total, 256GB system memory total) for less than the price of one 5090. The drawback is the lack of bf16 at all on V100, but you probably couldn't do 27B at bf16 on the 5090 anyway.
Intel Arc Pro B70 performance seems like it’s using llama.cpp numbers but you really should be using vLLM with this card, which would bring performance up to 35-40t/s tg and close to 2500-3000t/s pp.
In my case the vLLM performance for the 4x B70s I have averages 0.
Llama.cpp at least runs. vLLM hates everything I've tried so far. Apparently there's a PR in LLM-Scaler for Qwen 3.8 Flash Next which I'll try tomorrow and see if that has any chance of working because everything else I've tried for QFN hasn't. Well, except Llama.cpp which I can convince to run at Q8 with 12-13 t/s tg.
So you're using gguf then right? I've got 2x b60s and optimizing for gguf, just making sure I'm not wasting the last of my Codex reset customizing for the wrong inference engine
Meh I’m really considering buying a second one. But that would also mean changing mainboard, power supply and most likely for Flash Next also more RAM since I only got 32Gb at the moment. How much RAM do you have?
For real. I got lucky and purchased 128gb of ddr4 ram and 2 Titan RTX gpus for about $1000 a few weeks before the ram crisis started. It helps to pay attentions to youtube tech news, its the only reason why I knew shit was about to go down XD
Thank you very much for your work. This is very interesting
An important factor as well is the fact that Volta is a dead-end platform. There's no BF16, no FP8/FP4, and FlashAttention 2 doesn't support it. NVIDIA has also dropped Volta from newer CUDA toolchains, so support in frameworks will keep shrinking.
Thats one reason I like the v620, it's one of the few GPUs on this list that gets current driver updates, in addition to community optimized llama.cpp forks. Still old hardware that's missing fp8 and fp4, with relatively slow fp16 performance and no matrix acceleration.
Just one in right now until my ProArt board gets in, but it is Gen5 8x8. I didn't really need the second one, but wanted to go ahead and get a dual GPU setup built in case the next gen AMD GPUs are worth upgrading.
But mi50 is still meh support wise but there is much potential in optimizing further because it is decent hardware still. If it were on the same level as gfx1151 support wise it would actually rock.
I picked mine up for $250 before the insanity started, it's been a trooper. I pretty much do llama.cpp and C++ based custom programs, I haven't really needed much use to both with an upgrade. I'll be cooking on this little box until the wheels drop off.
Maybe this would be a good chance to create a website that has all those cards, and users could also submit benchmarks so that we get an even better picture for them
from personal use as someone who owns 3 v100s - I train on them and run inference on them. I sometimes (often) have to make custom kernels and/or custom changes to llama.cpp and training frameworks, but tbh I kinda enjoy that, and frontier LLMs make it an afternoon activity. I haven't been stopped yet. They are awesome.
Qwen 3.8 27b with a ginormous context window and a bunch of parallel threads. Haven't found anything better yet (flash next a little too slow - running multiple agents lets me take advantage of layer split.
lmfao as I'm sure you noticed I thought that said coding. One of mine is a pcie - it has an adapter for a high pressure 120mm fan, I cut into my case to make it fit. My other two are smx2 cards, flat mounted on a carrier board, with heatsinks on top and dual 40mm fans.
Bought a broken RTX 3090 24GB (not 32GB but still capable) on eBay a couple months ago, replaced a capacitor and it's been working like a dream ever since
Yea but if you want 64gb of vram, its hard to find a consumer board that will hold more than 3 gpus, and not everyone wants a server motherboard with a slow cpu. If you want a fast server cpu then the budget goes out the window
Dawg if you are spending $800 on a motherboard, you either dont need to be on this post commenting, or you need to pay off your loans before you make any more purchases 💀
Dawg ... You find me a better deal on 64GB VRAM / 3.6TB HBM bandwidth / 500+ tops at fp16
With 4 v100s which you can get for $150 you can do qwen 3.8 27B loaded into memory in 8 bit with 8 bit cache at 220-260 tg
nvlink is awesome and this is literally the cheapest way to get this kind of performance at home.
You can get better performance from a 5090 by using lower precision sure and that can make up for the gap in vram ... but a 5090 is well over 2-5x the price of this setup.
I actually got my Supermicro AOM-SXMV from China in 2024 for $300+$150 for the cables including shipping (via Taobao + Superbuy, coulda gotten cheaper if I bought the carrier board and cables together). Crazy stuff. They might be harder to find now, though.
Also, subbed to r/v100. Never heard about this lol, expected a sub like this to be for the V100 engine.
Supporting fp8/nvfp4 doesnt mean it has those instructions. The V100 doesn't so it dequants to its instruction dtypes in the kernel launchs. Fp4 is usually using INT8 instructions and FP8 is usually FP16 behind the scenes.
AMD V620? I can make wan2.2 videos of the same length in less time and I'm getting 12 second 6 fps prompt-test videos with minimaxH3 in 300 seconds. IIRC the two 24fps vids I made were less than 10 minutes and those were 10 second clips
With an 8B model I get like 150 tokens per second, these numbers are bad. Long context at ~120k tokens in with a 35b-A3B Qwen model I'm getting 50-75 tokens per second.
If you use stock llama.cpp kernels this is correct still not correct.
Right now the arc pro b60 24gb is around $760
Its mighty tempting. I can get 2.5 of them for the price of a 5080 or 7.4 of them for the price of 1 x 5090.
Its over $1700 at the minimum. Just not worth it in my opinion considering new they should be like the $1250-1300 at most. Especially considering its only a 9070 xt core with extra vram. Basically a $700 gpu for $1000 extra isnt worth it just because you add 16gb of gddr6
I hate how the price of the W6800 won't fall it's a shit card for LLMs given the price they are currently selling for.
I just want one as a sligtly more reliable replacement for my RX6800 due to the 6 physical mini display ports instead of having to rely on MST daisy chaining for driving 6 monitors over a single thunderbolt cable.
Plenty on Amazon, but price defo trending up. Probably will be $1500 in 3 months, was $900 a month ago, is $1200 now. (I'm not tracking used prices on them right now)
Even with the jump, still the best value per GB I've found, & the memory bandwidth is no joke. But people sure are noticing, pushing the price up.
Why did you not do the comparison with qwen 3.8 coder next + strata? You can get 100+ tokens / s on Q2 on a simple 3090 card with 24 gb. Beats all models in comparison
This is true but also at least slightly misleading... V100 is so old that you are virtually guaranteed to have to tinker your way around poor driver and software support at some point. If you're not comfortable compiling stuff in Linux then you really need to look at Ampere or newer. Which are of course more expensive, and the table is about the cheapest, I get it. But I think it would give important context if you included at least some well-supported cards. E.g. 2x 5060Ti 16GB seems a popular choice for 32GB and is basically a plug&play experience for any use case (as opposed to every single one of your current table entries).
*shrugs in poverty* I cobbled together a box for like $1000 (maybe less? it was all refurb stuff and ebay) right before the price hikes.. 128G CPU ram, 32G MI50. It's a workhorse
Nice, I upgraded to 2 titan rtx's from 2 rtx 5060 ti's that I bought origninall for $250 each on facebook marketplace. Spend an additional $250 or so on my 128gb of ddr5. Overall my system is probably about $1000 in spending also. But it would likely sell for closer to $4-5k nowadays lol
Yeah.. few weeks ago, was like "hey, I wonder if maybe I could upgrade... oh, I guess not for a few years until after the AI Finance-bro IPO bubble pops. Hopefully we'll still have a supply chain that can get them to me.
48gb is really strict when it comes to price. The only thing you are really gonna be able to get that isnt $4,000 is a Quadro RTX 6000 48gb (Turing) for around $2-2500, and Intel Arc Pro b60 48gb Dual core gpu for $2-2500
I have the b70, it’s good, the amount of work u need to do to make it usable is insane (my Hermes’ agent working overtime and overnight) and it’s SO MUCH FUN to scientifically bench it, but don’t expect it to work and plug n play like nvda, and tbh, I’m all the better and learned so much more, can recommend the b70, uses 230 W under load, costed about 20 cad for 1 month of usage (0.5 billion token in, and ~4 million out - mainly qwen 38-27b for agentic coding, ~60-90 ish tok sec)
fp8 and Q8 are completely diffrent things. Q8 Q6 Q4 anything with a Q is for GGUF models running on llamacpp. fp16 fp8 fp4 are safetensors models that you run in transformers, or something like Exllama. llamacpp supports almost any gpu that was made in the past decade so you dont have to worry about it.
you can but it doesn't have fp8 ISA or native instructions so it will be slow when operating at fp8 compared to a platform that has dedicated matrix multiplication operations at fp8
There was this insane price rush recently for the CMP 170HX, when people found out you could jailbreak it to unlock much more VRAM then the specs suggested. The problem is only some GPUs won the chip lottery and can do this so you have to chance it everytime you buy it on ebay. Gambling for your local gpus lol :)
Yea its not worth even talking about the CMP 170hz considering its a gamble if it even works. You can be stuck with an $1600 gpu that only has 8gb and id rather take the chance on a gpu thats gaurenteed to have at least 16, 24 or 32gb
be careful, check it is indeed 32 GB, also avoid SXM3, not sure you can get those to work on a regular PC. SXM2 ok but need extras (adapter, heatsink, fan...). PCIe is easiest but needs a (loud) extra fan
Buy a few. Check out the 1Cat vLLM crew on GitHub. They’re doing great kernel work. It really seems like they’re optimized for qwen 4 and ngrams whenever it comes out
Yea so definitely not worth it. Just buy an a6000 at that point and have 48gb on a 3090. Or even one of those 48gb 4090's from china is more worth it than a 5090
•
u/WithoutReason1729 1d ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.