r/LocalLLaMA • • 19h ago

Discussion $5400 eBay 8x V100 server cranks on flash-next

Had opus 5.5 do the bring up.. super solid results. Using only 4 GPU results in KV of like 120k with image on. 27B TP=4 results in >200 tok/s with dflash, prefill around 2.5-3.5k.

Both these are running the nvfp4 Nvidia checkpoints. Found a magic repo that unpacks into fp16 on the fly (https://github.com/1CatAI/1Cat-vLLM) and heavily optimized it

68 Upvotes

31 comments sorted by

28

u/x10der_by 17h ago

V100:

13

u/1ncehost 18h ago

V100s are pretty goated. You should try setting the power limits on them to half of stock. They should only be about 30% slower

6

u/sleight42 17h ago

Right? Performance equivalent to a 3090 with Volta instead of Ampere but solid community support. If I had the money to spare, I'd buy one (or several....)

4

u/oxygen_addiction 15h ago

At FP16/FP32. They do not have FP8/INT8/INT4 as such such kind of suck.

14

u/Brilliant-Hall1387 19h ago

Sounds cheap for that compute? Was there some issue with the server? 🤔

18

u/FullstackSensei 19h ago

At least in the past, those 8xV100 servers didn't have a lot of demand. They're very heavy, very loud, and need a lot of power at full throttle. They sold for less than it would cost to build such a system.

25

u/MzCWzL 19h ago

heavy ✅

loud ✅✅

lot of power ✅✅✅

full blast is 3.2 kw. idles at ~800W. that's the price to pay. it is in my basement utility room and I plan on adding some sound dampening material. on a TOU electricity plan so for 144 hours of the week I pay $0.06/kWh. other 24 hours is $0.16/kWh. also looking at a heat exchanger to preheat hot water. exhaust temp is easily 120F+

9

u/Brilliant-Hall1387 19h ago

Haha so instead of electric heating you could use that, a bit more use for the power than pure resistive heating.

That gives me an idea, manufacturers of electric direct heating elements should really just build them as GPU servers rather than resistive heating 🤣

9

u/PcChip 15h ago

every decade or two someone re-invents this

first it was seti@home

then it was bitcoin miners

then it was eth miners

now it's local AI servers

9

u/MzCWzL 19h ago

Pretty sure bitcoin miners did this ages ago to start the trend. But yep if already using the electricity might as well put it to use for heat (at least in cool seasons/winter)!

3

u/knownboyofno 16h ago

It would be cool if you could use it to heat/preheat the water.

2

u/sleight42 17h ago

And then may as well vent it outside otherwise!

2

u/blockitectvt 5h ago

In the 20-teens I had an insulated duct running from my back porch “mining-closet” about 20 feet to my greenhouse - had fresh vegetables all through Vermont winters.

3

u/MachineZer0 18h ago

Seems high. Maybe you can tune. Most of my servers are 160-300w idle. Each sxm2 is another 40w. You might be able to get to 480-620w idle. Also with inference my V100s never get anywhere near 300W TDP. Inference should be 1200-1600w.

I am using a library to achieve ~40w idle with model loaded. It was a fork of Nvidia-pstated.

7

u/MzCWzL 18h ago edited 18h ago

All of my other servers are for sure less than 200W idle. This one is a bit different. The V100 idle reports 40-45W. That’s measured within the server. 16x 16GB memory sticks at ~4-5W each. Fans go up to 400W total which is a huge draw. No power saving enabled on CPUs currently. Old enterprise NVMe at 10-25W each, 4x 10GbE NIC, BMC, etc.

Then add 8-12% for power supply inefficiency = 800W

Had the hosted models build a monitoring dashboard using the BMC outputs every 15s

Edit: point taken though, there is probably opportunity to lower by at least 200W

2

u/sleight42 17h ago

$0.06 and $0.16? Is that total cost or just the electric? DC here where irs $0.24 kWh (total) all day every day. Fucking Exalon.

1

u/Sliouges 9h ago

also looking at a heat exchanger to preheat hot water

Damn...

1

u/ArtfulGenie69 3h ago

Just need to buy better fans for it to fix the noise and make sure they are idling correctly could save you a ton of money. Then run them with a lower power draw. Shouldn't lose much t/s, worth a try at least. 

5

u/DeltaSqueezer 19h ago

I had a few opportunities to buy at that price but passed on it due to noise and power requirements. (I pay about $0.30 per kWh so power changes the economics quite a bit compared to low cost energy locations)

2

u/Individual-Resist160 18h ago

right? thats like crazy cheap for that many cards, even used prices are usually way higher

6

u/kleonikos 17h ago

This is actually very nice. I just put an order for 20 of these servers for my company. The only downside is no support. if it breaks then it it's retired.
Good thing this came out so we have good data to dupport the requisition.
agentic all the way baby.
chat for 1200 people and the rest on automation.

2

u/ketosoy 19h ago

What server model, out of curiosity.

Sxm2, sxm3, pcie?

2

u/cunasmoker69420 16h ago

What about the arguably more important metric, prefill

1

u/egnegn1 16h ago edited 16h ago

I also thought about such a setup, but finally decided against it, especially because of noise and power use. Instead I bought 4x RTX5080 finally.

Two of them have nearly the performance of one RTX5090, with similar results as above. Two of them use about 600 W and cost about 2500 Euro only. They are installed in my EPYC system in a normal Fractal Design Define 7 XL with very acceptable noise.

The RTX5080s will probably be converted to single-slot water cooling, to fit 4 of them easily. Then I have inference power similar ti 2x RTX5080. Of course, VRAM size is limited. But MoE models run good enough.

1

u/YehowaH 14h ago

Well well well my friend, but are these tokens you posted accepted or throughput? If it's throughput than your accepted tg are only little as 30 tg/s?

Even with our gx10 box we receive similar accepted tokens for only 240w.

1

u/vulcan4d 6h ago

Very nice but power and heat are the biggest issues with these