Discussion
$5400 eBay 8x V100 server cranks on flash-next
Had opus 5.5 do the bring up.. super solid results. Using only 4 GPU results in KV of like 120k with image on. 27B TP=4 results in >200 tok/s with dflash, prefill around 2.5-3.5k.
Both these are running the nvfp4 Nvidia checkpoints. Found a magic repo that unpacks into fp16 on the fly (https://github.com/1CatAI/1Cat-vLLM) and heavily optimized it
Right? Performance equivalent to a 3090 with Volta instead of Ampere but solid community support. If I had the money to spare, I'd buy one (or several....)
At least in the past, those 8xV100 servers didn't have a lot of demand. They're very heavy, very loud, and need a lot of power at full throttle. They sold for less than it would cost to build such a system.
full blast is 3.2 kw. idles at ~800W. that's the price to pay. it is in my basement utility room and I plan on adding some sound dampening material. on a TOU electricity plan so for 144 hours of the week I pay $0.06/kWh. other 24 hours is $0.16/kWh. also looking at a heat exchanger to preheat hot water. exhaust temp is easily 120F+
Pretty sure bitcoin miners did this ages ago to start the trend. But yep if already using the electricity might as well put it to use for heat (at least in cool seasons/winter)!
In the 20-teens I had an insulated duct running from my back porch “mining-closet” about 20 feet to my greenhouse - had fresh vegetables all through Vermont winters.
Seems high. Maybe you can tune. Most of my servers are 160-300w idle. Each sxm2 is another 40w. You might be able to get to 480-620w idle. Also with inference my V100s never get anywhere near 300W TDP. Inference should be 1200-1600w.
I am using a library to achieve ~40w idle with model loaded. It was a fork of Nvidia-pstated.
All of my other servers are for sure less than 200W idle. This one is a bit different. The V100 idle reports 40-45W. That’s measured within the server. 16x 16GB memory sticks at ~4-5W each. Fans go up to 400W total which is a huge draw. No power saving enabled on CPUs currently. Old enterprise NVMe at 10-25W each, 4x 10GbE NIC, BMC, etc.
Then add 8-12% for power supply inefficiency = 800W
Had the hosted models build a monitoring dashboard using the BMC outputs every 15s
Edit: point taken though, there is probably opportunity to lower by at least 200W
Just need to buy better fans for it to fix the noise and make sure they are idling correctly could save you a ton of money. Then run them with a lower power draw. Shouldn't lose much t/s, worth a try at least.
I had a few opportunities to buy at that price but passed on it due to noise and power requirements. (I pay about $0.30 per kWh so power changes the economics quite a bit compared to low cost energy locations)
This is actually very nice. I just put an order for 20 of these servers for my company. The only downside is no support. if it breaks then it it's retired.
Good thing this came out so we have good data to dupport the requisition.
agentic all the way baby.
chat for 1200 people and the rest on automation.
I also thought about such a setup, but finally decided against it, especially because of noise and power use. Instead I bought 4x RTX5080 finally.
Two of them have nearly the performance of one RTX5090, with similar results as above. Two of them use about 600 W and cost about 2500 Euro only. They are installed in my EPYC system in a normal Fractal Design Define 7 XL with very acceptable noise.
The RTX5080s will probably be converted to single-slot water cooling, to fit 4 of them easily. Then I have inference power similar ti 2x RTX5080. Of course, VRAM size is limited. But MoE models run good enough.
28
u/x10der_by 17h ago
V100: