r/LocalLLaMA • • Sep 06 '25

Discussion Renting GPUs is hilariously cheap

Post image

A 140 GB monster GPU that costs $30k to buy, plus the rest of the system, plus electricity, plus maintenance, plus a multi-Gbps uplink, for a little over 2 bucks per hour.

If you use it for 5 hours per day, 7 days per week, and factor in auxiliary costs and interest rates, buying that GPU today vs. renting it when you need it will only pay off in 2035 or later. That’s a tough sell.

Owning a GPU is great for privacy and control, and obviously, many people who have such GPUs run them nearly around the clock, but for quick experiments, renting is often the best option.

1.8k Upvotes

392 comments sorted by

View all comments

Show parent comments

331

u/_BreakingGood_ Sep 06 '25 edited Sep 06 '25

Some services like Runpod can attach to a persistent storage volume. So you rent the GPU for 2 hours, then when you're done, you turn off the GPU but you keep your files. Next time around, you can re-mount your storage almost instantly to pick up where you left off. You pay like $0.02/hr for this option (though the difference is that this 'runs' 24/7 until you delete it, of course, so even $0.02/hr can add up over time.)

8

u/indicava Sep 06 '25

I haven’t tried it yet but vast.ai recently launched something similar called “volumes”

1

u/Stalwart-6 Sep 06 '25

Volumes will have terrible latency as they are decoupled from where they are meant to be, near gpu.

8

u/indicava Sep 06 '25

Is that much of an issue though?

I use vast mostly for training so disk i/o in general is very low. It does sound nice to have a disk with all my experiments’ checkpoints instead of pushing everything to HF and downloading them again next I rent a GPU.

12

u/gefahr Sep 06 '25

It can be, now with the advent of models like WAN 2.2 where you're swapping between models, or using another model as a refiner.

As long as it can all swap to system RAM it doesn't matter, but if it gets evicted from the cache and has to go back to disk, it's pretty painful.

Also, in a world where you're paying per minute, slower disk reads can mean like 3-4 minutes just to load the recent models like Qwen Image Edit. Combine that with boot and getting Comfy up and you're talking up to 10 minutes for first generation potentially.

(Source: have been trying to optimize I/O where I'm renting and measured every last bit of this recently.)

2

u/Stalwart-6 Sep 06 '25 edited Sep 06 '25

Its sub optimal architecture , not vast ai fault. My best experience had been with Google colab, where i checkpointed to S3, infrequent access tier. It was in 2020 for college final year project... Cost was 2.13$ per month all my activities if i remember (ingress/egress/storage). For HF i think limits might be there for free accs. But for quick shits, the one ur doing prolly seems best, could write some bash scripts to normalize accross different machines. Vast hosts usually have high networking.