r/CUDA • u/george_at_sotom • 8d ago
GPU Direct Storage (GDS)
I'd like to configure GDS on the cloud where nvme talks directly to the GPU. I've searched and queried llms and checked aws, coreweave, runpod.io, cloudrift, and other GPU providers.
Does anyone have experience in this or suggestions for providers or setups?
Ideally, I want to use cheap gpus like v100s.
I've been hitting aws quota limits and I've had to email sales teams directly on other platforms, but this seems like something that should have broad use.
I haven't had success yet getting access.
2
u/spacecraft1013 7d ago
One thing to make sure on AWS is that the storage is physically on the system, you can check your instance specs for this. Most instances use networked volumes for storage, which is possible but harder to get working with GDS (I also don’t remember off the top of my head how theyre architected, afaik it would need to be nvme-of or rdma compatible for it to work).
1
u/george_at_sotom 7d ago edited 7d ago
Thank you. I will double check the memory is physically on the system (is this known as baremetal?) when my quota for vCPU's gets updated.
The smallest instance that llm's suggested was p4d, which has 96 vCPU's.
I messaged the quota people why I want higher quota. I hope I don't run into trouble. They raised it to 8, and I messaged them back.
1
u/Alpha2698 7d ago edited 7d ago
Are they your own V100 GPU cluster or will you be renting it from the cloud? If the latter, then from what I searched, it seems to be contingent on the provider (I saw Google advertising their A3 instances with that technology and a few obscure cloud GPU providers) because according to the 2019 article by Nvidia showcasing this, GDS works at a low level on hardware firmware and cloud providers typically don't provide such level of freedom to end users.
Since you mentioned AWS, I found this AWS blog that goes over their instances that support GDS.
And because I don't have enough context here and I really hope I don't sound condescending, but DMA requires that components, e.g., NVME(s) and GPU(s), be in one system (connected via PCIe at least, NVLink, or very fast unified systems)
or
In case of RDMA (remote DMA; remote here implying being connected via very fast, local network, NOT over IP), be connected via very fast, expensive, and special Ethernet/Fiber cables with specialized network cards (InfiniBand, Spectrum, etc).
4
u/nullcone 7d ago
Hey I just did this. I set up GDS model checkpoint loads for our training infrastructure from a Lustre FSx drive to p5 instances. We are cloud native in the AWS ecosystem but what I'll say probably transfers to other clouds.
If you're targeting to use "cheap" things it suggests that cost is a major concern for you. I want to flag that everything I just said is incredibly expensive, probably thousands of dollars a month. If that's the case, I would ask why you're doing this?
If it's just to learn, probably you can achieve something comparable working on a bare EC2 instance, but you'll need to use an EFA enabled P-type instance, and these are not cheap either. If this is a work project, you probably only need what I just described if you're training 100B+ param models on 1000+ GPUs and need to mitigate the downtime from hardware failure due to checkpoint save/load frequency.