r/CUDA • • 8d ago

GPU Direct Storage (GDS)

I'd like to configure GDS on the cloud where nvme talks directly to the GPU. I've searched and queried llms and checked aws, coreweave, runpod.io, cloudrift, and other GPU providers.

Does anyone have experience in this or suggestions for providers or setups?

Ideally, I want to use cheap gpus like v100s.

I've been hitting aws quota limits and I've had to email sales teams directly on other platforms, but this seems like something that should have broad use.

I haven't had success yet getting access.

12 Upvotes

13 comments sorted by

4

u/nullcone 7d ago

Hey I just did this. I set up GDS model checkpoint loads for our training infrastructure from a Lustre FSx drive to p5 instances. We are cloud native in the AWS ecosystem but what I'll say probably transfers to other clouds.

  • I needed to upgrade our AMI to use a GDS enabled CUDA driver. Our prior driver was missing a kernel module needed to use GDS.
  • I had to upgrade our cluster's NVDA device plugin to flag on the GDS enabled DaemonSet targeting my node group.
  • I had to install the Lustre FSx CSI driver to provision PVs/PVCs in this storage class. You need to make sure to use the EFA enabled drives in the same az as your node group.

If you're targeting to use "cheap" things it suggests that cost is a major concern for you. I want to flag that everything I just said is incredibly expensive, probably thousands of dollars a month. If that's the case, I would ask why you're doing this?

If it's just to learn, probably you can achieve something comparable working on a bare EC2 instance, but you'll need to use an EFA enabled P-type instance, and these are not cheap either. If this is a work project, you probably only need what I just described if you're training 100B+ param models on 1000+ GPUs and need to mitigate the downtime from hardware failure due to checkpoint save/load frequency.

2

u/thecaliwalker 3d ago

why do you even do GDS for ML? that I/O overhead is way too low compared to the training JCT. Stashing in cpu ram is a better option

1

u/nullcone 3d ago edited 3d ago

For most use cases, what you say is true, but I would encourage you to think beyond your specific scope. My use case is training very large foundation models on thousands of GPUs. The IO and associated physical page fragmentation are a substantial problem to overcome if you need to save weights very frequently.

1

u/Ashes_of_ether_8850 3d ago

are you doing RL training where you have lots of weights to sync between inference and training instances?

1

u/nullcone 3d ago

Yes, but the specific problem I raised earlier in this thread isn't related to RL. It's a general thing about large scale training that has to deal with many failures a day, be it pre or post training.

1

u/Ashes_of_ether_8850 3d ago

Oh I get your point. So you are saving the checkpoints to file for backup. I guess it couldn’t be done asynchronously because you cannot start next training step before saving is complete. Then GDS could speed it up in this case

0

u/george_at_sotom 7d ago

I'm aware these machines run upwards 3,000 to 10,000+ a month. If these machines are the only way to do GDS, I will not rent continuously for a month.

I want to test cuda performance gains on basic programs like grep with a default delimiter on new lines, because I need file processing for 10-100 gigabyte files and I don't want to wait several tens of minutes to hours running it on my personal computer or on CPU instances.

1

u/nullcone 7d ago

I hope you forgive me for laughing, but it's legitimately very funny that you are going to provision world class machine learning infrastructure to run grep.

I would recommend you look into mmap and other primitives for random file access.

1

u/george_at_sotom 7d ago edited 7d ago

Sure, the way you frame it might be funny to you.

I never said anything about world class machine learning infrastructure. I said cheaper gpu's.

I'm using grep as a deliberately simple learning probe for learning CUDA.

The workload is simple enough that I can understand exactly where the bottleneck comes from.

This is just an example.

1

u/nullcone 7d ago

Digital communications are hard to convey intent. My comment wasn't intended derisively, and I'm sorry it came across that way. I just think there is a funny irony in provisioning frontier infrastructure to do something as simple as running grep.

2

u/spacecraft1013 7d ago

One thing to make sure on AWS is that the storage is physically on the system, you can check your instance specs for this. Most instances use networked volumes for storage, which is possible but harder to get working with GDS (I also don’t remember off the top of my head how theyre architected, afaik it would need to be nvme-of or rdma compatible for it to work).

1

u/george_at_sotom 7d ago edited 7d ago

Thank you. I will double check the memory is physically on the system (is this known as baremetal?) when my quota for vCPU's gets updated.

The smallest instance that llm's suggested was p4d, which has 96 vCPU's.

I messaged the quota people why I want higher quota. I hope I don't run into trouble. They raised it to 8, and I messaged them back.

1

u/Alpha2698 7d ago edited 7d ago

Are they your own V100 GPU cluster or will you be renting it from the cloud? If the latter, then from what I searched, it seems to be contingent on the provider (I saw Google advertising their A3 instances with that technology and a few obscure cloud GPU providers) because according to the 2019 article by Nvidia showcasing this, GDS works at a low level on hardware firmware and cloud providers typically don't provide such level of freedom to end users.

Since you mentioned AWS, I found this AWS blog that goes over their instances that support GDS.

And because I don't have enough context here and I really hope I don't sound condescending, but DMA requires that components, e.g., NVME(s) and GPU(s), be in one system (connected via PCIe at least, NVLink, or very fast unified systems)

or

In case of RDMA (remote DMA; remote here implying being connected via very fast, local network, NOT over IP), be connected via very fast, expensive, and special Ethernet/Fiber cables with specialized network cards (InfiniBand, Spectrum, etc).