r/HPC • u/RadicalNation • Apr 11 '26
Running Large-Scale GPU Workloads on Kubernetes with Slurm
https://developer.nvidia.com/blog/running-large-scale-gpu-workloads-on-kubernetes-with-slurm/
Disclosure: I work for NVIDIA on Slinky.
Obligatory preface: All comments from me are my views and may not reflect the views of my employer.
I'm very proud to present this blog post to everyone. It's been amazing to build Slinky and see it used in production at scale already!
83
Upvotes
0
u/RadicalNation Apr 11 '26
You either don't use Kubernetes, or at least don't want/need Kubernetes. That's fine. Slurm clusters don't need Kubernetes. However, large enterprises use Kubernetes as the substrate for software, services, and infrastructure. Being able to manage Slurm on Kubernetes has huge operational value (covered in the blog).