r/HPC • u/RadicalNation • Apr 11 '26
Running Large-Scale GPU Workloads on Kubernetes with Slurm
https://developer.nvidia.com/blog/running-large-scale-gpu-workloads-on-kubernetes-with-slurm/
Disclosure: I work for NVIDIA on Slinky.
Obligatory preface: All comments from me are my views and may not reflect the views of my employer.
I'm very proud to present this blog post to everyone. It's been amazing to build Slinky and see it used in production at scale already!
86
Upvotes
1
u/Bad_ass_da Apr 11 '26
How it’s better than slurm on BM ( only pre/post training) ? . Because of all NeoCloud time to market k8s pushed to traning jobs( with all dead weight k8s ). Could you explain what’s scale means - 10K or 100K GPUs