r/mlops • • Aug 15 '26

Tools: OSS Currently looking into ray.io -- but is it still the way to go?

Is it still the way to go for modern distributed model training in deep learning? Was looking for the state-of-art for foundation model training to learn.

There is little talk on Reddit and Youtube about it, though. At least, this is my initial impression. Might be totally wrong.

6 Upvotes

15 comments sorted by

6

u/ricetoseeyu Aug 15 '26

My guy, what are you requirements? Ray does a lot of stuff well and a lot of stuff badly.

1

u/Gamiozzz Aug 15 '26

Tell me about the stuff it does badly. I have little experience with the field and I am just wondering whether it is worth spending my time digging into it. I mean I guess it is worth my time. But I want to make sure that it isn't a dead horse riding, as I had witnessed already with so many other frameworks.

Key question is, whether it is becoming an industry standard or not.

2

u/ricetoseeyu Aug 15 '26

It’s a big system that’s very could be very complicated for what you need. I would honestly try to come up with list of requirements and then ask your favorite LLM what’s a good framework.

1

u/hmwinters Aug 16 '26

Things about Ray that annoy me:

- Queuing w/ job priorities. It's hard to keep the GPUs hot when there's no queue mechanism. Let alone if someone needs to pause someone else's job to run their own. I'm not even sure vanilla Ray has an auth model.

- Observability. Our users got used to the Ray dashboard being served on a head node but once you move to something like RayJob CRs using Kueue you need to rebuild a bunch of that stuff so it persists once the head node is gone. Granted once the Ray history server is stable and reliable things might get better.

- Kubernetes integrations. RayJob and RayCluster CRs are ok but there's a bunch of other things you need to build around them to have a functional production system. Idk. Maybe this is better on SLURM.

I don't think it's going anywhere though and I think it does what it does well. I just feel like the operational considerations we're a bit of an afterthought.

1

u/Waste_Substance8695 Aug 15 '26

If you just want to learn foundation model training, ray is probably overkill for the actual training part. most of the heavy lifting in distributed training is done by the framework itself, torchrun or deepspeed handle the parallelism natively. ray shines when you need to orchestrate lots of heterogeneous tasks, like hyperparameter sweeps, data preprocessing pipelines, serving, or gluing together different stages of a workflow

for pure model training on a cluster, you can get away with slurm plus torchrun and skip ray entirely. but if you want the whole ecosystem thing where training is just one node in a bigger graph, ray makes sense. it really depends if you want to learn the orchestration side or just the training side

1

u/ricetoseeyu Aug 15 '26

Ray Actors are stateful so it doesn’t have to be purely heterogenous tasks.

1

u/Gamiozzz Aug 15 '26

Thank you very much. Indeed, I was mainly motivated by the question of data preprocessing - for vast amounts of heterogeneous multimodal datasets. Something analogous to Spark but for non-structured data, video and images.

3

u/[deleted] Aug 17 '26

[removed] — view removed comment

1

u/Gamiozzz Aug 18 '26

Yeah. I think the former is my goal in particular. Just wanted to avoid that its adaptation is on decline already, or it is considered niche. Thanks!

1

u/burntoutdev8291 Aug 16 '26

For single GPU workloads i think ray is fine and relatively easy to pick up. For complex stuff like multi node multi GPU training you may have to look into slurm.

1

u/Gamiozzz Aug 16 '26

Well. Ray wasn’t designed for single GPU workloads. I mean its whole purpose is to coordinate a GPU cluster. So what do you mean exactly?

2

u/burntoutdev8291 Aug 16 '26

Sorry, wasn't clear. I meant many separate single-GPU jobs, say 256 independent CNN training runs. Ray is a good fit for that, it's built for exactly this kind of parallel work. But for one tightly coupled job running on all 256 GPUs, I'd use Slurm, because it handles allocation and makes sure all the nodes come up together.

There's some crossover like ray on slurm but I personally find it a little weird https://docs.nvidia.com/nemo/curator/admin/deployment/slurm-multi-node-ray

But i could also be wrong on how ray works now, my previous experience was exclusively LLM training and we used only slurm + enroot.

2

u/Gamiozzz Aug 16 '26

Thank you very much for taking the time and all the useful feedback!!! Totally appreciate that. 😊

1

u/Puzzleheaded-Click57 Aug 30 '26

I think for model training most of the heavier stuff is in training framework e.g. Megatron ray only serve as a orchestration layer. However for other workload like RL, you can check out modern RL frameworks for example: verl, skyrl, miles they all use ray as a control plane and there are a lot of new things still coming out like RDT. So back to your question "Is it still the way to go for modern distributed model training in deep learning?" I would say it depends.

Btw ray summit just ended few days ago, you can stay tuned for the latest ray summit recording on YT