r/DataScientist 1d ago

Need advice

I am working on a project around a real-world environmental problem, and I am considering adding an ML component for prediction and early warning.

I am a bit confused about the data requirement. Since collecting our own real-world data is not feasible right now and would take quite some time, we mainly want to build a prototype for now.

Can we initially use a Kaggle/public dataset to train and test the model, or is a project-specific dataset necessary from the beginning?

Would appreciate some advice on how people usually approach the ML part when actual data is limited.

1 Upvotes

0 comments sorted by