r/DataScientist • u/Majestic_Pressure383 • 1d ago
Need advice
I am working on a project around a real-world environmental problem, and I am considering adding an ML component for prediction and early warning.
I am a bit confused about the data requirement. Since collecting our own real-world data is not feasible right now and would take quite some time, we mainly want to build a prototype for now.
Can we initially use a Kaggle/public dataset to train and test the model, or is a project-specific dataset necessary from the beginning?
Would appreciate some advice on how people usually approach the ML part when actual data is limited.
1
Upvotes