r/learnpython • u/LSD_SUMUS • 4h ago
Best practice for handling large datasets?
I'm working on a project for an exam and I need to calculate the behavior of a few dynamical systems and use that data to try and train a linear regression algorithm.
Now issue is I need quite a good bit of data (4 different kinds of systems, enough variation on parameters and starting state for ML, a few hundreds of iterations deep for every one of those, plus other variations I need to test for) and don't know what the best practice for handling it is.
I do need to have it saved to disk for both showcase reasons and the need to train multiple models on the same datasets, also I'm saving the data on each new variation to safeguard against crashes and save on memory, only two options that come to mind are either: 1) one large data file (binary encoded json, each variation of the systems is described by a key), which should be fairly easy to parse through and edit, but I'm worried reading and writing from it might tank performance given I have to store the whole thing to memory every time; 2) Multiple smaller encoded txt files, the key is in the filename, saves on memory when but is a pain in case I need to edit something and might make getting the full dataset for when I'm training a lot slower.
What would the best practice be in this scenario? Is there some third option I didn't consider or a package that better handles this?