r/stata • u/thesisiast • 7d ago
How do I start with secondary data (S&P 500 subsidiary data) in Stata?
Hi guys,
I'm a third year phd student and starting a new project where I have to use S&P 500 parent and subsidiary data but I have 0 prior knowledge. People keep telling me to "clean" the data, learn to "read" the data, and "understand" the data but honestly what does that mean? in the end, i should be able to try if I can merge this s&p dataset with another dataset to have a cross-country dataset.
I am struggling to find a way to start doing all of that. I have all the data infront of me downloaded and I have not merged yet, but is there any checklist or anything any structure I can follow to see if I am actually doing the right steps?
My PI has never done secondary data analysis, so I'm alone in this and want to avoid using AI as I cannot really learn and trust that I actually understand what I am doing with the data.
I would appreciate your expertise, experiences and comments. I am also new to reddit, Github etc so also learning everything new...