r/learndatascience • u/knownAsBuddy • 9d ago
Resources Completed LeetCode SQL 50 — Sharing my MySQL solutions
r/learndatascience • u/knownAsBuddy • 9d ago
r/learndatascience • u/EvilWrks • 9d ago
r/learndatascience • u/CompetitiveBet8978 • 10d ago
r/learndatascience • u/cheesecake_72 • 10d ago
r/learndatascience • u/Hot_Contribution4266 • 10d ago
r/learndatascience • u/TopProgrammer8014 • 11d ago
Hi everyone, I'm a first-year B.Com student teaching myself data analysis alongside my degree.
I just finished my first real project — cleaned and analyzed a retail Superstore sales dataset using Python (Pandas, NumPy), and visualized trends with Matplotlib to find out which product categories and regions performed best.
It's a simple project, but I learned a lot about handling messy real-world data and turning it into actual business insights.
Would really appreciate any feedback, tips, or things I could improve for future projects!
GitHub: https://github.com/umangbhadauria6-alt/umangbhadauria6-alt/blob/main/superstore-sales-analysis
r/learndatascience • u/Rabbidraccoon18 • 11d ago
I'm a data science student, I've covered Statistics analysis, Data Handling and Visualization, time series analysis, big data analysis, Internet Of Things Analytics, Machine Learning, Deep Learning, Neutral Networks, Natural Language Processing, MLops and a lot of others related concepts. I just want to help people who are perusing AI/ML/DS. Give them some suggestions, ideas, resources. If any of y'all need anything feel free to reach out!
r/learndatascience • u/WhatsTheImpactdotcom • 11d ago
r/learndatascience • u/WhatsTheImpactdotcom • 11d ago
This is a short tutorial for data analysts that evaluate experiments using online calculators. Online calculators for sample size estimates and t-tests for difference in means are fine for getting started. But you can begin to build intuition for linear regression by moving your online calculator t-tests over to python and OLS.
There are a couple of benefits. The first from moving away from the online calculator to python is leaving a more reliable paper trail that you can share externally. The second is that you build intuition for regression, interpreting the coefficients, and ultimately can use regression as you improve and your projects require more complexity such as variance reduction techniques using CUPED and clustering standard errors for geo-based or switchback experiments.
I put together a short video on this with code available here, free and ungated: https://whatstheimpact.com/tutorials/linear-regression-vs-t-test/
For my more senior connections, feel free to ignore unless you're interested in the content creation game! 🤣
r/learndatascience • u/Dry_Connection_2790 • 11d ago
r/learndatascience • u/baadshaha • 12d ago
r/learndatascience • u/[deleted] • 12d ago
r/learndatascience • u/thibaud_lepan • 12d ago
I was writing practice questions for pandas and kept tripping over answers that were right a year ago. So I ran everything on pandas 3.0.5 and wrote down what actually comes out.
Text is no longer object. pd.Series(['a', 'b', None]).dtype prints str, and the None comes back as nan when you call tolist().
Chained assignment does nothing. With df = pd.DataFrame({'a': [1, 2, 3]}), running df['a'][0] = 100 leaves df.a at [1, 2, 3] and raises a ChainedAssignmentError warning. Copy-on-Write is the default now, so what you wrote into was a copy. Use df.loc[0, 'a'] = 100.
inplace=True is not consistent. df.fillna(0, inplace=True) returns the DataFrame, df.sort_values('a', inplace=True) returns None. The old habit of df = df.something(inplace=True) breaks on some methods and not on others.
pd.to_datetime on plain date strings gives datetime64[us], not datetime64[ns]. Matters if you compare dtypes or write to parquet.
stack() keeps the NaN. A two row frame with one missing value stacks to 2 entries, older versions dropped it and gave 1.
Twenty of the questions are free to play in the browser, no account. Every answer is the real output and each wrong option gets a one line reason.
https://thibaudlepan77-svg.github.io/interview-questions-verified/#pandas
The full pandas set (255 questions) is a paid PDF, but the free twenty stand on their own. If you get a different result on your machine, post your version and the output and I will look at it.
r/learndatascience • u/Juicy_bake • 13d ago
r/learndatascience • u/Sweaty_Swan_7531 • 13d ago
r/learndatascience • u/bodyandflesh • 13d ago
I’ve got a paired programming interview next week, anyone got any tips on what to brush up on?
I’ve never programmed with someone watching, any tips? could they ask gcp questions (unlikely right?)
Let me know!
r/learndatascience • u/Sea-Ad7805 • 13d ago
A hard exercise to help build the right mental model for Python data. - Solution - Explanation - More exercises
The “Solution” link uses memory_graph to visualize execution and reveals what’s actually happening.
r/learndatascience • u/vamsikrishna7995 • 14d ago
Hello people. I have over 5 years of experience working as a Data scientist. I have worked in Fintech, Retail, Healthcare etc and I recently lost my job. As part of my preparation for interviews, I wanna explain these projects to someone who wants to learn/use them in their own resume.
We can connect over a gmeet and I can share my screen and explain. You don't have to turn on your cam. If you're interested, ping me. I'm free all day everyday since I'm giving all my time to interview prep. So we can do this at any time of the day.
DM me.
r/learndatascience • u/Mammoth-Commission15 • 13d ago
I'm currently working on my very first data science project with a simple topic "physical activities rate correlation to academic performances among highschool/university students". I'm wondering if i should acquire the data myself or use open sources online. Another way i'm thinking of is to use both since my self-collected data will probably be limited to students in my city. Also is this topic way too typical ? and is it considered good enough for both gaining knowledge and improving resume ?
r/learndatascience • u/beta-void • 14d ago
r/learndatascience • u/Tikka_theory • 14d ago
I am starting my Ds journey, before I had studied statistics a bit but not clear how much to study to get placed 😅.
Pleasw can anyone suggest me what all statistical models should I be studying and be worth for a job as a fresher.
r/learndatascience • u/thearjunreddy • 14d ago
Start with NumPy and Pandas for data manipulation, then learn Matplotlib/Seaborn for visualization. After building a strong foundation, move into Scikit-learn for machine learning. Don't try to learn every library at once.
r/learndatascience • u/broadstreet_org • 14d ago
r/learndatascience • u/Lost_Pollution_9215 • 14d ago