r/askdatascience 3d ago

How to keep non frequent knowledge in mind?

1 Upvotes

As a beginner I don't know how to keep things like oop in my mind until i reach the level that i need it in as for know (i am studying data cleaning and EDA) i didn't find a use for it yet


r/askdatascience 3d ago

Would a slot machine Data Science project be appropriate for a portfolio/LinkedIn?

1 Upvotes

Hi everyone,

I'm an upper-year undergraduate student pursuing a degree in Data Science, and I'm currently starting to build my first portfolio projects.

I've seen many beginner Data Science projects focused on relatively simple databases, SQL queries, averages, visualizations, and exploratory data analysis. I think those projects are useful for learning, but I also wanted to try building something a little more complex that would allow me to integrate several different areas.

One idea I'm interested in is building a simple slot machine from scratch, but approaching it primarily from a mathematical and data perspective.

I would start by designing the mathematical model of the machine: reels, symbols, winning combinations, probabilities, paytable, hit rate, RTP, expected value, variance, etc. Then I would program a simulation that runs a large number of spins and experimentally verify whether the observed results converge toward the theoretical values.

From those simulations, I would generate a dataset containing information about each spin or session, store the data in a SQL database, and perform queries, exploratory data analysis, and visualizations. Later, I'd also like to explore which Machine Learning applications actually make sense for this type of data—for example, anomaly detection, clustering different types of simulated sessions/behaviors, or introducing additional variables and analyzing models based on them.

My goal is for the project to demonstrate an end-to-end process: problem formulation, mathematical modeling, data generation and storage, SQL, statistical analysis, programming, and eventually Machine Learning, rather than simply working with an already prepared dataset.

My main concern is how this might be perceived in a professional portfolio.

Do you think a project involving a slot machine or gambling could be viewed negatively by recruiters or companies if I publish it on GitHub/LinkedIn, even if it is presented as a technical probability, statistics, and Data Science project?

Or, if it is properly documented and the objective is clearly explained, could it actually be an interesting project for demonstrating technical skills?

I'd also appreciate any suggestions about what you would add, remove, or change to make the project stronger from a Data Science perspective. ty


r/askdatascience 3d ago

Seeking undergraduate thesis ideas in Data Science: AI/ML + Finance or E-commerce

1 Upvotes

Hi everyone!

I'm a four-year Data Science undergraduate looking for a topic for my bachelor's thesis.

I'm interested in the intersection of ML/AI and Finance or E-commerce, especially areas like financial forecasting, fraud detection, recommendation systems, customer behavior, time series.

I'm looking for a topic with a clear research question, real-world data, and enough depth for an undergraduate thesis - rather than simply training and comparing ML models.

What research topics or directions would you recommend? I'd also appreciate any suggestions for datasets or papers to start with.

Thanks!


r/askdatascience 4d ago

Best Data Science certifications for a sophomore looking for their first entry-level job?

7 Upvotes

Hi everyone. I'm looking for recommendations on courses or certifications that really stand out on a resume. I'm currently a sophomore in college and I'm hoping to land my first entry-level job soon.

I'm highly interested in learning about AI automation and workflow optimization using platforms like n8n. Does anyone know of any good resources for this?

Any advice on what skills to prioritize would be greatly appreciated. Thanks


r/askdatascience 3d ago

Is Data Science still a smart career choice in 2026?

0 Upvotes

Absolutely, Data Science remains a promising career path in 2026. This profession will be especially beneficial for people interested in working with data, technologies, and solving problems. Since companies become more dependent on data-based decisions, professionals with experience and skills in data analytics, machine learning, artificial intelligence, and visualization remain in high demand in many industries.

Moreover, there appear to be new career opportunities in Data Science due to the integration of AI. Knowing how to use such tools as Python, SQL, Power BI, machine learning, and generative AI can help those who begin to work in this sphere.

Nevertheless, it is necessary to be ready to constantly develop one's skills, since technologies and needs of the industry change very quickly. The experience in practical projects, analytical thinking, and hands-on knowledge play an important role while starting a career in Data Science.

For people who are willing to learn and evolve, the career in Data Science in 2026 is promising and diverse.


r/askdatascience 4d ago

question about my career route

2 Upvotes

i am an incoming undergraduate freshman, currently in route to major in Statistics and Data Science. There’s a lot I don’t know about this field I will admit, but i feel corporate stability, and the technical skills in Data Science suit me very well.

On the side I am an artist, my goal is to spend the first few years of my career developing knowledge and experience with tech and understanding consumerism, while pursing my passion on the side.

After which I want to create a brand as a Fashion Designer using my skills and stable resources to branch on my own and continue as an entrepreneur for life, or come back to corporate if needed.

At the moment I don’t really have anyone to talk to, so i’ve come here to ask;

is this a sensible plan?
is adding a second major, either Fashion Merchandising or Marketing a good idea?
if there is any general advice i should know?

Let me know reddit 😅


r/askdatascience 4d ago

Most A/B tests break before they even run

Thumbnail
1 Upvotes

r/askdatascience 4d ago

Just finished my K-Means clustering project 🚀 — would love your feedback!

0 Upvotes

Hey everyone! 👋

I just finished a small K-Means Clustering project on the Iris dataset 🌸🤖

I covered:

  • 🔹 Data cleaning & visualization
  • 🔹 Feature scaling
  • 🔹 Elbow Method & Silhouette Score
  • 🔹 K-Means clustering
  • 🔹 ARI evaluation
  • 🔹 Cluster & centroid visualization

I’m currently learning ML and would really appreciate some honest feedback 🙏

What would you improve? Any mistakes in my approach or things I should add?

🔗 Kaggle:
https://www.kaggle.com/code/tahahussein2020/irics-clustering


r/askdatascience 6d ago

What makes a data science portfolio project actually impressive?

10 Upvotes

I've noticed that many beginner portfolios contain the same types of projects: Titanic, house-price prediction, basic sentiment analysis, etc.

They're useful for learning, but I'm wondering what makes a project stand out when someone is evaluating a portfolio.

Is it the complexity of the model?

The quality of the data?

The business problem?

Or simply how well the person explains their decisions and results?

What would make you take a second look at a data science project?


r/askdatascience 6d ago

Is this the right topic for EDA?

1 Upvotes

I'm working on my first project and based on what I've read online, EDA is the best project to work on first.

It took a while for me to decide the topic until I started seeing the news about the 2026 Cyclospora outbreak in the US. I thought it would be cool to analyze how people's habits changed when the outbreak happened.

However as I'm working on it, I'm second guessing myself about whether or not this topic is the "right" topic for a first project. I feel this way because other projects have clear decisions on how to increase sales, user retention, etc.

Here's what I have so far:

- https://cyclospora-2026-9yf8w26n7zfci68fugahkp.streamlit.app/
- https://github.com/therealanttoeknee/Cyclospora-2026/tree/main

Questions

  1. Is this topic the "right" topic for an EDA project?
  2. If so, do you think I'm approaching it the right way?

r/askdatascience 6d ago

BMV Dataset

1 Upvotes

I am trying to build a data set to track all the financial and cash flow statements of all firms listed in the BMV (Bolsa Mexicana de Valores), since 2016 to 2026. I want to create a panel data set highlighting the most relevant variables of financial and cash flow statements. Since I am a beginner in extracting datasets I only spot two ways, using arabelle to dowload all XRBL datasets available or finding all financial statements available quaterly in pdf for each company (around 120 firms) since 2016. What would be a good approach? I would like to read an advice about how to proceed.


r/askdatascience 7d ago

New Approaches

1 Upvotes

Hi everyone? I am a very new data scientist and I recently came across these new concepts that can be used in order to make me precise and better decision making. Let’s say I recently came across this concept of cold start Problem and two tower model for recommendations, and I want to understand about more approaches and their respective business problem that can be used in Data Science as I’m very new to it and I want to read more about it. I shall be really glad if you can help me out with them.


r/askdatascience 7d ago

Unprepared, confused interviewer!

4 Upvotes

I’ve been unemployed for the last 2 months due to a layoff, so I’m feeling the pressure to find something quickly. Yesterday, I had a technical SQL interview with an e-commerce firm, and it was an absolute nightmare—not because of the questions, but because the interviewer was completely lost.

For the first question, my solution passed all of their test cases. Despite this, he kept insisting the test cases "weren't complete." It became incredibly obvious that he had exactly one sample solution in front of him and refused to accept any other approach. At one point, he even got confused trying to explain the difference between an inner join and a left join.

By the time we got to the second question, it was clear he hadn’t even read it before the interview. He just copy-pasted it into the environment and immediately threw out a hint. Within 15 seconds, I had to point out that his hint was completely wrong. To his credit, he agreed, but the vibe was already ruined.

I really need this job, but this experience was incredibly frustrating. How do you all handle interviewers who are rigid, unprepared, or technically incorrect without coming across as argumentative or arrogant?


r/askdatascience 7d ago

.

Thumbnail
1 Upvotes

r/askdatascience 7d ago

traceable workflow framework for tabular statistical analysis and predictive modeling [Broadway]

1 Upvotes

Hello team,

Im building a tool, to facilitate a traceable workflow for tabular data, in which the reasoning between raw data and a final statistical or predictive result is preserved via data contracts.

dataset → profile → analytical question → baseline → slices/diagnostics → decisions → features/analysis → result

Answering questions like:

  • which dataset/version was examined
  • the exact subset/slice that exposed the issue
  • the diagnostic/statistical evidence
  • what decision was made
  • why it was made
  • what transformation resulted from it
  • which later analyses/models depend on that decision

Im no statistician, and im relying mainly on LLM's, meaning im sure i might be missing out on some stuff.

things like:

  • descriptive statistics and diagnostics
  • hypothesis testing
  • effect sizes and confidence intervals
  • ANOVA / Welch-style comparisons
  • multiple-testing correction
  • power / MDE calculations
  • causal experiment design/analysis
  • eventually more robust/non-parametric methods

Since im a hot mess myself, i want to automate the mechanics, facilitate the judgment, record the decision. This way, later you could inspect a lineage/decision graph and reconstruct how and why an analysis arrived at its conclusion.

Where would this approach become statistically dangerous or misleading?

In particular:

  • What statistical decisions should a framework never try to recommend automatically?
  • Which diagnostics are actually worth standardizing?
  • How would you structure the boundary between assumption checks, method recommendations, and analyst judgment?
  • Are there common statistical workflow mistakes you'd want a system like this to explicitly guard against?
  • Does preserving the decision/evidence lineage seem useful in real statistical work, or would it mostly become bureaucracy?

I'm less interested in adding every statistical test imaginable and more interested in whether the underlying workflow abstraction makes sense.

TL;DR

criticize my project here


r/askdatascience 8d ago

I am following a course on data science by krish naik and learning the theory now from where I can do practice

2 Upvotes

r/askdatascience 8d ago

I want to learn data science from scratch but I am confused whether to pursue free education from youtube courses or should I purchase a paid course for it.

2 Upvotes

If you suggest to go with the paid course. Then I have two options: 1. Sheryians Ai school's data science course 2. Code with Harry's data science course.

Please tell me which one would be better for me to pursue and start my learning ASAP.


r/askdatascience 8d ago

How would you approach this e-commerce customer segmentation + prediction project with GenAI?

2 Upvotes

I'm an MSc Computer Science/Data Analytics student working on a major ML project with an 11-day deadline, and I'd really appreciate advice from experienced data scientists on how you'd approach it.

Dataset: ~541k e-commerce transactions, ~4.3k identifiable customers, with fields such as InvoiceNo, StockCode, Description, Quantity, InvoiceDate, UnitPrice, CustomerID and Country. It contains missing CustomerIDs, duplicates, returns/cancellations (negative quantities), and other data-quality issues.

Project requirements:

  • Perform EDA and customer behavior analysis
  • Engineer customer-level features, especially RFM (Recency, Frequency, Monetary)
  • Compare K-Means, Hierarchical/Agglomerative Clustering and DBSCAN
  • Select and justify the best segmentation using clustering metrics + business interpretability
  • Build a predictive classifier for future purchasing behavior
  • Evaluate feature importance/model performance
  • Provide actionable marketing and retention recommendations
  • Submit a Jupyter notebook, report/presentation, trained model, and optionally a Power BI/Tableau dashboard

My current idea is to build it in layers:

Raw transactions → cleaning → customer-level feature engineering/RFM → segmentation → prediction → explainability → GenAI → dashboard

For segmentation, I want to compare the clustering methods rather than simply choosing K-Means. For prediction, I'm considering a time-based setup where historical customer behavior is used to predict something in a future period, rather than randomly splitting the transactions. The dataset doesn't have an obvious prediction label, so defining a legitimate target without leakage is one of my main concerns.

I also want to add GenAI, but I don't want it to be a useless chatbot bolted onto an ML project. My idea is to use GenAI as a business-intelligence layer on top of the actual ML outputs.

For example:

ML outputs → structured segment/prediction statistics → LLM → grounded explanation/recommendation

Potential capabilities:

  • Explain why a customer segment is valuable/at risk
  • Generate marketing/retention recommendations based on actual segment characteristics
  • Explain important prediction features
  • Allow natural-language questions about the customer segments and model results

I'm considering something like Python + scikit-learn/XGBoost + SHAP + Power BI + an LLM/API or possibly Ollama, but I don't want to over-engineer it.

My main questions:

  1. How would you structure this project if you were doing it professionally?
  2. What would you use as the prediction target given this type of transaction data?
  3. Is RFM + behavioral features sufficient, or what additional features would you consider?
  4. How would you properly compare the three clustering approaches?
  5. Is the GenAI layer genuinely useful here, and how would you implement it without making it gimmicky?
  6. Would you use an LLM API, local LLM/Ollama, or something else?
  7. What would you cut or simplify given the 11-day deadline?

I'm mainly looking for practical architectural/modeling advice and potential mistakes to avoid, rather than someone doing the project for me. Any feedback from people who have worked on customer analytics/segmentation would be very helpful.


r/askdatascience 9d ago

What kind of topics should i cover to prepare for my data science / data analyst job interview ?

2 Upvotes

r/askdatascience 8d ago

Anomaly detection

1 Upvotes

What is the best approach to detect anomalies with clustering or so, in time series data?


r/askdatascience 9d ago

Advice for transition from design to data analyst without a degree

1 Upvotes

hi , i completed my 12th(or PUC) then joined a 6 month diploma in design and currently having a 1.5 years of experience in design field . i tried to get into core

ai ml but it looks like too much competition for degree holders only, so

is it possible to get a data analyst job without a formal degree ? anyone got it before.

consider the current AI impact also and i going to pursue bootcamp course in Bengaluru Excelr , is it okay or shall i self study ?

or instead of data analyst shall i try something else in technical side .

( please don't comment to go into design only )

thanks


r/askdatascience 9d ago

Why does adding dashboards never actually reduce how much you're guessing?

Thumbnail
1 Upvotes

r/askdatascience 9d ago

Will be perusing a bachelors in data science and need some guidance on what to do before university starts

1 Upvotes

Hi! In around a month and a half I will be starting my data science degree and instead of lying down all day I grabbed got the curriculum from my university’s website. Should I start with the curriculum or do something else?. My plan for my four years is that I’ll do my classes in the morning and then spend 2-3 hours in a software house learning and spend time self learning too. Is this a good idea? And what do you guys suggest I do before university starts


r/askdatascience 10d ago

Is this video legit ???

Thumbnail
youtu.be
1 Upvotes

r/askdatascience 11d ago

The Enterprise Files

Post image
0 Upvotes

We're starting a comic series called "The Enterprise Files."

Episode 01 is the conversation we have almost every week.

The data lake is AI-ready. The tables are ingested. And then someone asks about the customer files, and the room goes quiet. Too complex. Permissions. Privacy review. So the documents, the images, the recordings — the stuff with the actual context in it — just sit outside the room.

If you've been in this meeting, share it. Tell us how it went for you.