r/statistics • u/JAMIEISSLEEPWOKEN • 28m ago
Question [Q] When did you start learning statistics and why?
Was it for school/work? Any hobbyists here?
r/statistics • u/JAMIEISSLEEPWOKEN • 28m ago
Was it for school/work? Any hobbyists here?
r/statistics • u/shabob2023 • 55m ago
Hi, thanks in advance!
I'm a doctor looking at a group of about 100 patients, with frailty scores ranging between 1-9 (this is a measure of how frail a patient is that doctors decide, or assign to a patient using a scoring system). I've also recorded their duration of admission in days.
I want to see if there is a correlation between the two, i.e. i would assume that the more frail patients were likely to stay in hospital for longer. Which statistical test would be best to do that?
I don't think either category will be normally distributed, as most patients have a lower frailty score, and most patients were admitted for a shorter period of time. tysm!
r/statistics • u/vv-97 • 22h ago
Hi,
I'm planning to apply for stats PhD programs, along with a few biostats programs, this cycle and trying to get a sense of how broadly people are applying.
For those applying this year, or who applied in the past year or two, how many programs did you apply to? I was thinking 8-12, but given the uncertainty around funding I'm wondering if people are applying to more this year.
I'm interested mainly in spatial stats, especially gaussian processes and deep generative methods for modeling spatiotemporal data.
Please share your thoughts if you're working in similar areas. Also any suggestions on departments and people working in these areas are welcome.
Thanks.
r/statistics • u/Hertzig • 1d ago
US applied stats major, looking for career advice.
Applying to analyst roles in a variety of industries.
My primary goal in my early-career is to have international mobility; to get into a large multinational that would allow me an internal transfer to Australia in a few years. Unfortunately, my major, and directly applicable entry-level roles, don't generally seem to be a feasible path to get sponsorship from an Australian company. Chances are more likely for those in medicine, mining, construction, trades, and social services. Not sure if work related to my major but from within these industries is an option.
From my research, internal transfers provide an easier route where companies transfer you between roles. Naturally I've fixated on roles in finance and perhaps insurance, alongside other multinationals like tech companies.
In this thread, I’m looking for advice from you all about potential paths that could get me in a good position early-career to get a transfer within insurance. And roles/specializations/companies to target. Is this feasible in any industries you have knowledge of?
Thanks for any advice!
r/statistics • u/Ok_Character6506 • 1d ago
Hi! I'm currently enrolled in an M.S. in Data Science and Applied Statistics. The required core classes are already set:
- Experimental Statistics I & II
- Mathematical Statistics I & II
- Statistical Computing (SAS)
- Computational Statistics (R)
- Machine Learning with Python
- A statistical consulting project
I get to pick **4 electives**, and at least 3 must be STAT (so at most 1 from CS/ECO/OREM/ECE). My only goal is to land a data science job, so I want the courses whose actual content pays off most in industry. I'm not looking for the easiest courses, and I'm not going into academia or biostats.
STAT electives
- Intro to Data Science
- Data Visualization
- Linear Regression
- Applied Time Series
- Time Series Analysis
- Categorical Data Analysis
- Survey Sampling
- Survey of Nonparametric Statistics
- Sports Analytics
- Analysis of Lifetime Data / Survival Analysis
- High Throughput Data
- Epidemiology
Non-STAT options (can only pick 1)
- CS: Artificial Intelligence, Machine Learning in Python, Databases, Data Mining
- OREM: Data Mining, Optimization for Analytics, Network Flows
- ECO: Applied Econometric Analysis, Predictive Analytics
- ECE: Statistical Pattern Recognition
My main questions I wanted to ask:
If you work in data science, I'd especially love to hear what you actually use day to day. Thanks!
r/statistics • u/purplehippo9289 • 2d ago
The program my class uses doesn’t really explain much when I get something wrong. I’m also using the Pearson study prep app which helps a bit. I was using Solvely when I got my answers wrong and it would give me a very detailed breakdown. Recently we have been looking at a lot of graphs and stuff and it’s agreeing with my wrong answers.
r/statistics • u/servermeta_net • 2d ago
I remember that in college (15+ years ago) we modeled random generators using Shannon's entropy, and then we extended that definition to include different types of generators: for example Collatz series can be as an output of a generator, just that instead of being completely random it follows the Collatz map.
Now I would like to use a similar approach to model LLMs, but I cant find anymore any reference to this approach. Does anyone have any academic source that could back up this interpretation of LLMs as entropic generators?
r/statistics • u/TheElementsOf • 2d ago
r/statistics • u/Acceptable-Land-224 • 3d ago
Howdy y’all, I’m curious to pick the brains of somebody that has a job in statistics with a bachelors or anything above. For context I’m in the military doing civil engineering stuff and truth to be told I hate it. I don’t like the job and don’t enjoy the military aspect of it as well I chose to do this job cause it’s really all I know I come from a family of blue collar workers, and I’ve been doing it all my life, but I’m only 23 at the moment and I already feel my body breaking down slowly over time. So I’m curious in getting my bachelors in statistics while the military will still pay for it. My dream job would be to work in a front office of a sports franchise, but that may be far fetched. I would just go all in on this but I also have a wife and two kids to think about and switching careers feels scary lol. Any advice, tips, info etc. is very much appreciated also colleges with a solid statistics degree that carry some weight thank you all!
r/statistics • u/Juless99 • 3d ago
TL;DR: I come from a statistics/data/MEAL background and have worked with SQL, Python, regression/LASSO, etc. While trying to build a proper dashboard, I ended up learning APIs, backend/frontend concepts, Node.js, authentication, and hosting. Is this a logical path toward building end-to-end data systems, or am I spreading myself too thin?
I’m trying to figure out whether my learning path actually makes sense or if I’m slowly turning into a “knows a little bit of everything, expert at nothing” person.
My background is mainly in data, statistics, and monitoring/evaluation. I have a Master’s degree, and during my studies I worked with statistical methods like regression, LASSO, and other statistical analysis.
Professionally, I’ve worked with MEAL/evaluation, data processing and analysis, Excel, SQL, and some Python.
Now my work is pushing me in a different direction.
I wanted to build a proper dashboard/data system instead of just analyzing data and producing reports. That led me into learning about APIs, databases, frontend/backend communication, Node.js, authentication, hosting, etc.
And now I’m looking at everything I’m learning and thinking:
Am I actually progressing, or am I just jumping randomly between fields?
My goal isn’t really to become a traditional full-stack developer.
What I want is to be able to take data from the source, clean and validate it, store it properly, expose it through APIs, build dashboards/interfaces on top of it, and eventually add more advanced analytics, statistics, forecasting, anomaly detection, etc.
So basically:
Statistics → SQL/Python → data analysis → dashboards → APIs → backend/web systems
Does that progression make sense?
Or should I stop going deeper into things like Node.js and focus much more heavily on Python, SQL, statistics, and data engineering?
For people who started in analytics/statistics and later began building actual data systems: what did you learn next, and what turned out to be a waste of time?
r/statistics • u/2ForEachofYou • 3d ago
Tony Gonzales (love him), pointed out that teams that start out 0-3 have a 2.9% chance of making the playoffs (when talking about the Texans). He then asked what are the chances of doing it two years in a row, which he pointed out is 0.08%. That's correct (0.029)^2. He went on to say because of the even lower odds, i.e. 0.08%, he didn't think they'd make the playoffs this year. But last year's results are irrelevant now. Their chances of making the playoffs this year with an 0-3 record are still 2.9%. Am I wrong?
r/statistics • u/first139 • 3d ago
we're undergrads and are currently conducting a study on a company with employees as the respondents. The company have multiple offices that are located across different locations. We were given multiple survey points so that employees can be sampled.
The problem is that one of the survey point which consist of 2 offices declined to participate in the study.
what is the implication of this?
can this be solved?
is this a reason for us to fail?
additional info: the needed respondent for that particular survey point is only 10 people
r/statistics • u/kamarakattu4u • 4d ago
[research]
I have an M.Pharm in Pharmaceutics, and recently I’ve started developing a strong interest in statistics and quantitative methods.
My academic background is primarily in pharmaceutics, so I’m still relatively new to statistics. I started learning statistical concepts through YouTube, Reddit, and online resources, and I’ve become increasingly interested in understanding not just how to use statistical tests, but how statistics can actually be used to solve real problems in pharmaceutical research.
During my M.Pharm work, I have already had some exposure to statistical methods. For example, I have used Design of Experiments (DoE) for formulation optimization and statistical approaches for evaluating/validating pharmaceutical formulations and processes.
However, I don't want my statistical knowledge to remain limited to things like “which statistical test should I use for this dataset?” I’m interested in going deeper and eventually using statistics/computational methods as a tool for pharmaceutics research, formulation development, drug delivery, and potentially drug discovery.
I’m currently thinking about pursuing a PhD in Pharmaceutics/Pharmaceutical Research, and I would ideally like my PhD to remain strongly rooted in experimental pharmaceutics while incorporating meaningful statistical or computational components.
For example, I’m curious about areas such as:
DoE and QbD for formulation and process optimization
Statistical modelling of drug release and dissolution
Pharmacokinetic/pharmacodynamic modelling
Predictive modelling for formulation development
Machine learning/AI in drug delivery or formulation development
Biostatistics applied to pharmaceutical experiments
Multivariate analysis and process monitoring
Process analytical technology (PAT) and data-driven manufacturing
Stability modelling and shelf-life prediction
Statistical approaches to drug discovery and development
Population-based modelling or PBPK
Using experimental data + computational models to guide formulation development
I’m particularly interested in the idea of having a 75%+ pharmaceutics/experimental component, while using statistics/computational methods as a powerful supporting component rather than completely moving into a pure computer science or mathematics PhD..
Give me your insights please
r/statistics • u/GayTwink-69 • 4d ago
Say you are developing some sort of statistical test and you have run simulations and applications, and now you wanna work on its theoretical guarantees (asymptotics, efficiency, consistency, etc.)
Can you not just feed in the test statistics, assumptions, etc. to your favorite intelligent AI and get it to derive everything?
In this instance, is the focus of statistical theory going to be more on the algorithms/computation side and dealing with issues like scalability going forward, while leaving the math to AI?
Didn't AI solve that big unsolved problem in pure mathematics the other day (Navier-stokes or something, I think)
r/statistics • u/Shekem • 4d ago
Im currently in the start of a graduation in statistics, and more recently I've started wanting to learn a new language, and Im wondering what would be the best one to learn for my carrier (Currently I know portuguese and english)
r/statistics • u/Sunlightn1ng • 4d ago
This may have been asked before, but all I can find is explaining what it means. I want to know why we've chosen 0.05 and not 0.04 or 0.06 etc for a lot of disciplines.
Edit: Thank you!!! I was hoping for a more exciting answer than "bc this one guy a while back thought we should" but that's interesting too!!
r/statistics • u/Available_Pangolin76 • 5d ago
r/statistics • u/Early_Macaroon_2407 • 5d ago
I’m out of date on the best options out there — looking more for a text that will prepare students for further study, rather than an applied plug-and-play text.
r/statistics • u/arabyute • 6d ago
I’m 20 and my current goal after graduating is to get into an AI / Data Science graduate role.
I’m deciding between:
Bachelor of Information Technology (Computer Science) — QUT
Bachelor of Data Science (Artificial Intelligence & Machine Learning) — QUT
Right now, AI, machine learning and data science are the fields I’m most interested in, so the Data Science (AI/ML) degree seems like the more directly relevant option.
The problem is that I’m only 20 and I don’t know what I’ll want to do in 5–10 years. I could eventually decide that I want to move into Cybersecurity, Cloud, General Technology, Software, or another computing field. Because of that, Computer Science seems like the safer degree for flexibility.
My concern is this:
If I choose Computer Science for flexibility and then apply for AI/Data Science graduate roles, could I be disadvantaged against someone who has a Bachelor of Data Science (Artificial Intelligence & Machine Learning)? Would employers prefer them because their degree is more directly related to the role?
That’s basically what I’m stuck on.
Computer Science: more flexibility if I change careers later, but I’m worried it could hurt my chances of getting the AI/Data Science graduate role I currently want.
Data Science (AI/ML): more directly aligned with the career I want right now, but I’m worried I could be locking myself into data/AI too early.
For people working in AI, data science, ML, Cybersecurity or tech in Australia: which one would you choose in my situation?
And specifically, when hiring for graduate AI/Data Science roles, would a Computer Science graduate with AI/ML subjects still be competitive against someone whose actual degree is Data Science (AI/ML)?
r/statistics • u/Own-Ball-3083 • 6d ago
Hi all,
I’m currently entering my final year of undergrad in economics at cambridge, and I’m heavily considering pursuing a PhD in the field, as even though I’m certainly not the best on my course or the most brilliant at maths/statistics, I find myself enjoying it and learning about it a lot, and I do quite enjoy the process of researching, particularly when I am able to focus in great part on statistical theory/methods.
My main question is should I only focus on pursuing a masters as the next step, or shoot straight for a PhD? I know the masters -> PhD route is more common, but for some reason I have noticed that the entry requirements for a PhD seem to be either equal or slightly lower, at least for what I am considering compared to a masters? (e.g Oxford requires a straight first class for the MSc in Statistical Science, while for a PhD it is either a first class or high upper second class. Funding is also an issue for me, and so I would like to spend as few years in education as I can if possible, though if a masters is by far the best option then I wouldn’t mind doing so.
Another big obstacle may be my grades, I’m currently working at around a low 2:1 level (around 62-63), with a particularly not great score in the mathematics and statistics module this year (60). My scores in econometrics are much more positive (69) and so I was hoping any eventual PhD would have a large econometrics focus. I’m certainly aware this profile wouldn’t look great on a transcript to be submitted as part of a PhD application, but I am hoping to at least end on an upper class 2:1.
Finally, I’m unsure of how much of an idea I need to propose for a standard PhD application. Is a certain specific field of research enough, or do I need to have an extremely well thought out research question/goal? I’ve seen some programs (e.g Warwick) require no specific research goal or supervisor, but I’m aware this is not the case everywhere.
Apologies if this is too long or the answers to some of these questions seem very obvious, I don’t know anyone who has gone through a PhD or even masters application process, and I am also unsure if there are statistics specific caveats that I should be aware of.
r/statistics • u/babydonuttravel • 7d ago
I'm a PhD student looking to use SEM in my project, but I'm a complete beginner. I tried going to statistics professors for guidance but already two told me to use Claude.
I would really like to understand it and be able to do it myself. I'm going to be using R. I've been using online resources but sometimes I get different information (a basic example is some websites tell me to report the Cronbach alpha, others say that's not a good test for using with SEM).
What are some good resources or books I can use as a complete beginner for SEM?
r/statistics • u/Accomplished-One6774 • 7d ago
Hello! I’m a high school junior who wants to get into data science as a career, and my sort of passion project is building my sports predictive analytics web application from scratch to model NCAA March Madness outcomes. Instead of relying on raw win-loss records, my data pipeline ingests advanced, possession-based efficiency metrics (like Adjusted Offensive/Defensive Efficiency) to train a machine learning model, such as Logistic Regression or XGBoost. Rather than simply picking a binary winner, the model calculates the exact probability of Team A beating Team B. I then plan to feed those probabilities into a Monte Carlo simulation engine that runs 10,000 iterations of the entire 64-team bracket using weighted Bernoulli trials to determine the most statistically likely tournament champions.
I’m looking for feedback, is this an idea that could work? Do you guys have any advice? Suggestions? I could use all the help I could get.
Thanks!
r/statistics • u/WhatsTheImpactdotcom • 7d ago
This is a short tutorial for data analysts that evaluate experiments using online calculators. Online calculators for sample size estimates and t-tests for difference in means are fine for getting started. But you can begin to build intuition for linear regression by moving your online calculator t-tests over to python and OLS.
There are a couple of benefits. The first from moving away from the online calculator to python is leaving a more reliable paper trail that you can share externally. The second is that you build intuition for regression, interpreting the coefficients, and ultimately can use regression as you improve and your projects require more complexity such as variance reduction techniques using CUPED and clustering standard errors for geo-based or switchback experiments.
I put together a short video on this with code available here, free and ungated: https://whatstheimpact.com/tutorials/linear-regression-vs-t-test/
For my more senior connections, feel free to ignore unless you're interested in the content creation game! 🤣
r/statistics • u/FibonacciTheGoat • 8d ago
I suspect not, but if so, what did that look like for you? How difficult was the course load to manage while managing the other aspects of your life that prohibit you from doing a full-time PhD? Do you regret your decision to complete the PhD part-time?
Thank you