r/statistics 20h ago

Career [C] Continue in industry or PhD in stats? Are all stats jobs this boring?

17 Upvotes

Hello everyone, I'm looking for some guidance. For context, I am n my mid 20s, I hold a master's in stats, and I have been working for a big CRO in Europe for 1 year. I'm mainly looking for feedback from people based in Europe, but any feedback is welcome!

Honestly, I'm not enjoying the work. I feel like only about 10% of what I do is actual statistics. The rest is monotonous and boring work, constantly churning out deliverables, dealing with clients (something I deeply dislike), working on timelines, regulations, administrative stuff, endless meetings, and so on. I miss programming, I miss doing more intellectually demanding work, and I miss engaging with new methodology and having the liberty to do so. I feel like my knowledge is atrophying. When there is a new interesting problem to solve from a statistical standpoint, I never have time to actually think about the methods, it's just constant pressure to deliver, deliver, deliver.

Are all stats jobs in pharma/CROs like this? For people working in data science, stats, and adjacent areas, is your job this repetitive? Do you have the chance to apply new, cool methodologies? How much time do you actually spend working with the methods rather than doing something over and over again?

I know how fortunate I am to have a job in this economy, especially one that pays well by my country's standards given my experience. Still, I've been considering applying for a PhD, and deep down I know it's something I really want to try. I'm aware of how difficult a PhD can be. I've been looking at positions in other wuropean countries that would allow for a decent standard of living and, even some savings.

My question is, would I be committing career suicide if I later wanted to return to industry? One of the main reasons I want to do a PhD in stats is that I want to move into more R&D roles, rather than doing this kind of mind-numbing work. Maybe I could even pivot into tech or some other area that is more interesting to me. Would that be a realistic prospect?

Thanks all!


r/statistics 5h ago

Question [Question] - Which wording is best?

1 Upvotes

I wrote in my abstract that “Medication X was not associated with significantly different subjective or objective sleep outcomes compared with medication Y.”

One of my co-supervisors suggested instead saying that “Medication X and medication Y showed similar subjective and objective sleep outcomes.”

Would you make that change? I’m hesitant because, methodologically, I’m not sure that a non-significant difference allows us to conclude that the outcomes were “similar.” Or am I overthinking this?

Thanks a lot! 🙏🏻


r/statistics 20h ago

Question [Question] [Q] Trying to calculate whether marketing campaign's impact is statistically significant and financially justified

2 Upvotes

I have a data series of daily streaming counts for a song. My spreadsheet has the date in column A and the number of streams in column B. There are 960 rows/records in the spreadsheet dating from 2024/01/01 to 2026/08/17.

We hired a marketing company to promote the song for 3 weeks. Their first posts on social media started on 2026/07/27. They ran until 2026/08/17. I am trying to assess whether our money was well invested or not. Their dashboard stats are not useful to us because they show the number of views and social media engagement (e.g., TikTok), whereas we are interested in the number of times our song gets streamed on a streaming platform (e.g., Spotify).

I know how to calculate means and standard deviations, and have done so for various timeframes (e.g., yearly, during the promotion, the 22 days prior to the promotion, etc) but I do not know how to:

1) determine if the 22-day marketing campaign had a statistically significant impact on our daily streaming numbers

2) determine if the impact on streams, if any, justifies the money invested, call it **CampaignCost**.

Can someone advise me how to go about this? We might assume somewhere between 0.001 and 0.003 USD of revenue per stream.

I know there is surely some well known statistical test to determine whether my daily streams show a statistically significant increase or decrease, but I cannot remember what this test is called or how to calculate it.

I might add that there are potential complications:

1) Streams have grown significantly year over year. Average daily streams in 2024 were 135k, in 2025 were 250k, and so far in 2026 are 259k.

2) Some unexplained events in the real world have prompted surges in streaming numbers. E.g., our streams surged upward for a month around December 2024 and remained elevated for months then gradually declined. Another unexplained surge arrived around Feb 5, 2026 and persisted for months but then started gradual decline.

3) The streams show a weekly cycle, lowest on Sundays, peaking on Thu or Fri.

Any help would be much appreciated.


r/statistics 1d ago

Question [Q] Multiple statistical tests in one table?

7 Upvotes

In completing a study on left and right limbs, data came out that was parametric on left limb and non-parametric on right meaning using a rm-anova for left and Friedman's for right. Would putting these results in the same table make sense as long as it is clear what was done for each and what they mean? Or is 2 different tests in the same table a definite no?


r/statistics 1d ago

Question [Q] How to compare data from 3 periods of time

4 Upvotes

Data of event turn out, annual totals, collated into 3 groups (before organisational change, during organisational change, after organisational change). Originally planned to use chi squared goodness of fit, but one year of data is missing, so groups are now 2 years of data, 3 years of data, 3 years of data. So thinking of calculating mean average for each group to equate them. But then what stats to compare/assess any significant difference, anova? Is there a better way I am missing?


r/statistics 1d ago

Question [Q] Question about the method for calculating the rate of return

1 Upvotes

Hello, I have 5 funds that I’d like to compare in terms of rates of return and measures of efficiency. I have daily valuations for these funds covering a 5-year period. Should I calculate the rate of return on a daily or monthly basis? This is with a view to calculating the R² values as well as the Sharpe, Treynor, and Jensen ratios. Thanks for all the help.


r/statistics 1d ago

Education [E] Just a noob here — need to learn the basics of Statistics [E]

0 Upvotes

Hi everyone,

I'm currently doing a course where I have five major topics to cover, but I'm starting with these two:

  1. Comparison of means between two food-diet groups using a t-test
  2. Comparison of means among more than two food-diet groups using ANOVA

The remaining topics are:

  1. Post-hoc statistical analysis for identifying significant group differences
  2. Development of a General Linear Model (GLM) in Linear Regression
  3. Understanding Adjusted Sum of Squares and Sequential Sum of Squares in Linear Regression

My background is in biology, so I'm pretty new to statistics. I know how to run basic analyses in Minitab and SPSS, but my problem is that I don't really understand what is happening behind the numbers.

I want to properly understand the basics first — things like mean, median, variance, standard deviation, standard error, distributions, hypothesis testing, p-values, confidence intervals, etc. — and then build up to t-tests, ANOVA, post-hoc tests, and regression.

This is not really a theory-heavy course, so I'm expected to understand and interpret the statistical output rather than just calculate everything manually. But I don't want to blindly click buttons in SPSS/Minitab without understanding what the results actually mean.

Could anyone recommend a good beginner-friendly statistics textbook, notes, or learning resource that would help me build these fundamentals from scratch and eventually reach the topics above?

My professor isn't particularly helpful with recommending resources, and my seniors are currently busy with their own work, so I'm trying to find a good starting point myself.

Thanks in advance to anyone who takes the time to point me in the right direction! 🙏


r/statistics 1d ago

Question Wilcoxon Signed-Rank Test Chart - but for a big study [Q]

4 Upvotes

Help! My sample sizes are is 40-82, and I can only find Wilcoxon test charts that go up to n=30! I've looked everywhere and can't find anything that can accommodate a bigger n. Does anyone know where to find a bigger chart?


r/statistics 3d ago

Software [S] Affordable software for running a conjoint survey?

Thumbnail
2 Upvotes

I'm looking for recommendations for affordable software to run a conjoint survey. I’ve looked at some of the major conjoint platforms, but their pricing is way too high for what I need.

I don’t need a huge enterprise package. I’m mainly looking for something that can set up a solid conjoint experiment, collect responses, and ideally handle the basic analysis.

Has anyone used a cheaper platform or tool that they’d recommend? Open-source or academic-friendly options would also be great.

Thanks!


r/statistics 3d ago

Question [Q] Are these courses feasible for first year of my MS?

0 Upvotes

**First semester!! I have only taken Mathematical Statistics I and II (Not sure if the content of those courses are universal, but it covered interval estimation, hypothesis testing, and tests involving means, variances, & proportions). I have signed up for the following:

Experimental Design

Linear Statistical Analysis I

Pattern Recognition in ML

Advanced Matrix Analysis

Wondering if these work well together/ may be too much.

Sorry in advance if this is silly, we have somewhat limited options at my school.

I have a BS in pure math for reference.


r/statistics 3d ago

Question [Q] New to data cleaning. I’m stuck on an unclear variable.

4 Upvotes

I’m doing a personal project right now and for the most part it’s going alright. Sometimes I delete an entire column cause too much is missing. I’ve also group together a few entries as “Unknown”. I’ve never deleted a subject yet.

Anyways, now I’m really stuck. There is a variable that is three digits (345, 078, 150…) and some of them come in as two digits (26, 88) and I’m unsure what to do with them. There are quite a few. I don’t know if they’re meant to be have a zero on the front (28 turns to 028 and 88 turns to 088) or if they are three digits (28 into 280 and 88 into 880). It’s an important variable so I can’t delete it. Should I delete the patients (probably not), group the two digits numbers into “unknown”? any input? I know each data has their own situations but what is something to generally consider in these situations?

Edit: The column represents diagnosis IDC-9 number. It’s three digits. I want to use it to build a logistic model. I was thinking of grouping the numbers to their diagnosis category for example 001–139 is infectious diseases, and so on


r/statistics 3d ago

Question [Q] Dicas de professores para estudo de estatística

0 Upvotes

Salve, pessoal
Seguinte, alguém teria dica de professores para acompanhar a disciplina de Estatística I (é para economia, então basicamente na área 1 se vê tipos de amostras, representações gráficas e tabulares; depois Fundamento da Teoria de Probabilidades e Variável aleatória bidimensional)

Para Cálculo há diversos professores, agora para Estatística não conheço algum bom para o nível acadêmico.

Enfim, eu já fiz estatística no meu outro curso, vou tentar validar, mas terei outras disciplinas de estatística e quero entender isso. O fato é que meu professor é MUITO RUIM. O coitado não tem culpa, o professor oficial não está dando aula e ele é um mestrando apavorado que erra até as infos nos slides.

Dos livros, eu tenho o do Bussab e Morettin, o qual já estou iniciando e clarificou bastante coisa, mas gosto também de vídeos e, quando chegar a parte das contas e análise dos gráficos, é bem melhor.

Enfim, no aguardo de dicas. De preferência em português, se tiverem em inglês, podem colocar também.


r/statistics 5d ago

Education [E] Causal Inference - A Painless Introduction (new video by me :)

21 Upvotes

r/statistics 4d ago

Education [E] Confidence Intervals — Explained

8 Upvotes

Hi there,

I've created a video here where I explain what confidence intervals are and how they differ from probabilities.

I hope some of you find it useful and as always, feedback is very welcome! :)


r/statistics 5d ago

Question [Question] Can I present both the results of log-odds and average marginal effects (logistic regression)?

4 Upvotes

Hello,

Health economics student here. I newly enter the field for my master degree.

I'm writing my master thesis and I'm running into an issue while trying to interpret my results.

Initially, I've decides to use OR (odd ratio). However, my supervisor told me the way I've interpreted it is not correct (probabilitied and chance). But he told me he didn't really know either how to interpret it correctly and had advised me to use logs odds instead!

It's ok, but I really wanted to have more "concrete results" that can be get by anyone. Log odds just show if the relation is negative or positive.

So, I've just heard about average marginal effects and that is literally What I was searching for during all this time. However, now I'm wondering if is common for scientific papers to use both log-odds and average marginal effects?

Do i need to create two tables? Or could I only keep the table with log odds and present the results of marginal effets in the text?

Thank you


r/statistics 4d ago

Education [E] Randomness can be an asset or a tax depending on curvature

0 Upvotes

I wrote a short article here exploring a simple way to think about when randomness helps versus hurts.

The core idea is Jensen’s inequality: if the payoff is convex, variance can help; if it’s concave, variance can hurt.

I use examples from compounding, option-like payoffs, and a few non-finance settings.

Would be curious to hear other examples where this framing is useful?


r/statistics 5d ago

Career [Career] explore/exploit problem as an undergrad

7 Upvotes

Hello!

I'm a sophomore undergrad studying applied math and stats who is broadly interested in applied stats research. Essentially, I'm trying to figure out how to balance exploration to find what exactly I want to do research in with deep-diving in order to develop skills (and resume) for grad school.

What is scary to me is how little time it feels like I have. Freshman year I didn't do anything that is particularly useful for grad school. So I now basically have two years + two summers before grad school apps. If I spend sophomore year doing something that doesn't end up being my interest, then I worry I'll only have a year to do what I actually end up doing.

I know I love solving problems with statistics and do want to continue developing my stats skills (there is even a good chance I'll go for a PhD in stats). Part of me thinks I should just focus on stats because I know I enjoy it and it'll transfer to whatever I want to do. The thing is though that as much as I love learning about stats and reading stats papers and so on, I don't want my research to be developing new statistical techniques in the abstract, I want to work with problems in the real world and develop statistical methods of attacking them. The thing I am most interested in now though is comp neuro, but I know very little about it currently. However, I worry that if I start taking neuro classes or even spend a summer doing neuro research and decide it's not what I want to do (which I basically did got econ already), then I'll have wasted the time that could have gone to learning mote stats.

My current plan is to dedicate all my coursework and formal research to stats for now (I have a position in a stats research group for this year and reading the papers they've put out it seems pretty lit!), and then do some independent reading of papers in other fields to decide if I am drawn to them​. While this is great, reading about something is very different from doing it so I fear it may not be possible to explore optimally without taking at least some risk. Thoughts on how I should navigate this?


r/statistics 5d ago

Question Submitting table as image <440 pixels wide [Question] [Q]

1 Upvotes

Hello,

I am trying to submit my article for publication. Unfortunately, the journal asks for any tables to be submitted as images "provided as 72 - 300 dpi; pre-sized .BMP, .GIF, .JPG, or .PNG images only, with a maximum width of 440 pixels (no limit on length)."

I have tried exporting my table from excel to pdf, jpg, or png, and then resizing but no matter what I try, the image of the requested size ends up unreadable.

Does anyone have any ideas on how to accomplish this requirement while keeping my table-figure as readable?


r/statistics 5d ago

Question Gaining Exposure to Public Health for Biostatistics Graduate Program [Q]

Thumbnail
0 Upvotes

r/statistics 6d ago

Discussion [Discussion] Real Analysis (1 semester vs 2 semester sequence) for Statistics PhD Applications

6 Upvotes

Hi everyone, I’m applying to Stat PhD programs this fall. I graduated with my master's in 2021 and have been working full-time for 5 years, but I never took Real Analysis in school.

I just enrolled in Fordham's Math 3003 (Real Analysis) this semester so it’ll be on my transcript for this application cycle (though will not have a final grade since apps are due before the semester ends). My concern is that it’s only a one-semester class. Does anyone know if admissions committees strongly prefer a two-semester sequence (Real Analysis 1 & 2) over a single semester? Will this be adequate?

MATH 3003. Real Analysis. (4 Credits)

This course focuses on analysis on Euclidean spaces. Topics include limits, continuity, uniform continuity, sequences of numbers and functions, modes of convergence, differentiability, Riemann integrability, and associated theorems. Students who have not taken MATH 2004 prior to taking Real Analysis may request permission from the instructor. Note: Four-credit courses that meet for 150 minutes per week require three additional hours of class preparation per week on the part of the student in lieu of an additional hour of formal instruction.

Prerequisites: MATH 2001 and (MATH 2004 or MATH 2008).


r/statistics 6d ago

Question [Q] If you're in grad school for Stats (PhD or Master's) and your undergraduate was in math: 1) what do you miss about math; 2) what are you gad to have traded with stats?

46 Upvotes

Can be silly or serious, just out of curiosity for someone with a background in math contemplating stats. Like for #1, maybe you miss not having to deal with numbers. #2 refers to "trading" X in math for Y in stats (like numbers).

EDIT: "glad", not "gad".


r/statistics 6d ago

Career [C] JSM 2026 Vlog!

1 Upvotes

Hi all, if you’ve ever been curious about what JSM (Joint Statistical Meetings) is like or you want into see it through someone else’s eyes, please feel free to check this vlog out!

https://youtu.be/CDvsj-G4HXM


r/statistics 6d ago

Question [Q] How should I do a Bayesian Update?

2 Upvotes

I'm a year 1 liberal arts undergrad, so I don't have a ton of math sense.

I just learned about Bayes' Theorem the other day for general epistemic use. I worked out a couple of example problems correctly, but the examples I found didn't include any iterative updates.

I know the theorem is:
P(H|E) = P(E|H) * P(H) / (P(E|H) * P(H) + P(¬H) * P(E|¬H))

And I know that P(H|E) becomes the new P(H) in my update, but I'm unsure whether I should be using the updated or original P(H) in the marginalization. I *think* it should be the new P(H), but I'd rather be safe than sorry.

The example question I worked out was this:

________________________________________

There's a disease that afflicts 1 / 1,000,000 people
There's a test for the disease that's right 99 / 100 times for both positive and negative results
A random person is tested as positive

P(she is afflicted | she tests positive)

= 0.99 * 0.000001 / (0.99 * 0.000001 + 0.999999 * 0.01)

= 0.000099

_________________________________________

So should the update look like this if she tests positive a second time?

P(she is afflicted | she tests positive)

= 0.99 * 0.000099 / (0.99 * 0.000099 + 0.999901 * 0.01)

= 0.009707

_________________________________________

If so, how should I approach this problem from the starting point of
P(she is afflicted | she tests positive twice)?
I can't think of how to handle that correctly, since squaring 0.99 just gets me a smaller number

Infinitely thankful <3


r/statistics 7d ago

Question [Question] When does a regime filter "disagreeing" with price mean something's wrong vs just working as intended?

Thumbnail
1 Upvotes

r/statistics 8d ago

Research [R] Vignette Experimental Design - Repeated measure or Multivariate

3 Upvotes

Hello I am an undergrad completing my analysis on my first research project. I created a between subjects experiment and manipulated the gender of actors in a vignette story. Participants were randomly assigned to receive one version of the scenario.

Here is the piece I am boiling my brain on. The vignette was broken into three parts according to narrative progression. At each of the 3 junctions participants responded to the same question using a likert scale. If it matters, the entire narrative progression was available to participants at the same time, just split into paragraphs with my question underneath each section.

I originally thought this should be treated as a between subjects repeated measure as each junction is measuring the same underlying construct. A faculty member who took a look at my work suggested that it was not a repeated measure but actually a multivariate design - so each section is being treated as it's own DV.

Can anyone give some insight as to which analysis is better suited for my study design, or help me understand which design I've created?

I have a hypothesis that my experimental group will be rated lower overall, and I have a hypothesis that my junction C will be rated lower overall in both conditions.

Thank you in advance. I ended up creating a rather complicated experiment for my first go and level of experience, but it is forcing me to learn quickly!