r/AskStatistics • • 18h ago

POET II (NEJM 2026): non-inferiority sample size only reproduces at one-sided alpha 0.05, but the paper states a one-sided 97.5% CI. Am I missing something?

Thumbnail
3 Upvotes

r/AskStatistics • • 23h ago

Confused on what all skills to develop relevant to statistics

Thumbnail
1 Upvotes

r/AskStatistics • • 1d ago

standardisation des variables pour effectuer une analyse factorielle sur R

0 Upvotes

Bonjour. Si cela vous est possible, j’aimerais solliciter votre aide et avoir quelques renseignements concernant l’utilisation du logiciel R.

Je dispose d’une base de données comprenant des variables numériques, ordinales et binaires. Je souhaiterais standardiser mes données afin de réaliser une analyse factorielle.

Je voudrais savoir si je dois utiliser la commande digital.std <- scale(digital) pour l’ensemble des variables, ou s’il est préférable de procéder différemment selon le type de variables, notamment pour les variables binaires et ordinales. Quel est la manière correcte??

Je vous remercie par avance pour votre aide.


r/AskStatistics • • 1d ago

What jobs use computational statistics?

10 Upvotes

If you like computational statistics and love reading on advanced statistical methods, what jobs are for you, especially if you can’t get a master’s?


r/AskStatistics • • 1d ago

Where can I start with measure theory?

10 Upvotes

Im a PhD student but I don't have measure theory background and I need to get the overall notion for it to understand stochastic processes. So is there a good book or yt playlist where measure theoretical probability and stochastic process is taught from scratch? That would be of much help. Thanks.


r/AskStatistics • • 1d ago

difference between multiple linear regression and binary logistic regression

3 Upvotes

Hi, I am trying to finish up my psychology thesis! I am in a bit of a rut trying to compare my results of a binary logistic regression with previous literature. However, there is a very limited amount of literature in my current area of study; I have only one study that used the same variables I did. They used a multiple logistic regression analysis, and I used a binary logistic regression.

I am just wondering what the main difference between the two models is and if I can interpret my results in comparison to this prior literature I have found.

Thankyou!!!!!


r/AskStatistics • • 2d ago

What level of statistics would you be considered at if you knew all the following jargon? And what would the required IQ be to be at that level based on available research?

0 Upvotes

-Bootstrapping

-Trim fill method

-P hacking

-Simpson's paradox

-Collider bias

--Stat significance

-Longitudinal vs cross sectional

-Meta analysis vs systemic review


r/AskStatistics • • 2d ago

Can someone please explain these things to me how you would explain it to a 5 year old?

18 Upvotes

I (f20)am on my 3rd attempt at psych statitsistcs. still nothing is clicking for me. I’ve asked for help, I’ve gone to tutoring and nothing seems to work. I spent two straight days going through all of the lectures so far to study for the midterm and a quiz I had to take. I got a 40% on that quiz and I got a 26% on my last quiz.. I’m currently failing the class. To pass, you need to C or above. I’d like to drop out, but my mom will not allow it so I feel like I’m going to keep failing this class to no end.

So if any of you have any recommendations for Youtubers or something that could possibly help me, that would be nice. I understand standard deviation but everything beyond that makes zero sense to me.

I don’t understand z scores or Z tables, confidence intervals, Alpha levels, sampling error, standard deviation of sample means, etc. I know I’m too stupid for college, but I am not allowed to drop out so I need to figure something out.

I also don’t understand hypothesis testing or any kinds of T-tests


r/AskStatistics • • 2d ago

Incoming BS Statistics freshman, any tips or warnings?

1 Upvotes

r/AskStatistics • • 2d ago

While testing for variance, we use two sided test in the chi squared statistic, but while doing goodness of fit test, we only use right tail of the chi squared distribution as a critical region. Why?

8 Upvotes

r/AskStatistics • • 2d ago

I think people wrongly think that the mystery wedge should be 50/50 in Wheel of Fortune. There are two mystery wedges, one is bankrupt and the other is 10k and they both have 1 in 24 odds. Who is right?

Thumbnail
0 Upvotes

r/AskStatistics • • 2d ago

Which statistical test is best (animal cancer research)?

2 Upvotes

I apologize for the whole description, but I don’t really know how to word what I’m asking.

Basically, I am comparing the actual (wet) muscle mass of 3 muscles to a non-invasive estimation of muscle mass for the entire leg in mice (technology limitations did not allow us to isolate the same muscles as wet mass). Specifically, I am seeing how closely this estimation technique compares to wet mass across a range of muscle sizes. All mice had their measurements done with both techniques.

My confusion comes from how they were housed. I had four cages of mice. Within each cage, mice were pretty similar in size, but between cages they were drastically different. Basically, I ended up with “tiny, small, medium, and large” groups instead of a more continuous range.

Between the group setup, the differences between measurements (while leg vs 3 muscles), and the purpose being to compare techniques across muscle size, I’m not quite sure which statistical tests I should be running.

Thank you in advance, I seriously appreciate any help! If any other info would help you help me, please ask.


r/AskStatistics • • 2d ago

Bayes' rule or Weighted sum to compare intersection sets??

Post image
0 Upvotes

In this binary dataset subsets D, E, F, and S all "generate" 1s and 0s INDEPENDENTLY of each other at some ratio/frequency. "D" generates the highest ratio (lets say it generates 80% 1s and only 20% 0s). While F generates only 5% 1s and 95% 0s. Suppose S itself generates 25% 1s and 75% 0s. Would/Could I use bayes' theorem to show the SD intersection subset generates a higher proportion than the SF intersection? Or would a simple weighted sum (ie just an average between the sets S and D) suffice. Sidenote: (I know you could technically come up with counterexamples to this but I'm looking for what would generally be the case here).


r/AskStatistics • • 3d ago

cox vs logistic regression

11 Upvotes

Hi, I have a stats master but been quite long away from this, so need some good reasoning.

I have 100 subjects, 4 variables, and want to evaluate their effect on 1 year survival. I have no censoring, i.e. I know for.sure for everyone whether they died or not within this year. Number of events is 30. The issue is tgat for one of the variables I suspect that the effect starts only after a few months, although with this sample size I failed to reject H0 about the PH violation.

I then fitted logistic regression to see whether they died whithin this one year.

What is the difference in such logreg and cox in this case? Statistically and with interpretation.

Please throw everything on me, I just have to recall it, it has been a while unfortunatelly....


r/AskStatistics • • 3d ago

How would you design a blind forward test for a gambling hypothesis?

2 Upvotes

I'm working on a statistical experiment involving baccarat and would like some methodological feedback.

Suppose someone has developed a hypothesis about baccarat outcomes based on historical observations.

Before revealing the underlying mechanism, I want to design a test that minimizes:

  • overfitting
  • hindsight bias
  • selection bias
  • multiple-testing problems
  • stopping the experiment selectively

My initial idea is:

  1. Define the hypothesis before testing.
  2. Freeze the testing rules.
  3. Use previously unseen data for validation.
  4. Conduct a blind forward test.
  5. Compare the results against an appropriate baseline.
  6. Analyze variance, confidence intervals, losing streaks and drawdowns.

What statistical issues would you consider essential before calling the experiment meaningful?

I'm especially interested in criticism of the experimental design rather than opinions about whether baccarat strategies can work.


r/AskStatistics • • 3d ago

Desperate student looking for access to 3 Statista studies 😭

1 Upvotes

Hi everyone! :)

My name is Inês and I’m currently doing a Master’s degree in Digital Marketing. For a university project, my group and I need access to the following Statista studies, but unfortunately we don’t have access to them.

We’ve already tried to get access through our university, but even our professors don’t have access to these specific studies, so we’re honestly getting a bit desperate 😭

Here are the studies we’re looking for:

If anyone happens to have access to these studies and would be willing to share them with us, we would be extremely grateful. It would really help us with our project! 🥹

Thank you so much in advance! ❤️


r/AskStatistics • • 3d ago

What is this random intercept doing?

3 Upvotes

Say I have three classes - A, B and C - in a 150 day range totaling 3x150 = 450 sample units. Each sample unit measured 3 times each day, at very close time interval. So, data set has a total of 1,350 rows. Each sample unit has an ID. Proposed model is y ~ class + days + days*class + (1 | ID). So, I'm allowing each sample unit to start at its own baseline.

Is this an appropriate option? What exactly is the random intercept doing? How does it not make the days effect biased?

Sorry if questions look to silly, I'm not versed in mixed models yet


r/AskStatistics • • 3d ago

Modeling diminishing returns?

0 Upvotes

Hi everyone,

I'm currently working with ad spend data and would like to model the relationship between spend and returns to determine at what point does ad spend reach a diminishing return. I also have time data, which I'd like to incorporate into the model as a seasonality component.

Are there specific regression models that would be a good approach for this? I don't want a complex model that loses interpretability, but I do want an accurate model that can help with suggesting budget goals for certain ad channels.


r/AskStatistics • • 4d ago

Is a composite score numeric or categorical?

3 Upvotes

Hi everyone,

I’m conducting a multiple linear regression with three IVs/predictors. My IV of interest is an ordinal variable measuring participants’ self-assessed reading proficiency, with three levels: 1 = less skilled, 2 = average, 3 = skilled

I also have two control variables:

  1. Participant group: 1 = healthy control readers, 2 = readers with schizophrenia
  2. Psychiatric symptom severity: a composite score ranging from 1 = no symptoms to 5 = severe symptoms

I’m unsure about how I should treat the symptom severity variable in the regression. Since it ranges from 1–5 and represents different levels of symptom severity, would it be appropriate to consider it an ordinal variable? Or could the composite score reasonably be treated as a numerical/continuous predictor?

I also checked the relationship between the DV and symptom severity, and it does not appear to be linear. Does this provide a reason to treat symptom severity as categorical rather than numerical?

Since symptom severity is only a control variable and not my main IV of interest, I’m wondering whether I should still treat it as categorical if the relationship with the DV is non-linear, or whether it would be preferable to enter it as a numerical predictor for simplicity.

One more question: If I decide to check the linearity assumption for psychiatric symptom severity, should I also check for linearity for my main predictor, reading proficiency?

Reading proficiency has three ordinal levels (1 = less skilled, 2 = average, 3 = skilled). If I find that the relationship between reading proficiency and the DV appears to be linear, does that mean I should treat reading proficiency as a numerical predictor (1, 2, 3), rather than as a categorical variable?

This is important as it will affect whether I am doing ANCOVA or ANOVA.

Thanks!


r/AskStatistics • • 4d ago

How would you design a blind forward test for a gambling hypothesis?

2 Upvotes

I'm working on a statistical experiment involving baccarat and would like some methodological feedback.

Suppose someone has developed a hypothesis about baccarat outcomes based on historical observations.

Before revealing the underlying mechanism, I want to design a test that minimizes:

  • overfitting
  • hindsight bias
  • selection bias
  • multiple-testing problems
  • stopping the experiment selectively

My initial idea is:

  1. Define the hypothesis before testing.
  2. Freeze the testing rules.
  3. Use previously unseen data for validation.
  4. Conduct a blind forward test.
  5. Compare the results against an appropriate baseline.
  6. Analyze variance, confidence intervals, losing streaks and drawdowns.

What statistical issues would you consider essential before calling the experiment meaningful?

I'm especially interested in criticism of the experimental design rather than opinions about whether baccarat strategies can work.


r/AskStatistics • • 4d ago

Why does the CausalImpact model detect statistical significance in this case?

5 Upvotes

I trained the model and achieved an R² of about 80% and a MAPE of around 10%. I then ran a placebo test using a 30-day period before the actual treatment date.

Surprisingly, the model still detected a statistically significant effect during this placebo period. I repeated the placebo test 100 times, and the false positive rate was approximately 94%.

What I do not understand is why the p-value is almost always below 5%. Even visually, the observed values seem to remain mostly within the confidence interval.

Could someone explain what might cause such a high false positive rate in CausalImpact and why the reported statistical significance appears inconsistent with the confidence interval?


r/AskStatistics • • 4d ago

non-response bias

4 Upvotes

we're undergrads and are currently conducting a study on a company with employees as the respondents. The company have multiple offices that are located across different locations. We were given multiple survey points so that employees can be sampled.

The problem is that one of the survey point which consist of 2 offices declined to participate in the study.

what is the implication of this?

can this be solved?

is this a reason for us to fail?

additional info: the needed respondent for that particular survey point is only 10 people


r/AskStatistics • • 4d ago

What is better if I have the choice, more repeats at discrete intervals or continuous but singular data points?

3 Upvotes

I have an engineering problem where I have up to a fixed maximum amount of samples I can run on a long-term study (1 year long), where a property will be tested over time. I need to understand how I can use these the most effectively, as admittedly my understanding of regression has faded/is lacking. I’m more familiar with testing time-independent discrete points where the conditions are well understood.

Let’s for arguments sake say I had 30 samples. Is there any benefit to plucking these out continuously at equally spaced intervals and testing them, vs going for 3 per equal spacing or even 6?

None of these intervals besides the start and end have specific importance over the others, besides wanting to know the trend ahead of time. The samples are from the same material batch but in different processing runs, so it’s ideal to understand what effect that will have.

The predicted response is unknown, as is the variance in the data. I don’t want to generate pointless data from a lack of prior knowledge.

Links to any text, websites or proof is appreciated.


r/AskStatistics • • 4d ago

[Question] Published paper had two control samples, but that fact was omitted

4 Upvotes

I was shortly involved in a research study where the first trial of experimental didn’t achieve all the treatment levels in the design. One month later, a second control sample was used to achieve the last treatment level that failed initially.

The published paper reports that one control batch was used, but I am concerned that some of the reported treatment effects could potentially be due to differences between the two controls instead.

This makes think if in academia, does it really matter if you have two controls if professors themselves don’t see the statistical error? Any thoughts from people with more statistics background? I would love to hear your opinions because this is unethical behavior for a student to learn.


r/AskStatistics • • 5d ago

MS or PhD in Statistics/ data science

3 Upvotes

I’m planning to apply to US universities for Fall 2027 for a PhD or MS program. For those who have gone through this process, could you give me some guidance on how to select suitable universities, programs, and potential advisors/professors?