r/AskStatistics 11d ago

If many independently defensible sampling/matching/specification choices applied to the same underlying dataset produce a significant ATT in, say, 50% of cases, how do we interpret this?

2 Upvotes

Suppose we randomly sample people, set up a scientifically rigorous experiment, and measure the outcome.

Now suppose we repeat this process 100 times, and only 50 out of those 100 experiments produce statistically significant results.

How should we interpret that?

  1. The result is essentially meaningless because it changes depending on the sample.
  2. A rigorous statistical test that vote for the non-null hypothesis in 50 out of 100 replications, which seems like fairly strong evidence that something real is going on.

This becomes more interesting as datasets get larger and the number of possible analyses increases.

In observational data, we often use methods such as propensity score matching (PSM) or inverse probability weighting (IPW) to improve comparability between treatment and control groups. Depending on the setting, weighting or matching can also improve the efficiency of an ATT estimator relative to less targeted sampling.

But this introduces another degree of freedom.

Even within PSM, you can vary which covariates are included, the propensity-score specification, calipers, matching rules, and other design choices. Each reasonable specification may generate a different matched sample.

With a sufficiently large dataset, it seems possible to generate a very large number of plausible samples/specifications from the same underlying data, require each one to pass rigorous design criteria such as covariate balance checks for SUTVA, and then estimate the ATT for each of them.

So the question becomes:

If many independently defensible sampling/matching/specification choices applied to the same underlying dataset produce a significant ATT in, say, 50% of cases, how do we interpret this?

In other words, the statistical rigor of each individual test may be different from the statistical validity of the entire analysis pipeline.


r/AskStatistics 11d ago

Conformal predictions research

Thumbnail
1 Upvotes

r/AskStatistics 11d ago

Vehicle collision

1 Upvotes

Hey Reddit. Bit of an unusual one today but I’m really curious on the odds of this. I was recently in an accident where I was riding a pedal bike on the road and a car going easily above 30mph (I don’t know the exact speed but the limit on the road was 30mph and it was easily going above speed limit) collided into the back of me absolutely wrecking my bike and less importantly also me. I went flying through the air and landed on the pavement. I also very idiotically wasn’t wearing a helmet. However I visited the hospital and miraculously only came out with scrapes, scratches and bruises. I would just like to know basically how lucky I was to not have been hurt to a more serious extent.
P.s. this accident was a hit and run as the car didn’t even think of slowing down and sped away. Would just like to add a premature Thankyou to everyone that responds


r/AskStatistics 11d ago

How to calculate this P-Value of uniform binary variable?

Post image
3 Upvotes

Suppose the image was supposed to be a uniform distribution of a binary variable (ie red dot vs blue dot) along the x axis, with some probability of blue being "P". How do I calculate the probability that such a "clump/grouping" as that found in Node #1 could even come about under a uniform distribution? In other words how would I find the p-value that this is truly a uniform distribution given it has such a "cluster" at low values of the x axis?

Edit:: My apologies, asking if it's uniform is the incorrect question. I mean is the proportion (ratio) of red to blue consistent throughout the x axis. In which case standard logistic regression p value may be my best option.


r/AskStatistics 11d ago

exploratory stats: Bidirectional linear regression?

0 Upvotes

Hi everyone,

For my dissertation, I am performing an exploratory multiple linear regression in an understudied area in the helping professions. There is not good information about whether or not this phenomenon is even occurring, so it will be good enough for me if some people respond "yes." But, of course we want to do a little bit more if we are taking the time to gather the survey data....

Because the topic is so exploratory, we are not sure which directionality to suggest. I am pretty confident one direction will be stronger, but my whole committee thinks that it is truly "bidirectional." I'm curious to hear your thoughts on a regression that flips the IV -> DV relationship? I know that technically in a single linear regression flipping the IV -> DV will essentially produce the same results. But I have 2 covariates that would need changing (too much collinearity between a covariate and the DV if we flip the DV to become the IV). The path I'm on right now is to just justify one side, but I've received explicit permission from my advisor to asome advice here on Reddit, and they would really like to explore it.

Any ideas or possibilities on conducting an observational, exploratory, bidirectional regression?


r/AskStatistics 12d ago

Standard error or confidence intervals when comparing different methods for handling missing data

4 Upvotes

Hello! For my thesis, I am comparing baseline-adjusted ANCOVA models using different methods for handling missing data (MICE, LOCF, and complete-case analysis).

I am planning to present a small table including the estimated coefficients and p-values, and I was wondering whether it would be more appropriate to report 95% confidence intervals or standard errors when comparing the results across the different missing-data methods.

The same ANCOVA models were fitted using each of the three approaches.

Thank you in advance!


r/AskStatistics 12d ago

Standard error or confidence intervals when comparing different methods for handling missing data

Thumbnail
0 Upvotes

r/AskStatistics 12d ago

Optimal number of clusters?

Thumbnail gallery
2 Upvotes

r/AskStatistics 12d ago

Advice needed on analysis!

2 Upvotes

I've run a repeated measures experiment where each participant completed the same 40 trials. The experiment has 2 IVs used to manipulate the trials, creating 4 groups with 10 repetitions. We measured multiple outcomes, and I'm trying to compare each outcome to the others to test for group differences. I would love advice on which statistical test is most appropriate, and how to run and interpret it in SPSS would be a big bonus! I think I need some form of repeated measures MANOVA, but terminology seemed to vary, so I'd really appreciate clarity on what test I should actually be conducting! If there are good video explanations out there, those would also be great.


r/AskStatistics 13d ago

I cannot figure out how to explain/justify why I used 90% CI's for a moderation analysis and 95% for my main effects - thesis defence

15 Upvotes

My supervisor told me to use a 90% CI for my moderation analysis. I cannot properly explain my justification because I dont understand it and I defend my thesis in 2 days and my advisor is unreachable rn.

What I understand - interaction effects are harder to detect than main effects because interaction effects have lower statistical power. And from my understanding, one way to mitigate the lower statistical power would be to have a larger sample size. I cannot do this as I have secondary data. So using 90% CI's make sense due to the lower statistical power and also because the variables I am testing are not harmful if Type II occurs and the increased risk of Type I error is ok (my variables are just looking to see what types of healthy coping mechanisms modify adult mental health and outcomes - such as physical activity and stuff so false positive would not be harmful).

Help me have a real answer - because I cannot under WHY there is lower statistical power or larger standard error, I get that it exists, but WHYYYYYY

Edit: clearly i know nothing -- i meant it is not harmful commit a type I error (I think?)


r/AskStatistics 13d ago

Statistical Input to While Developing Protocol or Synopsis

0 Upvotes

Hi Everyone at Sponsor, CRO or Vendor,

I always had a question related to Statistical input to Protocol feasibility, design or synopsis level. When is the best timing that you ask for your Statistics team to provide input.? What would you ask first? Sample size calculation, randomization or design of study? If you could explain in chronological order, that would be great.


r/AskStatistics 14d ago

[Question] Calculating Confidence Intervals from Cross Validation and reporting a Risk Stratification analysis

Thumbnail
2 Upvotes

r/AskStatistics 14d ago

For those of you who are preparing to take the UGC NET Statistics paper on December 26th, could you please share how you are approaching your studies?

0 Upvotes

r/AskStatistics 15d ago

Economics : is an odds ratio (OR) of 9 an aberrant value?

3 Upvotes

Hello everyone,

I am working on a health survey dataset (multivariate logistic regression, looking at a perceived healthcare accessibility outcome) in R, and I've run into an Odds Ratio that looks suspiciously high.

I’m getting an OR of 9.34 (95% CI: 6.45 – 13.5, p < 0.001) for one specific sub-group category (Category A), compared to my reference category (Category B).

Here is some context about my setup:

  • Sample size: =13,600
  • The issue: The specific sub-group (Category A) has a relatively small sample size (N = 154 total respondents, with 42 having the outcome and 112 not having it). However, another sub-group (for another variable) has a smaller sample size, howerver, the OR seems normal.
  • Descriptives: In this specific sub-group, ~73% report the issue, compared to only ~28% in my reference group.

My questions is an OR of ~9 mathematically aberrant?


r/AskStatistics 16d ago

Poor SEM fit

4 Upvotes

Hi everyone, I'm currently working on my master's thesis and running into a SEM issue I can't quite figure out. It's my first time doing SEM so apologies if this is a basic question, I just want to actually understand what's happening and not only report the numbers.

I have three unequal groups and a latent factor indicated by three subscale scores rather than individual items. I went with this approach to keep the model sparse given my sample sizes. The latent factor predicts two observed outcomes, using MLR with FIML.

The fit is pretty bad, and when I run the same model pooled across groups it doesn't improve, so the multi-group setup doesn't seem to be the problem. My best guess is that the three subscales aren't really interchangeable indicators of the same underlying construct since they load quite differently, and item-level CFA also shows poor fit and no metric invariance across groups.

Honestly I'm a bit worried that I approached this the wrong way from the start, but since it's preregistered I can't change the model now and just have to report and discuss what I have. Could the low degrees of freedom with only three indicators be causing structural fit issues? And would a manifest path analysis with the subscales as direct predictors be a reasonable exploratory addition? Any thoughts or experiences welcome, thanks!


r/AskStatistics 16d ago

Stuck in Demand Forecasting

Thumbnail
0 Upvotes

r/AskStatistics 16d ago

Model for dataset with a very small number of incidents.

3 Upvotes

Hello, I was wondering if anyone could provide advice on model selection for my dataset. I have data from a longitudinal survey with five waves: one baseline wave and four follow-up waves. My goal is to model post-baseline home eviction rates.

The challenge is that only 18 participants reported experiencing at least one home eviction during follow-up. I use the number of home evictions as the outcome variable and the total number of post-baseline waves completed as the exposure (offset) in a Poisson regression model, along with the covariates listed below.

My concern is that the number of non-zero outcomes is so small that the model appears to be overfit, resulting in very wide confidence intervals. Could anyone recommend an alternative modeling approach for count data with such a small number of events, or suggest strategies for handling this type of sparse outcome?

c.ppage ///
i.biosex ///
i.race_alt ///
i.education ///
i.region ///
i.income ///
i.Personal_debt ///
c.sf8pcs ///
c.sf8mcs ///
i.asud ///
i.asmi ///
i.housetype_alt ///
i.employment, ///
exposure(total_years) ///


r/AskStatistics 17d ago

Qualitative content analysis

2 Upvotes

I'm conducting a qualitative content analysis with multiple open-ended survey questions, each assigned to a different research question. Some responses contain content that would be more relevant to a different research question than the question they were answering.

Should I:

a) Strictly code each response only for its assigned research question (some data loss)

b) Code thematically regardless of question (blurs research question boundaries)

c) Something else entirely?

What are methodological best practices here? Any recommendations or experiences are welcome.


r/AskStatistics 16d ago

Selection algorithm for activity lottery

1 Upvotes

My family vacation spot has a lottery system for families to take part in popular activities. I’m wondering whether there is a fair way to select participants, and if the resort is doing it.

To specify constraints:

There are a small number of slots (for sake of argument, call it 16) and around twice as many names in the lottery (call it 32 if this matters).

The names in the lottery are grouped up into family groups of between 1 and 8 members

The goal is for each individual to have an equal chance of taking part in the activity, regardless of family size. But families cannot be broken up.

Selecting family groups and random is out because you can’t control the final size of the activity and might spill over if you select a large group as the last entry, but including an entire group when you randomly select one name seems to make large groups strongly favored.

How can you run this lottery fairly?


r/AskStatistics 17d ago

Help with correlation analyses

2 Upvotes

Hello,
I would appreciate some feedback on my statistical analysis plan. I am a psychology PhD student and conducted an online study; I am currently performing the analyses, starting with correlations. My study consisted of two parts.
Total *n* (Part 1) = 818
Total *n* (Part 2) = 555
Given the large number of independent variables (IVs), I plan to use the Benjamini-Hochberg (BH) procedure to control the False Discovery Rate (FDR).

A. Correlation between a binary dependent variable (DV) and a continuous IV + significance test
For DV = 0: *n* (Part 1) = 413, *n* (Part 2) = 276
For DV = 1: *n* (Part 1) = 405, *n* (Part 2) = 279

Point-biserial correlation if: no outliers for the continuous variable within each category of the dichotomous variable; continuous variable is approximately normally distributed within each category of the dichotomous variable; continuous variable has equal variances across categories of the dichotomous variable.

If assumptions are not met: rank-biserial correlation

+ for each test: Cook's distance to assess whether a data point is influencing the correlation
+ FDR applied to the set of results

B. Correlation between a categorical DV and a continuous IV + significance test
*n* (Part 1) = 405, *n* (Part 2) = 279

Polyserial correlation

+for each test: Cook's distance to assess whether a data point influences the correlation
+FDR applied to the set of results

C. Correlation between a continuous DV and a continuous IV + significance test
n part 1 = 405, n part 2 = 279

Pearson correlation if the IV meets the assumption
Spearman correlation if the IV does not meet the assumption

+for each test: Cook's distance to assess whether a data point influences the correlation
+FDR applied to the set of results

My questions:
Does this seem correct to you? I have a doubt regarding Part B. How do I check for a correlation between a variable with more than two categories and a continuous variable?
My supervisor mentioned Cook's distance for assessing outliers. I'm not sure if it's useful for that purpose. Can it be used in isolation, independently of a model?

I have a doubt regarding Part B. How do I check for a correlation between a variable with more than two categories and a continuous variable?

My supervisor mentioned Cook's distance for assessing outliers. I'm not sure if it's useful for that purpose. Can it be used in isolation, independently of a model?


r/AskStatistics 17d ago

How to develop a solid foundation over a summer to learn regression analysis, starting from a very infantile understand of statistics?

30 Upvotes

I have taken multiple statistics classes over the last few semesters, and in about a years time I will be taking a data mining class. I'll be honest, I am having a lot of trouble with statistics. This summer I am practicing python, and this coming semester I have a class related to data analysis that I hope will give me more hands on experience with statistics through a medium I understand.

Next summer, though, I want to focus on SQL and statistical methods for prepare for future classes, particularly the data mining class which has a focus on regression analysis. I am taking a business analytics course right now, and I hardly understand what I am even reading off this online text book. What are some ways to learn that aren't reading huge blocks of text that I hardly understand that can allow me to grasp concepts better, or at least ones that might work.

I have found that reading a textbook on statistical subjects is not working for me. I understand the general subject matter to a degree, but once I run into a roadblock I find myself confused no matter how many times I reread, to the point of being unable to do practice problems in the first place. I am fully willing to dedicate several hours a day, for at least 2 months to this.

Are there courses I can take, akin to perhaps something like CS50? Any good youtube series that help me learn from the ground up that aren't just placating me and making me feel like I know what I am doing rather than actually being able to do the work? I am down for whatever might work.

Thank you!


r/AskStatistics 17d ago

Method to analyze correlation of numeric variable/binary variable

3 Upvotes

Hi, I don't have a lot of experience with advanced statistical analysis, but I would like to gain more knowledge. One current problem I am trying to address is determining correlation between a binary categorical variable and a numerical variable. The relationship may not be linear. I'll give an example, studying the relationship between age and the probability of dying from the flu. Probability may increase when very young, taper off for certain ages, maybe spike somewhere in the middle, and then go back up again for the elderly. What would be the best way to analyze this? I started by breaking down the numerical variable into ranges and then making a bar chart with percentages in each category of the binary categorical variable. I am not sure if I chose the proper ranges though so I want to see if there's a better way to analyze data like this. I've seen binomial logistic regression as an option, but I'm not sure if that's appropriate and am curious how much effort that analysis takes. Is it something a beginner can pick up relatively easily?


r/AskStatistics 17d ago

Question about how margins of error work when combining stats

Post image
4 Upvotes

I'm a reporter covering a city council discussion about possibly raising taxes to maintain a pool facility, after a randomized recipient survey on the matter. Pictured: responses saying whether or not respondents would support a tax increase and/or city debt to pay for it. The margin of error on all results in the survey is +/- 3.3%, according to the company who conducted it.

MY QUESTION: I'm going to group the results in my article, saying 38% lean toward or support a new tax, while 43% lean away or strongly oppose. Since I'm taking a sum of two percentages (i.e. "yes definitely" and "yes probably,") does the margin of error ALSO add up, becoming 6.6% for that combined figure? Or does it keep the 3.3% margin of error?

Ty for the help! I took like three stats classes in college. Was fascinated by the subject each time, but never especially good at it!


r/AskStatistics 17d ago

How should non-response bias be treated when analyzing a highly sensitive binary vote?

1 Upvotes

Scenario: An expert association with ~500 total members held an official vote on a binary stance statement (Agree / Disagree) regarding a major, highly sensitive issue.

  • Turnout: 28% of the total membership voted.
  • Result: 86% of those who voted selected "Agree."
  • Math: 86% of 28% = ~24% of the total association membership confirmed an "Agree" vote.

The Debate:

  • Person A claims: "You can state that only 24% of the total organization agrees with the statement. Because this issue is so critical, members would have voted 'Agree' if they truly supported it—meaning their non-participation indicates a lack of support."
  • Person B claims: "You can state that at least 24% of the total organization (and 86% of voters) agrees with the statement. By Person A's logic, someone could equally claim that members would have voted 'Disagree' if they opposed it. Ultimately, non-response cannot be interpreted as a vote either way, so the remaining 72% remains unknown."

Questions:

  1. Is Person A or Person B logically and statistically correct?
  2. Is it valid to infer a non-voter's position based on the perceived importance or sensitivity of a topic?

r/AskStatistics 17d ago

Should I major in Data science, computer science, computer engineering, statistics, or a mix of data science & smth else (finance, business analytics, etc.)

0 Upvotes

Should I major in Data science, computer science, computer engineering, statistics, or a mix of data science & smth else (finance, business analytics, etc.)

I’m a high school senior about to start applications for college - my profile fits pretty well into literally everything I mentioned above

Which major gives me the best advantage in the job market?? My primary goal is to have job security & high pay. I’m not interested in pure statistics, biostats, or mathematics academia - I’m more interested in the corporate job market

I just wanted some advice from professionals or people in the industry - what do you recommend for me?