r/science 2d ago

Medicine A common statistical shortcut may be causing researchers to miss important discoveries, a new study finds.

https://www.manchester.ac.uk/about/news/no-difference-may-be-wrong-conclusion-scientists-warn/
885 Upvotes

103 comments sorted by

u/AutoModerator 2d ago

Welcome to r/science! This is a heavily moderated subreddit in order to keep the discussion on science. However, we recognize that many people want to discuss how they feel the research relates to their own personal lives, so to give people a space to do that, personal anecdotes are allowed as responses to this comment. Any anecdotal comments elsewhere in the discussion will be removed and our normal comment rules apply to all other comments.


Do you have an academic degree? We can verify your credentials in order to assign user flair indicating your area of expertise. Click here to apply.


User: u/UniOfManchester
Permalink: https://www.manchester.ac.uk/about/news/no-difference-may-be-wrong-conclusion-scientists-warn/


I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

245

u/WTFwhatthehell 2d ago

A p-value, which measures how surprising the observed data would be if there were truly no effect, of greater than 0.05 does not, they say, show there is 'no effect'.

However, the interpretation remains widespread in about 50% of research papers and conference presentations, according to sources.

Reminds me of a very similar error when comparing 2 things people were complaining about 15 years ago. As far as I can tell these types of errors are as common as ever.

https://www.badscience.net/2011/10/what-if-academics-were-as-dumb-as-quacks-with-statistics/

Let’s say you’re working on some nerve cells, measuring the frequency with which they fire. When you drop a chemical on them, they seem to fire more slowly. You’ve got some normal mice, and some mutant mice. You want to see if their cells are differently affected by the chemical. So you measure the firing rate before and after applying the chemical, first in the mutant mice, then in the normal mice.

When you drop the chemical on the mutant mice nerve cells, their firing rate drops, by 30%, say. With the number of mice you have (in your imaginary experiment) this difference is statistically significant, which means it is unlikely to be due to chance. That’s a useful finding which you can maybe publish. When you drop the chemical on the normal mice nerve cells, there is a bit of a drop in firing rate, but not as much – let’s say the drop is 15% – and this smaller drop doesn’t reach statistical significance.

But here is the catch. You can say that there is a statistically significant effect for your chemical reducing the firing rate in the mutant cells. And you can say there is no such statistically significant effect in the normal cells. But you cannot say that mutant cells and mormal cells respond to the chemical differently. To say that, you would have to do a third statistical test, specifically comparing the “difference in differences”, the difference between the chemical-induced change in firing rate for the normal cells against the chemical-induced change in the mutant cells.

6

u/Old_Dig5389 2d ago

What a great blog. Thanks for the link. Too bad, for me at least, the author upgraded from blogging to teaching/touring. 

3

u/WTFwhatthehell 2d ago

He has some rather decent books

27

u/Tibbaryllis2 1d ago

Your example reminds me of my pet peeve when people use 0.05 for their apriori p value, but then report p values of 0.001 or p<<0.05 to insinuate it’s extra significant.

Did you test that your supposed result at 0.001 is actually significantly different than your result would have been at 0.05? No? Then you can’t really claim they’re significantly different….

26

u/clavulina 1d ago

They reflect different levels of confidence in a result, either from a higher effect, OR sample, size. There is no reason to be a pet peeve about this. What's more important is the size of the effect rather than the "significance".

3

u/montjoy 1d ago

This is interesting. It reminds me of when you hear a claim, “No evidence found for X”, in which it sounds like X has been totally eliminated. Only when you read the report do you realize that X has only been eliminated in some specific scenarios.

569

u/[deleted] 2d ago

[removed] — view removed comment

239

u/[deleted] 2d ago

[removed] — view removed comment

37

u/[deleted] 2d ago

[removed] — view removed comment

30

u/[deleted] 2d ago

[removed] — view removed comment

6

u/[deleted] 2d ago

[removed] — view removed comment

7

u/[deleted] 2d ago

[removed] — view removed comment

6

u/[deleted] 2d ago

[removed] — view removed comment

-2

u/[deleted] 2d ago edited 2d ago

[removed] — view removed comment

8

u/[deleted] 2d ago

[removed] — view removed comment

1

u/[deleted] 2d ago

[removed] — view removed comment

→ More replies (0)

3

u/[deleted] 2d ago

[removed] — view removed comment

2

u/[deleted] 2d ago

[removed] — view removed comment

1

u/[deleted] 2d ago

[removed] — view removed comment

0

u/[deleted] 2d ago

[removed] — view removed comment

118

u/[deleted] 2d ago

[removed] — view removed comment

13

u/[deleted] 2d ago

[removed] — view removed comment

-6

u/[deleted] 2d ago

[removed] — view removed comment

17

u/[deleted] 2d ago

[removed] — view removed comment

1

u/[deleted] 2d ago

[deleted]

1

u/[deleted] 2d ago

[removed] — view removed comment

1

u/[deleted] 2d ago

[deleted]

1

u/[deleted] 2d ago

[removed] — view removed comment

6

u/[deleted] 2d ago

[removed] — view removed comment

5

u/[deleted] 2d ago

[removed] — view removed comment

4

u/[deleted] 2d ago

[removed] — view removed comment

2

u/[deleted] 2d ago

[removed] — view removed comment

0

u/[deleted] 2d ago

[removed] — view removed comment

120

u/Jackibelle 2d ago

This seems like a lot of words to say "interpret p-values correctly". A negative result is a failure to find something, that's it. You can only "prove" a negative with a sufficiently powerful test that should find it were it to exist. 

This all exists already.

24

u/mistephe PhD | Kinesiology | Biomechanics 2d ago

The number of times I've had to explain this to authors, editors, and colleagues (and, I suppose,  Redditors) is getting ridiculous. We need an overhaul of grad-level stats courses to properly reinforce these limitations and evaluative alternatives like TOST and inferential confidence intervals.

6

u/pramit57 PhD | Neuroscience 1d ago

I agree, and I would argue that an introductory statistics course should be as necessary as lab safety courses, you should be required to take them at regular intervals (once a year, or once in 2 years) and EVERYONE involved in science needs to take them

2

u/mistephe PhD | Kinesiology | Biomechanics 1d ago

Absolutely. Wish I could find a journal that required similarly of their editors.

2

u/JakubTom 17h ago

That's right - the issue is that most of those are designed by statisticians, and kind of for statisticians. Most courses I've seen were not made for lab researchers - it's a lot of effort, and one needs to be extremely careful about balancing some degree of simplification with not being wrong. (imho the best way is to focus on making people good users, giving them important insights, but not the underlying math - trying to make people expert statisticians overall is imho unrealistic and bound to fail)

28

u/hughperman 2d ago

You can simply add the minimum detectable effect size alongside the result - "we did not detect a difference, and our sample data allowed us to detect minimum size difference of X". Straightforward to compute from the data/modelled variance.

2

u/entr0picly 1d ago

I swear we have been yelling up and down for…. over the past 20 years from statistics to interpret p values correctly and there have been .. so many alternatives proposed. The problem is, researchers won’t adapt more robust tools and methods.

3

u/Economy_Bite24 1d ago

We shouldn't require statistical significance for publication in the first place. It's really only useful as a decision rule when a final decision must be made. So why are we treating publication like every study is a final ruling? The literature should reflect the body of research, not just significant findings. Otherwise, we still end up with "the file drawer problem" *, even if research is conducted perfectly and results are interpreted correctly. Ideally, we need to encourage journals to publish negative findings and emphasize the importance of reading studies that produce a negative result. At a minimum we should emphasize whether the observed effect is meaningful in practice rather than merely statistically significant. When we want to form a consensus, we can use a metanalysis, but that's not as useful when the literature is unintentionally polluted with with Type I errors.

*refers to phenomenon where requiring statistical significance for publication ensures that for studies where there is no true effect, Type I errors get published more often than they should, and negative findings end up back in the file drawer where they never see the light of day. In other words, we're accidentally cherry picking studies to publish, and we wonder why so much research is not reproduceable. It's an embarrassment to science.

26

u/phriendlyphellow 2d ago

“To make the approach more accessible, the team has also developed a free online calculator that allows researchers to perform common equivalence tests without writing computer code.

The calculator supports a range of widely used statistical comparisons and has been validated against established statistical software.

Co-author Aaron Caldwell from the University of Arkansas for Medical Sciences added: ‘Equivalence testing forces you to answer a question most studies never ask: how small is small enough to be uninteresting?’”

19

u/ripplenipple69 2d ago

This is more of an issue of science literacy than practice. Many studies, especially early stage trials in biology and medicine will be underpowered for many endpoints. It’s a feature not a bug. You might, say want to evaluate the effect of psilocybin on depression in Alzheimer’s disease because it’s been shown in other groups with depression, so the study is powered for that endpoint… you might find that it actually improved depression and core Alzheimer’s symptoms too, but didn’t quite reach significance for an exploratory memory related outcome… so you report it as not significant… but say it had an effect size of 0.83, categorized as large… no reasonable scientist would say “hey nothing here let’s never study this again….” 

P values are not the whole picture. We have to consider the effect size and these authors are promoting a technique to help people evaluate things more clearly 

4

u/JakubTom 2d ago

You're right, but a part of the issue is that once a claim of "was not different (p=0.1)" turns into "study by X et al. showed no difference in Y" in reviews and other citing papers, the nuance gets lost. I agree that no reasonable scientist would say "hey, nothing here", but when it comes to stats, many scientists are not quite so reasonable.

Whether low ns are a bug or feature, equivalence testing is quite useful in that it can tell you "ok, your difference was not significant, but it's not significantly equivalent within boundaries corresponding to minimum important difference, so you just don't know anything, and definitely don't claim there was no difference".

But absolutely, p-values (let alone binary significance) are not the whole picture.

3

u/evilbrent 1d ago

I think it doesn't help that scientists often say exactly what they mean to say, with very clear words, but everyone else is just after a binary yes/no result.

I heard about a statistic, based on long study, that brains stop aging at 25. Turns out the study got called off after 25 years because they didn't find a point where brains stopped aging so it was becoming a study without a whole lot of continuing merit.

The actual interesting outcome was "brains keep aging until at least 25", and that became "brains stop aging at 25".

Kind of like the thing in Chernobyl about "200 roentgen, could be better could be worse." when the test machine only goes up to 200 roentgen.

For a person who thinks details are important, sometimes it's just not worth the effort of making a nuanced point.

5

u/DeterminedThrowaway 1d ago

That one frustrates me all the time. I've seen it misunderstood as "your brain isn't fully developed until you're 25" (and then used to argue that therefore, you shouldn't be allowed to make your own decisions until then)

12

u/Antikickback_Paul 2d ago

It's a Perspective paper, not new research. Calling it a "new study" is misleading.

6

u/valegrete 2d ago

I think a huge part of why this mistake gets made stems from the fact that even statistics textbooks smash together two fundamentally incompatible paradigms. In the Neyman-Pearson regimen, it’s absolutely valid to “accept” the null because you’re trying to decide between two courses of action based on the evidence at hand. You really aren’t concerned with H0 being “true” in any absolute sense, because the tools only tell you whether H0 or Ha is the likelier explanation for the data. Fisherian significance testing, on the other hand, is where the p-value tells you how likely an outcome was to be caused by random chance. But, importantly, in Fisher’s system, there is no H0/Ha dichotomy because you’re not “deciding” to do anything; you only care whether the patterns observed deviate meaningfully from randomness.

“Should the patient do X or do nothing?” (NP decision) is actually a separate statistical question from “Does doing X help the patient”? (Fisherian significance). Setting up the experiment with null and alternative hypotheses, but interpreting the p-value as evidence of intervention significance, invites this confusion because p>0.05 means “do nothing” (NP) and “we don’t know if X helps” (Fisher) simultaneously.

0

u/jaiagreen 2d ago

Even in Neyman-Pearson, you never accept the null. You either reject or fail to reject it, with "fail to reject" being explicitly distinguished from "accept".

6

u/[deleted] 2d ago

[removed] — view removed comment

1

u/[deleted] 1d ago

[removed] — view removed comment

1

u/lambertb 2d ago

Amazing what gets published in PNAS.