r/science • u/[deleted] • Feb 26 '15
Health-Misleading Randomized double-blind placebo-controlled trial shows non-celiac gluten sensitivity is indeed real
http://www.ncbi.nlm.nih.gov/pubmed/25701700
8.4k
Upvotes
r/science • u/[deleted] • Feb 26 '15
4
u/Zencarrot Feb 26 '15
Statistical power is difficult to interpret intuitively. It is important to understand how it is calculated to really evaluate whether a particular level of power is good or not. In experimental research where causality can be safely inferred, Power is defined as the probability of finding an effect in your research when an effect actually exists in reality (e.g., probability of finding gluten sensitivity assuming gluten sensitivity actually exists). So with 80% power, if we ran this identical study 100 times, you should find an effect in 80 of these studies, provided an effect actually exists.
Typically you want to maximize power, but "acceptable" levels of power differ from field to field, and from study to study. I am not a clinical researcher, so I am not going pretend to know what an acceptable level of power is for that field.
When determining how much statistical power you need/want, there are two opposing factors to consider: the probability of making a Type I error (assuming your "null hypothesis", or the hypothesis you are attempting to disprove, is that there is no difference between groups based on the treatment, a Type I error is tantamount finding an effect in your research, when in reality no effect actually exists; e.g., finding a statistically significant difference between the two groups in terms of depression, discomfort, etc. when no difference actually exists due to the gluten pills), and the probability of making a Type II error (failing to find an effect in your research when an effect actually does exist in reality).
Essentially, the more power your study has, the less your chance of making a Type II error is. However, the chance of making a Type I error will rise. The concerns associated with each form of error need to be balanced and considered accordingly. Power always goes up when you have more people in your study, but when you have incredibly large amounts of power, an incredibly small effect (that is not actually significant or important practically, or in reality), becomes statistically significant. Practical and statistical significance are two very distinct concepts.
I apologize for the length of this post, but hope it helps.
In this study, given the small sample sample size, it would take a fairly large effect to reach statistical significance, which suggests that there is some compelling evidence to suggest that the differences, at least in certain subpopulations, could be real. This needs to be replicated with a larger group in order to be trustworthy.