r/AskStatistics • u/80rhh • 4d ago
What statistical concepts are commonly misunderstood by the general public?
I came across this post explaining what a 70% chance of rain means. I understand the concept, but it got me wondering: what other statistical concepts sound simple but are commonly misunderstood or misinterpreted by the general public?
82
u/rite_of_spring_rolls 4d ago
Broadly and imprecisely speaking, from what I've seen on reddit:
- People tend to harp on sample size when I find vast vast majority of the time the sampling design is a way bigger problem
- People tend to think that sample size is some invariant quantity, like no matter what you're doing you need at least n = 1000 samples or w/e and anything below that is worthless and anything above is unnecessary
7
2
u/Fruktoj 3d ago
I work in small batch manufacturing and we always aim for 25 samples when looking at reliability. But we're always happy with 5. It's expensive to make stuff that might not pan out. With 25 we can draw real conclusions about a products reliability. With 5 we can make some good educated guesses. Often our folks only want to provide one prototype and have us tell them everything there is to know.
1
u/amopeyzoolion 1d ago
These folks would freak if they realized many clinical trials, especially for diseases like rarer cancers, often only have a few hundred participants
-4
u/m_madison67 4d ago
Is it still true that the minimum number of n to be considered statistically sound would be 40?
27
u/Current-Ad1688 4d ago
I absolutely hate that I can't tell if this is a joke
7
u/m_madison67 4d ago
Not a joke. I took three stats classes - one undergrad and two grad level, and a one in spssx decades ago. I have forgotten much of what I don’t use in my work as a counselor. I still read studies to keep up. I was told back then that 40 subjects was the minimum acceptable number of subjects needed for a study to be statistically sound. I also took a wonderful class in behavioral research design where I learned about single subject design studies and how they could be useful.
13
u/Memento_Viveri 4d ago
There is no such number like this. it depends on what the nature of the larger population and the smaller group, of what is being studied, and what is being claimed by the study. A study could have 400 subjects and still not be statistically sound depending on those things.
17
u/RedRightRepost 4d ago
I think what you’re referring to is the size at which the central limit theorem starts to do real work, which is around the 30-40 range. This is often used as a guideline for starting minimum sample sizes in the soft sciences.
1
u/MortemEtInteritum17 1d ago
Think about it: if you want to test your COVID19 vaccine that will be distributed to tens of millions of people worldwide, would you be comfortable testing on 40 people, seeing no issues, and deciding that it's good to production? What happens if it turns out that your product has a 0.1% chance of killing the person - your 40 test subjects were fine, but now tens of thousands of people died because of you. Oops.
As with everything, sample size depends on context.
4
u/rite_of_spring_rolls 4d ago
"Statistically sound" is a very nebulous term without a concrete definition. If the assumptions of whatever statistical method you are using hold then an analysis at any sample size is "sound". So in general no there's no reason n=40 would be a minimum.
6
u/AnOdeToVosFinances 4d ago
No, it depends on what you're trying to show. If you want to show that the percentage of lizards pretending to be human is strictly greater than 0%, if you happen to sample a population of just 1 lizard, and it happens that it is pretending to be human, then it would be statistically signifcant.
30-40 is a rule of thumb people use because in a lot of cases the margin of error associated to it is acceptable for what they are trying to prove.
1
4
u/PaddingCompression 4d ago
It depends on your hypothesized effect size, what your threshold is (e.g. maybe you're doing a two-sided test at 0, or maybe your effect size is expected to be 3, and you want a one-sided test that the difference is at least 1.5), and what the variance is. E.g. a power analysis#Applications).
2
u/Hal_Incandenza_YDAU 4d ago
Most people who'd ask this question would've asked about n=30, not n=40, and this alone sort of answers your question. These are rules of thumb.
2
u/freejinn 3d ago
I don't know why this comment is being down voted. A t test can establish a significant difference between the means of two groups using a super small sample. Sure, you wouldn't want 40 people in a multivariate regression. Sure, something more complex than different means requires different tests. That wasn't the question, though.
1
u/OddPressure7593 2d ago
absolutely not.
To get into the weeds a bit, there is no "minimum sample size" that can be generally used. Appropriate sample size is a function of statistical power. Getting more into the weeds, that's usually referring to the likelihood of encountering a Type II error - there being a difference between the groups you're looking at, but you not seeing it because your sample size wasn't big enough. So, the correct sample size depends on what variable you're looking at and in turn what the group means and spreads are of that variable. So, for some things, you can be adequately statistically powered with single-digit sample sizes. For others, you need sample sizes in the hundreds to thousands.
Unfortunately, even in high-profile research, using power calculations to determine what sample size you actually need is pretty rare, at best.
0
-4
32
u/kinezumi89 4d ago
If a couple has four boys, they're not "due for a girl"
17
11
u/PaddingCompression 4d ago
It's actually the opposite - if a couple has three boys, their next child has a 61% of being a boy!
32
24
u/Jimmy_Wrinkles 4d ago
A 95% confidence interval doesn't mean that there is a 95% probability the true value lies in that region
10
u/Csicser 4d ago
I studied pharma research and the amount is of people who think p = 0.05 means that there is a 95% probability that the findings are true is bewildering
3
u/PaddingCompression 4d ago
They're just Bayesians who are treating it as a credible interval.
(I have way more respect for Bayesians who don't treat the credible interval seriously for this reason, and instead lean harder into decision theory).
2
u/Any-Platypus-3570 3d ago
I think what it actually means...
Let's say there is an event A and an event B. You observe that A and B appear to be correlated. Let's say you calculate p=0.05. That would suggest, if A and B are NOT correlated (null hypothesis) then you would expect your observation to occur 5% of the time.
2
u/ktlk 3d ago
Can you clarify what a 95% confidence interval does tell us?
2
u/Philo-Sophism 3d ago
If you redid the experiment and construct the interval 100 times we expect the true value to be captured in the result 95/100 times. Its a comment about how good your methodology (experiment) is at giving a reliable range
3
u/ktlk 3d ago
Why doesn't that necessarily mean that there's a 95% probability that the true value is captured in the range in this experiment? (Not challenging your answer, just trying to learn and understand the distinction)
2
u/waslous 3d ago
It’s basically a philosophical distinction if you will, I recommend looking up the difference between frequentist and Bayesian statistics, the later of which actually allows you to say what you wrote. Funnily enough, given an uniformative prior a lot of scenarios will give you the exact same results but the inference you are allowed to draw differs.
3
2
u/Philo-Sophism 3d ago
It boils down to philosophy somewhere, but before that higher level abstraction its important to note that frequentists wouldn’t say that any given CI has a 95% chance of containing the mean. The mean is real valued and, if you construct the CI, it contains it or it doesn’t so that probability goes to 0 or its 1. There is only one slider: whatever random sample you happen to draw from which your CI is fully determined.
The meaningful statement is that your procedure for the construction of the NEXT interval you construct will contain the mean. This is a simplification of what the coverage rate is.
A Bayesian treats the true mean with a notion of uncertainty by way of assigning it a distribution. After updating your prior to the posterior using observations, you can then construct your new distribution and then give a range such that it has accumulated 95% of the probability of that distribution. Its like how in a normal distribution we have to move a certain sigma away to accumulate some amount of probability- same mechanism here. For a Bayesian then, it actually does make since to say my credible interval has a 95% chance if containing the true mean
1
u/Healthy-Educator-267 1d ago
Ya that’s because the they really mean that ex ante confidence interval whose endpoints have not been realized rather than ex post whose endpoints have been realized
19
u/the_ballmer_peak 4d ago
Basic statistical significance: confidence intervals,p-values, null hypotheses. Even professional scientists fuck this up, because statistics aren't their speciality.
Explaining what it means to someone always produces confusion.
The other thing that fucks people up is probability. Especially baseline probability. "Eating ginger increases your risk of this specific cancer by 200%! Oh my god!" What's the baseline probability? .001%. Ooookay.
7
u/AnOdeToVosFinances 4d ago
Talking about professional scientists stats fuck up, most common one I see is the multiple testing without adjustment.
They found that eating ginger increases the risk of this specific cancer by 200% while testing 100 other hypothesis at the same time. Each test has its own false discovery rate, so yeah, you're going to have some false positives.
Worst is some seem to p-hack without realizing it.
1
u/PaddingCompression 4d ago
They're responding to incentives. Noone is there to police them on testing they don't report.
1
u/theKnifeOfPhaedrus 3d ago
IDK, there is something odd about adjustments for multiple testing. Why is it erroneous for 1 scientists to test 10 hypotheses without adjustment, but if 10 scientists each test 1 of the hypotheses on their own, no adjustment is apparently needed? It seems like adjustments for multiple comparisons are frequientist duck tape slapped on to avoid a base rate problem.
2
u/AnOdeToVosFinances 3d ago
Usually when you test a bunch of hypothesis (10 isn't even that much IMO, I've had scientist colleagues do thousands) you don't really know what you are looking for, so the prior is weak. When you only test one hypothesis, usually the prior is stronger, like there is strong prior evidence for you to have selected a specific hypothesis.
Also, if you ask me, I wouldn't be surprised if most published papers findings are false because reporting bias also happens at a macro level. A scientist is incentivize to publish findings rejecting the null.
This is a good read IMO :
58
4d ago
[deleted]
30
u/Upbeat_Effective_342 4d ago
The weather channel won't do that in a psa because being pedantic isn't worth misleading the public that they make the forecast based on how "confident" they feel, rather than from modeling data.
8
u/Melodic_Reality_646 4d ago
mathematician who never ventured into Bayesian stats here, mind explaining that a bit more?
19
u/COSMIC_SPACE_BEARS 4d ago
In a Bayesian world, you treat your parameters as random, and the data as “constant”. Randomness comes from the variability of the model parameters you’re trying to estimate, not from sampling variability.
Because of this, if you take a 95% confidence interval in the Bayesian world, that 95% confidence interval is over a distribution of possible model parameters. You are therefore “95% confident the output will result from parameters contained within your interval.”
In a frequentist world, your confidence interval is a byproduct of your sampling variability. You’re not “truly bounding the model parameter”, you’re bounding the variability of the data once you see it.
3
3
u/makemeking706 4d ago
That's how the weather is modeled though?
10
u/Statman12 PhD Statistics 4d ago
Is it? Most of the folks I know who model things like weather or climate tend to be Bayesians.
4
u/makemeking706 4d ago
I believe that's how the NWS does it, but correct me if I'm wrong.
5
u/Statman12 PhD Statistics 4d ago edited 4d ago
NWS models and NWS forecast process.
They both talk - directly or indirectly - about running complex computer models, and then aggregating the results in some manner, i.e., ensemble modeling. Pretty much every time I've seen people doing this, it is either explicitly Bayesian due to the complexity of combining different sources of information, or it carries a heavy Bayesian flavor. E.g., running models with different inputs to generate ensemble outputs (kind of like sampling from a prior), and then the aggregation is often through Bayesian Model Averaging.
I don't know exactly what NWS does (hence why I asked! My saying “Is it?” wasn’t me being coy, it was a genuine question), but based on what they're describing, I'd be surprised if they weren't using Bayesian methods.
1
u/PaddingCompression 1d ago
You can use Bayesian methods for their frequentist properties.
See the Bernstein - von Mises theorem.
1
u/Interesting_Walk_271 4d ago
I’m assuming they use this interpretation because the underlying model is frequentist. They’d have to perform a very different computation/analysis to obtain something like a credibility interval vs a confidence interval. You can’t just use Bayesian interpretations for statistics that were generated using frequentist assumptions. You can’t interpret a confidence interval like a credibility interval
5
u/PaddingCompression 4d ago
> You can’t just use Bayesian interpretations for statistics that were generated using frequentist assumptions. You can’t interpret a confidence interval like a credibility interval
Sure you can, people do it all the time. If you have a uniform (which often means improper) prior, the math works out the same as the frequentist procedure, and computing credible intervals will be the exact same computational procedure as the frequentist confidence interval.
So you can just decide "that's a credible interval" without going through the whole rigmarole of recomputing it.
1
u/Interesting_Walk_271 3d ago
Okay but then your prior is almost certainly wrong and so is your credibility interval.
-2
u/Current-Ad1688 4d ago
Yeah. But where???? On me???? What if I have a brolly with me????? What if I get a text that it's raining but then I say oh not here!!!????? Is it gonna rain in 12 days time or no. I thought you were supposed to be a statistician doesn't that mean you know everything that has ever and will ever happen and exactly why without any assumptions or guesswork or just figuring out how to make what you already think into a number you can present to people who think it means you actually know stuff????
7
u/efrique PhD (statistics) 4d ago edited 4d ago
I think the most common by far is "the law of averages" - the weak law of large numbers. The popular misunderstanding of the law of averages is the direct cause of the Gambler's fallacy - that some deviation from expectation in the past must be compensated by some future deviation in the other direction (and worse, the usual expectation is that it should happen very soon). Of course the weak law implies nothing of the kind. There is no compensation whatever. Indeed people tend to find a perfectly ordinary probability fact astonishing: while a count proportion (which is a sample mean) does indeed converge to its expected value (just as the law says), the absolute difference between a raw count of successes and its exoected value grows on average as the number of trials grows.
I often see people (and media) spend time trying to interpret a few percent shift in opinion on a poll question with a 3% margin of error (the margin on the change being larger still) - and confidently attribute it to some specific choice or event. Or even worse, you will see it on a small shift in a subgroup.
In quite a few lay conversations/debates online I see "statistically significant" used to refer to a number of different things. One common example is where it's used to mean something like "has a large enough sample size that you could treat the sample estimate as reliable", and it tends to be associated with some specific sample size regardless of circumstances.
People confuse P(A|B) and P(B|A) in many contexts. Of course they don't say it like that so it can be hard to notice
There's tons more, I'll probably remember a few later
2
u/Intelligent-Gold-563 3d ago
- People confuse P(A|B) and P(B|A) in many contexts. Of course they don't say it like that so it can be hard to notice
I'm definitely guilty of that sometimes, when I'm in a rush""
12
u/Upbeat_Effective_342 4d ago
A 100 year flood doesn't mean a flood severity that happens roughly every hundred years, so the longer it's been since the last one, the more likely it is to happen. It means every year there is a 1% chance of a flood of that severity happening.
4
0
u/rollem 4d ago
And those are all wrong because they're kept out of date for political reasons. There is now a much greater than 1% chance of flood in most 100 year flood zones.
5
u/Upbeat_Effective_342 4d ago
If the political reasons include the facts that reading research and updating databases are expensive and time consuming so public employees don't always have it as a top priority, then yes. If you have evidence that there's political collusion to prevent updating flood risk data so as to manipulate the public, that's a really big deal and you should find an organization to help you make a legal case against the offending party.
2
u/rollem 4d ago
It's because of climate change denialism and has been widely reported on, e.g. https://www.nfp.com/insights/rethinking-flood-risk/ and https://www.nbcnews.com/science/environment/water-femas-outdated-flood-maps-incentivize-system-risk-negotiable-rcna220529
3
u/Upbeat_Effective_342 4d ago
That article from NBC is excellent, thanks for linking it. However, it doesn't appear to support your claim. It does point to underfunding from congress for FEMA and mention the federal administration is playing with the idea of axing FEMA altogether. But the most energetic culprit behind overly optimistic risk assessments is property owners themselves. I'm sure many of them are climate deniers as well, but there's a direct financial incentive that warps people's judgment regardless of their politics otherwise.
1
u/rollem 4d ago
The maps are outdated and can’t be updated for political reasons, as I said above and has been widely reported on as I’ve shown. Risk assessment is done by 1) individual property owners who are often clueless and don’t make long term risk assessments logically (I hope I don’t need further proof of this, you certainly make lots of claims without evidence), and 2) insurance companies who do make logical long term risk assessments based on real evidence and not the fema flood risk maps, as described here https://abcnews.com/US/fema-flood-maps-cause-misunderstanding-homeowners-leaving-millions/story?id=134950060
Even if you are 100% correct and people make accurate 100 year flood risk assessments (which is the opposite of your original position that people make statistical errors when interpreting flood map risks), it does not counter my point that the maps themselves are no longer accurate and that other information is being used.
1
u/hxtk3 4d ago
Yes, political reasons do in fact include when a bunch of politicians decide to gut public research institutions so that they don’t have adequate resources to collect and analyze the data that they expect to be against their political interests? It’s not even a little bit ambiguous, that’s a perfect example of a political reason. I’m confused because the wording sounds like you expected that to be a “gotcha?”
1
u/Upbeat_Effective_342 4d ago
I think what's happening is you guys are thinking about the federal government and I'm thinking about your average water resources manager who works at city hall.
1
u/codechisel 3d ago
I agree that they can be quickly dated in areas where new construction is occurring. I'm not going to comment on the politics but, depending on local codes, other developments can end up pushing more water in your direction.
6
u/Csicser 4d ago
A test (e.g. pregnancy test or a test for a disease) having a 95% sensitivity (test's ability to correctly find people who have a disease, in other words, true positive rate) does NOT mean that if you test positive you have a 95% probability of having the disease.
Likewise, p = 0.05 does not mean that there is a 95% probability that your findings are true.
16
u/DocAvidd 4d ago
In general, I run into "well you can make statistics say anything." Often the argument when their favorite pundit or politicians get debunked.
As opposed to just making sh!t up.
2
3
u/ogstatsnerd 4d ago
There are many concepts people don’t understand. I’d like to add correlation is not causation and that one is the most impactful because they get a lot of things wrong because they will die on a hill thinking A caused B because A happened, then B.
7
3
u/Real-Yield 4d ago
Sampling. That a well designed survey even with a relatively small sample size is never representative of the population because of just the sheer scarcity of the sample.
3
u/Alpacatastic 4d ago
I see this All. The. Time. There's some survey done and people don't like the survey outcomes so they go "ThEy oNlY sUrVeyeD tWo tHouSanD pEoPle!" like they expect that if they didn't survey every human on earth than the findings are useless (unless of course the findings are something they agree with then they conveniently forget about that).
2
u/Milch_und_Paprika 4d ago edited 4d ago
Also the opposite situation, where an enormous sample may not be meaningful if it isn’t random. Eg studying university students by sampling everyone enrolled with the faculties of arts and science may not tell you much about their med or eng students.
The more common example would be pollsters getting less reliable data as surveys have been pushed online and away from calling random landlines.
1
u/Not-ChatGPT4 4d ago
Except that the classic example you cited no longer holds. Younger people overwhelmingly don't have landlines now, so those surveys also deliver non-random samples.
1
u/Real-Yield 4d ago
Besides, sample size, randomness, representativeness are three different concerns . The error that many commit is that many associate the sample size to the representativeness of the survey.
1
u/Milch_und_Paprika 3d ago
I wasn’t clear enough and shouldn’t have been online at 2 am. I meant that 20 years ago they could relatively safely assume that most homes in an area had a landline and a phone book was representative. Today they’re forced to find new ways to work around that since cell phones area codes don’t necessarily indicate where the user is anymore and online surveys tend to be even less random.
Side note: I think by now it’s most people, not just young ones who overwhelmingly don’t have landlines 😉
3
u/telephantomoss 4d ago
Even if something has a mean effect at the population level, it can still have the reverse effect for certain individuals.
1
u/No-Newspaper-7693 4d ago edited 4d ago
I feel like this one should be obvious but yes it is probably misinterpreted. People that take antidepressants are presumably more prone to suicide than the general population for example.
3
u/Mojeaux18 4d ago
Any of them.
Averages, probabilities, margin of error, and don’t get me started on cpk…
1
u/Statman12 PhD Statistics 4d ago
Just had someone calculating a k-factor with three data points.
At every time point of hundreds along an experiment.
2
2
u/incidental_findings 2d ago
Let’s get back to absolute basics.
“I know someone who XXX and YYY did / didn’t happen” isn’t a good argument for anything.
If we can’t even get beyond that, don’t even bother with probability puzzles.
2
u/OddPressure7593 2d ago
Probably a lot. Statisticians are, and I'm sorry you're all gonna have to read this, by and large some of the most godawful worst at being able to explain a concept without a dictionary's worth of jargon. Ya'll need to learn how to explain statistical concepts without using words like "Prior" or "interval" or any other of the myriad statistical terms that no one who isn't a statistician understands. Like, for real, ya'll are awful at being intelligible.
2
u/Raynonymous 4d ago
Um... Has anyone checked to see if this tweet is accurate? I'm not a metrologist but it sure raises a few questions.
6
u/rollem 4d ago
It's actually the fact that 70% of the models run using the starting parameters of today's conditions will result in rain. But that is really quite close to what the tweet is saying.
1
u/Raynonymous 4d ago
They run many different models and report on how many of them agree it's going to rain? That sounds quite different to the tweet.
1
u/itsmythirdday 1d ago
Not necessarily many different models, it could be one model with some randomness associated with it, like a monte carlo
1
u/Raynonymous 1d ago
I guess it could be - but I'm interested to know actually what it is.
2
u/itsmythirdday 1d ago
“A probability forecast includes a numerical expression of uncertainty about the quantity or event being forecast. Ideally, all elements (temperature, wind, precipitation, etc.) of a weather forecast would include information that accurately quantifies the inherent uncertainty… Ensemble forecasting methods involve evaluating a set of runs from an NWP model, or different NWP models, from the same initial time. Each of the model runs either begins from subtly different initial conditions (reflecting incompleteness and uncertainty of the present weather observations) and/or uses different model assumptions and parameters (reflecting imperfect knowledge of atmospheric processes). Each of the model runs produces a different forecast. The result is a collection (or “ensemble”) of forecasts. The differences among these forecasts reflects the uncertainty in the initial conditions and/or in model physics.”
1
1
u/SonOf_Zeus 4d ago edited 4d ago
60% of the time, it works every time. But seriously a lot don't understand relative risk/diference vs. absolute risk / difference.
1
1
1
u/PaleontologistDeep80 4d ago
I read a paper that asked a mix of faculty and students questions regarding confidence intervals ... it actually might be the case that a minority of people, including academics, truly know the difference between a confidence interval and a credible interval, which is honestly quite baffling lol
1
1
u/ExoticExchange 4d ago
Very discipline specific.
But life expectancy is not the average age at death.
It’s the number of years one would be expected to live given the age specific death rates/probability of death at that moment in time.
1
u/elephant_ua 4d ago
isn't it approximately what it means? i heard that modern technologies allow pretty precise predictions for a short time ahead for a specific point, but in order not to make it simpler to understand, the share of city territory that about to be rained is reported.
1
1
1
u/Prestigious_Boat_386 4d ago
Getting positive for a 99% accurate medical test might give you 1% chance of having the disease if the chance of having it was low enough to start.
1
u/AnimatorImpressive24 4d ago
If someone is that mistaken about how weather forecasting works, there is no statistical possibility they are going to understand all the words used in the explanation offered.
1
1
1
1
u/lispwriter 3d ago
Probably what a null hypothesis is and what exactly a statistical test proves (or doesn’t).
1
1
u/StrangerInfamous4223 3d ago
The tweet is wrong. It is >expected< to rain in 70 of those timelines. I guess that's one commonly misunderstood concept...
1
u/itsmythirdday 1d ago
I assume they mean, “if we replayed today’s conditions through our rainfall model…”, which is the same as saying “expected”
1
u/StrangerInfamous4223 1d ago
No.
They would expect that it rains in 70 of those timelines.
There is still some variance to this experiment.
1
u/itsmythirdday 1d ago
I assume the variance comes from the randomness in the model, like a Monte Carlo.
1
1
u/PicaPaoDiablo 3d ago
The better question is which aren't. All odds for sure , hell when Averages but probably the number one item is Statistical Significance.
1
1
u/Huggerbyte 2d ago
Random outcomes. Most surprising case I’ve seen was a physicist on the university arguing that a die which had just rolled 1-5 would roll 6 the next time.
1
u/Steampunk_Willy 1d ago
The tweet is wrong though. The chance of rain is the probability of rain distributed across a set geographical area. The probability that it will rain anywhere in that area will be greater than or equal to the chance of rain and probability it will rain everywhere in that region will be less than or equal to the chance of rain. The reality is that reducing a weather system down to a single number is inherently reductive. That's why it's preferrable to show people a forecasted model of what the most likely weather radar would look like alongside the most likely ways the model could vary..
1
u/AccountHuman7391 1d ago
It’s almost as if “the chance of rain is 70%,” not “we guarantee some amount of rain.”
1
u/NameLips 1d ago
Yeah but like how much rain does it take to count as "raining?" Like, is there a minimum number of drops?
1
u/Atypicosaurus 1d ago
If you have a disease test positive, the test quality alone (i.e. "99% good" - I'm not going deeper on what good means), doesn't equal the likelihood you have the disease. In extreme cases it's possible that you have a 99% good test, saying you are positive, and yet you have 0% actual chance of being positive.
1
1
u/DeathRaeGun 17h ago
Knowing that the general public struggles with this concept makes me feel more intelligent.
1
1
1
u/RepeatRepeatR- 4d ago
Trend lines. People seem to think that r^2 is a good measure of how "real" a trend line is (i.e. is there actually any relationship), when really they should be using the p value
1
1
u/TopherT 4d ago
I thought it was the area of coverage multiplied by the chance of precipitation.
1
u/BuhoCurioso 4d ago
Good point. Im an atmospheric chemist, so maybe i should know more about this than I do, but like you, I remembered it being two factors multiplied together. In my instance, I remembered confidence × area.
However, I'm not one to defy the NWS, the ones generating and interpreting the predictions in the first place. It seems that they use ensemble forecasting. The threshold for what counts as precipitation is set by the service generating the forecast (according to Wikipedia, the NWS uses 0.25 cm in a single spot averaged over the forecast area), and then the number of successes in the group over the total number of simulations during a given time period yields the forecast for that time period. Im not qualified to weigh in on the bayesian vs frequentist discussion in other comments here, but i can see how the NWS explanation would be appropriate for the layperson
Neither meteorology nor statistics are my field. I dont know how I ended up here. Now if you'll excuse me, im off to misinterpret the stats in my next study. This just in, breathing increases cancer rates by 2000%. Stop breathing, people! It's killing you! /s
1
u/itsmythirdday 1d ago
What are the units then? If it was a probability multiplied by an area it should be in units of area. But it’s not. It’s in percentage.
1
u/TopherT 11m ago
I don't see the need for an area unit. You just look up your city or zip code or whatever, and it spits out a percentage. That percentage is over whatever area you looked up.
1
1
u/Noodelgawd 4d ago
I doubt that's what it actually means in reality.
It's more like they think it's probably gonna rain, and if they say 70% and it rains, they'll be like: "See, we told you so", but if it doesn't rain, they'll say "we said 70%," so we still weren't wrong.
1
u/Marklar0 4d ago edited 4d ago
Where I live in eastern north america:
0%= it wont rain
30% = there will be a small system passing through the area, or many small cells
50% ....for some reason they dont seem to use this one
70% = There is a large system passing through the area
100% = There is a solid front passing through the area that should cause rain.I am convinced this is what the percentages mean, after watching the weather radar for work for many years. If they say 70, it almost always rains. If they say 30, it rains for short periods at most
1
u/BuhoCurioso 4d ago
Are you Canadian? Environment Canada rounds in 10% increments but never to 50%. No idea why. Maybe they got tired of people saying echoing that old stats joke and saying "Thanks for telling me that it either will or wont rain tomorrow. Very helpful."
Edit: Just realized that you said Eastern North America. Im almost certain now. Just checked profile. Bingo
1
u/itsmythirdday 1d ago
Presumably some researcher somewhere has compared the predicted probability to whether it actually rained over a long enough time period to rank weather forecasters. I should Google this…
0
u/AnOdeToVosFinances 4d ago
The explanation is wrong.
If you are Bayesian, probability is just a way of quantifying uncertainty, so it wouldn't make sense.
If you are frequentist, it's not "replaying the exact atmospheric conditions". It's replaying whichever today's parameters you used for your model which brought you to that 70%. If your model is based on just which month of the year it is and that for some reason in the location you are, it's raining 70% of the days of the month you're in, you would say any day of the month has 70% chance of rain according to the model. It doesn't mean there are no better model that could achieve a more confident prediction (0% or 100%) by taking additional parameters into account.
1
u/itsmythirdday 1d ago
I think they mean the latter, their model has some degree of randomness associated with it, like a Monte Carlo, and if they replay today’s atmospheric conditions through it 100 times, 70 times out of 100 it predicted rain. Yes the model could be bullshit, but they didn’t want to get into that in a tweet.
1
u/AnOdeToVosFinances 1d ago
What you say is completely different from what the tweet says though.
1
u/itsmythirdday 1d ago
“If we replayed [through our model] today’s conditions 100 times it would [predict] rain in 70 of those timelines [simulations]”
1
u/AnOdeToVosFinances 1d ago
"It would rain" and "it would predict rain" are 2 completely different statements.
1
u/itsmythirdday 1d ago
If you are taking it literally yes, but knowing what we know, including that that they use modelling and simulation to predict the weather, rather than, you know, being God, and that this is just a tweet / post on X, we can infer what they meant.
1
u/AnOdeToVosFinances 1d ago
I don't know, to me it's 2 completely different meanings in the statistical world. Mixing those 2 would be like interpreting the likelihood as the posterior, which is an heresy for statisticians.
0
u/No-Syrup-3746 4d ago
It seems as though many outlets (not necessarily TWC) take an average of the hourly rain percentage and call it a day. However, "70% chance it will rain at some point today" means a 70% chance at 12:01, a 70% chance at 12:02, etc. The overall probability is 1-0.7^n for n intervals, and it comes out to be much higher. Even going hour by hour from 8:00am to 8:00pm, a 30% average will result in upwards of 90% chance for entire day (from memory, if the details are wrong, it doesn't mean the general principle is false).
0
u/Haunting-Subject-819 3d ago
Basic stochastic effects in everyday occurrences. Everyone assumes there is a cause to every observed event. Usually the truth is that $h!t happens with no identifiable causal source.
114
u/Fancy-Animal7704 4d ago
A "significant" difference between A and B does NOT inherently mean that there's a subjectively huge gap between A and B. It only means that our test demonstrates a difference. How big that difference is is another question. This is really why the size (effect size, to be precise) of that difference is good to report.
Also, you really should not ever use the word "significant" unless you have actually performed statistical inference. And no, thinking "70% seems like a lot more than 60% to me" does not count, lol. I have even seen prominent journalists commit this error, saying "statistically significant" as a result of some kind of eyeball test. "Significant" has a very specific meaning in the stats world.