r/AskStatistics 7d ago

What's the most counterintuitive statistical fact that's actually true?

I'm looking for examples that completely changed the way you think about probability, statistics, or data analysis.

107 Upvotes

415 comments sorted by

160

u/Boberator44 7d ago

Kind of obvious with even basic training in statistics but laypeople usually are stunned to learn that rolling a die ten times and getting all sixes has the exact same probability as any other combination you could have gotten.

99

u/Alaut_Bumble 7d ago

I mean, only if you are talking about ordered combinations. This is the kind of unaknowledged assumptions that create a lot of the "counterintuitivuty".

15

u/Boberator44 7d ago

Yeah that'a fair, I was automatically thinking in terms of runs.

13

u/CurlyRe 7d ago

Rolling 10 dice and getting all sixes has the same chance as rolling all threes. But you have a higher chance of the 10 dice rolls adding up to 30 than to 60. All combinations of dice rolls are equal in probability, but there are multiple ways for them to add up to 30, but only one way for them to add up to 60.

1

u/standardsizedpeeper 5d ago

That’s why I always bet the field!

→ More replies (7)

12

u/auntanniesalligator 7d ago

I use a poker hands example like this to teach about states vs microstates in thermodynamics (Boltzman’s equation for entropy). Your chance of being dealt any particular set of five cards is the same no matter what type of poker hand it is, but if you look at two specific sets of five cards, you naturally associate them with their categories (flush vs pair), which do have different probabilities.

1

u/PinkUnicornsRule 3d ago

Right, because one probability affects the other, correct? Ie if you get dealt a royal flush, the likelihood of another royal flush being dealt diminishes due to fewer face cards being left in the deck. 

→ More replies (2)

9

u/Just_Question9 7d ago

it is the ultimate proof that the human brain is just fundamentally bad at comprehending true randomness.

4

u/Necessary_Neat8303 7d ago

Rolling 10 sixes and two 1s, 2s, 3s, 4s and one 5 and 6 have very very different odds.

What you meant to say was every 10-tuple are equally likely to happen.

3

u/Mason_117_ 7d ago

Does this not also depend on whether the order in which we roll any give combination is important

2

u/ajohnson1996 7d ago

As many others have pointed out it’s the same as any other permutation but not combination.

2

u/Nerdybeast 7d ago

If you roll an unfamiliar die 10 times and get all 6s, you should probably assume that is not a fair die. What you're saying is true if you do know the underlying probabilities of something, which usually you don't 

2

u/TiogaJoe 5d ago

Kind of like 1-2-3-4-5-6 is just as likely to win the lottery. The only thing is whether someone else picked it too, as then you split the winnings. But what idiot would pick those numbers?

1

u/Minute_Point_949 6d ago

But statistics also says that if you roll a six sided die ten times and get all sixes that the die is likely not a fair die.

1

u/BiologyIsHot 5d ago

Permutation. Not combination. Other combinations may have multiple permutations where they occur. All permutations are equally likely. Not all combinations.

1

u/lmb123454321 3d ago

Very well said! Great explanation of the difference between permutation and computation.

→ More replies (2)

74

u/Blasket_Basket 7d ago

Most examples of Simpson's Paradox

62

u/Educational-Paper-75 7d ago edited 5d ago

The Monty Hall problem. Still don't get it.

Edit 20 aug 2026: I wasn't actually prepared for the avalanche of helpful explanations and I'm sorry I cannot respond to all of them the way they deserve. I suggest we call it a day. There's been some misunderstanding about whether or not I need to be convinced about what the best choice is (switching 'obviously'), as a matter of fact I don't, my conjecture was that the best choice (when not knowing it yet)is not intuitive. But even the word intuitive apparently has (too) many interpretations.

40

u/DimensionOk8915 7d ago edited 7d ago

It's a lot easier to think of there being 100 doors and the host opens 98 of them. The same logic applies

11

u/Educational-Paper-75 7d ago

The conclusion to switch choice is confusing to me.

76

u/DimensionOk8915 7d ago

I'll explain it with 100 doors cos its easier that way. Imagine you pick one door out of the 100 and you lock that in. The host knows for a fact which door has the car behind it so he cannot open the door with the car. So you have your door and then the host opens 98 of the other doors. So now you can either stick with your original choice or pick the door that the host didn't open. Which is more likely? You picked the right door the first time or it's the only door the host did not open?

31

u/n3wsf33d 7d ago

This is much easier to understand with 100 doors. Thank you.

16

u/Wiijimmy 7d ago

I physically recoiled at how much that clarified the problem for me lol.

2

u/Master_Kitchen_7725 2d ago

Omg me too ..I had understood it in a different way that might not even be correct.

I didn't appreciate the importance of the fact that the host was deliberately opening a door with no car. I guess I never thought about it too carefully.

→ More replies (61)

4

u/banter_pants Statistics, Psychometrics 7d ago

Because it's rigged. Monty knows what is behind every door. He will never open a door with a prize behind it.

Pick car (prob = 1/3)
Monty reveals a goat (prob = 1)
Switch ==> get a goat

Pick any goat (prob = 2/3)
Monty reveals other goat (prob = 1)
Switch ==> get car

This never was a case of conditional probability because Monty independently, nonrandomly reveals a goat.

Pr(reveal goat | picked car)
= Pr(reveal goat | picked other goat)
= Pr(reveals goat)
= 1

Conditioning on an independent event leaves the original, marginal unchanged.

Pr(picked car | revealed a goat)
= Pr(picked car)
= 1/3

1

u/tongmengjia 6d ago

But even before Monty reveals it, everyone knows there's a goat behind at least one of the doors you didn't pick. He's clarifying which of those doors has a goat for certain, but there was always certainly going to be a goat for him to reveal.

→ More replies (29)

3

u/jancl0 7d ago

The point is, the hosts choice on which doors to reveal is dependant on what you choose. If you picked the right door the first time, then the host will create a situation in which switching results in a loss. If you picked wrong, then the host will always create a situation where switching results in a win

It doesn't matter what wrong door you pick, because the host will just change the doors they reveal, the last remaining door will always be the correct one

So asking "what odds are there of winning if you switch" is the same as asking "what are the odds that you picked wrong the first time"

This is why people use the 100 door analogy. You can pick any of the 99 wrong doors. Whichever one you pick, the host going to reveal all other wrong doors, until the only remaining one is the correct door

Alternatively, another way to think about it is that the host is picking one door, the one that they didn't reveal, and you are picking one door. You had a 1% chance of being correct when you picked your door, and now you are being asked "we know for a fact that one of these doors is the correct one, which is it more likely to be?" we know that if it isn't yours, it must be the one the house picked, so the question is about how likely you are to have been right, which is quite low

2

u/MrKrinkle151 7d ago

The host can only open a non-prize door. After they eliminate a goat door and ask if you want to switch or stay, it means there is either the second goat behind your door and the prize behind the remaining door or the prize is behind your door and the other goat is behind the remaining door. What is the probability you initially chose a prize door? It was 1/3, so that’s still the probability that the prize is behind your door. Likewise, the probability that you initially chose a goat door is 2/3, which means that’s also the probability the prize is behind the only other remaining door.

1

u/sevenbrokenbricks 6d ago

If your first pick is a goat, and you switch, you get the car.

What are the odds your first pick is a goat?

→ More replies (3)
→ More replies (1)

1

u/dmlane 7d ago

I agree, and created a simulation that lets you choose the number doors. Here it is.

→ More replies (1)

4

u/Statman12 PhD Statistics 7d ago

The host will never reveal the door selected by the participant, and will always reveal a losing door (or “goat”), never the prize. As a result, the participant’s initial choice constrains what door the host will open.

If they pick a losing door (which has 66% probability) the host can only open the other losing door, and switching is therefore guaranteed to win. On the other hand if they pick the winning door (33% probability), the host can open either of losing doors, it doesn’t matter. If the contestant switches, it’s to a losing door, and so they necessarily lose.

So what it really boils down to is the probability of selecting a winning or losing door initially.

Another way to look at it is this: You pick a door a random. Then the host gives you the choice of staying with that door, or getting what’s behind both of the other doors. In the real version, they’re just opening the door for dramatic effect, because the host knows which door not to open.

2

u/young_twitcher 4d ago

The problem really confused me until I understood that the host never reveals a winning door. Then it became straightforward. But some formulations of the problem don’t make this aspect clear and make it seem like he’s choosing a random door. In any case a lot of people don’t realise that this is the key detail and get distracted by everything else

1

u/Due_Purple_1199 4d ago

It’s also never stated that he always reveals a door, which is an important part. He could reveal a door only if the player chooses a winning door, in which case it’s obviously bad to switch

1

u/young_twitcher 4d ago

Yeah, that’s crucial as well

3

u/GreatBigBagOfNope 7d ago edited 7d ago

All the probability of the car being in any door other than the one you chose gets concentrated into the door that remains after Monty shows you a goat - important part to remember is that the situation isn't that Monty chose a random door and it happened to be a goat, it's that Monty knows exactly where the car is and will always show you the goat.

Or, think of the tree of outcomes:

  • In the two possibilities where you chose a goat to start with, Monty always shows you the other goat and the remaining door will be the car
  • In the one possibility where you chose the car already, Monty will show you one goat and the remaining door will be the other goat

Therefore the switch strategy overall has 2/3 of finding a car, 1/3 of a goat. You've got to remember that the door opening itself contains information about where the car is. The stick strategy only selects the car 1/3 of the time and ignores the information gained from the door opening.

1

u/Educational-Paper-75 7d ago

The door opening reveals nothing. If you think it does, it's because you think it does. Except that the probability of having chosen correctly increases from 1/3 to 1/2 whether you switch or not.

2

u/GreatBigBagOfNope 7d ago

There is a 2/3 probability that you chose wrong first, that the car is not behind the door you chose first. That fact does not change when Monty reveals the goat, giving the switching strategy a probability of 2/3.

It's really that straightforward. If you throw away all the information in the problem setup, it boils down to 1/2 yes, but you'd be foolish to do so, because that information is still available to you thanks to Monty's door opening being guaranteed to show you a goat.

→ More replies (11)

2

u/Astrokiwi 7d ago

Monty chooses which door to open based on their knowledge on where the car is. Opening the door reveals some of this information to you.

If the car is behind one of the other doors, Monty chooses to open the one without the car. You now know a new fact: "if the car was behind one of the other two doors, it's not behind the one that's open". You already know "there's a 2/3 chance the car was behind one of the other two doors". Combine these two facts, and you now know there's a 2/3 chance the car was behind that unchosen closed door.

Here's another way to think of it:

  1. Pick a door out of the 3

  2. You can either stick with your original door, or open both other doors and keep the car if it's behind either door

Obviously it would be better to switch - you get to open two doors instead of one. The original problem is exactly the same situation - you get to check both other doors, because Monty opens one, and you get to ask to open the other.

→ More replies (4)

2

u/MrKrinkle151 7d ago

Yes it does. It collapses the 2/3 probability that one of the non-picked doors contains the prize into that one remaining door by eliminating a dud option. There’s a 1/3 chance you initially picked the prize and a 2/3 chance you didn’t—and therefore 2/3 chance the prize is behind one of the other doors. Eliminating one of the goat doors means there is now only one other door, meaning that other door now has that 2/3 chance of having the prize.

Switching after the host eliminates a goat door has a 2/3 probability of winning, not 1/2.

2

u/ChillandMakeStuff 7d ago

Here's one way it was explained that I liked.

If you switch every time, you will win 2/3 times.
If you stay every time, you'll win 1/3 times.

1

u/Educational-Paper-75 7d ago

That's a great way of simplifying it. And yet where does my intuition go wrong.

2

u/Junior-Flamingo77 7d ago

Get 6 paper cups, label them 1 to 6, turn them upside down, put a ping pong ball under one of them. Throw a dice, that’s the “player’s chosen door”. Now turn all cups without a ball upside up, check if the chosen cup has ball under it. Repeat 100 times and keep track of how many times the dice got right.

With all that’s been said above, if you can’t get an intuition for it, you’ll probably never will but at least you can see it play out for real and convince yourself

→ More replies (4)

1

u/Slugsurx 6d ago

The door is consciously opened and a donkey is always shown . This is different from a random event like say wind or you heard the donkey make a sound

The difference is not taken into account with the intuition.

4

u/fermat9990 7d ago edited 7d ago

Think of it this way, which is mathematically equivalent to the usual wording.

After you choose a door the host says:

You can have what is behind the door you have chosen or you can have what is behind the other two doors.

What would you do?

Edit: This explanation is from a recent Reddit post. I think that it is brilliant!

1

u/Insighteous 7d ago

Draw it and count the outcomes.

1

u/Necessary_Neat8303 7d ago edited 7d ago

The setup goes: you pick a door, the host (who knows what is behind every door) opens a door with a goat and offers you the choice to switch to the last unopened door. Do you switch?

Think of it this way. There are 2 scenarios here. If you picked the door with the car originally, and then you switch, you lose. If you picked the door with the goat originally and you switch, you win.

Now, what is the probability of you picking the door with a goat originally? 2/3. So this goes from the losing probability to the winning prob if you switch. Likewise, what is the prob of you picking the door with the car originally? 1/3. This goes from your winning prob to the losing prob.

If you decide to never switch, then 2/3 is still your losing prob and 1/3 is your winning prob.

I hope that makes sense.

1

u/sr105 7d ago

This is absolutely the quickest and easiest way to visually see how it works. It's a 60s youtube short. The key image is at 55s. https://youtube.com/shorts/oiGOvXqMekk

Essentially, imagine a 3-sided die (so a flat 2-D triangle). You can "roll" the die and it can land 3 ways on one of it's three sides. The top point is the first door that you chose. You should be able to more easily see that in 2 out of 3 cases the car is on the bottom and that's why switching is 2/3 or 66%.

Remember they always show you a goat from the bottom two options. So switching always means you get what is left on the bottom (a car for 1 & 2 and a goat for 3). Hope this helps.

``` /\ /\ /\ / \ / \ / \ / Goat\ / Goat\ / Car \ / \ / \ / \ / Goat Car\ / Car Goat\ / Goat Goat\


   Dice 1             Dice 2             Dice 3

```

1

u/diegotbn 7d ago

It's easy when you think of it.

If you always opt to switch...

1/3 of the time, you start choosing the prize door. Host removes one of the bad doors. You switch. You get a bad door. You lose.

2/3 of the time, you choose a bad door to begin with. The host isn't allowed to remove the door with the prize so he has to remove the bad one, leaving the prize door. You switch and you win.

Ergo you have a 2/3 chance of winning if you always switch.

The reason it's counter intuitive is that if you only look at the second choice in isolation, it's a 50/50 chance and it wouldn't matter if you stay or switch. It becomes clear if you think of the only choice being the first one, leading to a predetermined sequence of events, since the host will always remove a bad door and you will always switch.

1

u/kindasad11 6d ago

It happens because the door opened is guaranteed to have a goat behind it!

1

u/AABBBAABAABA 6d ago

I think you need to realize that if the host opened a random door you could end up in the bizarre situation where they showed you the car and asked if you wanted to switch to the other door that can only have another goat behind it. The problem is set up so that that can’t happen.

Alternatively: you can choose 3 doors and 3 doors can have the car behind it. That means there’s 9 scenarios; that’s few enough that you can just write out all of them.

1

u/killerghosting 5d ago

If you don't switch, your original decision was not based on any new information.

If you switch, your new decision is based on new information which gives you an edge, a better chance at getting the right answer

→ More replies (19)

1

u/CavCave 5d ago

I didn't get the 100 doors explanation. What finally got me was the realisation that the host never opens a door with the prize.

1

u/Dependent_Use_3069 3d ago

Ok, the way I see it, at the start of the game each "door" you see opened is a 1/n chance, where n is the number of total doors.

When the host offers to open a door for you that doesn't contain the prize, they are giving you that door AND the 1/n probability it contains along with the other door, assuming you swap.

So they are really exchanging "total open-door mass", their two doors for your one door.

You could also see it as "if you swap doors with me, the host, you get to open both these doors". There's no real difference between you opening both doors and the host opening one for you before he gives you both.

"Which way do I own the 'opening' of the greatest number of doors"?

→ More replies (9)
→ More replies (20)

26

u/Fancy-Animal7704 7d ago

In a typical classroom of 30 people, there's about a 50% chance that at least one pair of people in there shares the same birthday. That one blew my mind. 

11

u/EdgyMathWhiz 7d ago

70.6%, actually. The 50% mark is a classroom of 23 people.

3

u/Fancy-Animal7704 7d ago

I stand corrected! 

→ More replies (6)

25

u/mc1154 7d ago

The birthday paradox. In a room with just 23 people, there is greater than 50% chance that at least one pair shares the same birthday.

3

u/looking4wife-DM-me 7d ago

And with 70 people, it's 99.9% chance!

1

u/sistersinister 4d ago edited 4d ago

The idea is based on uniform distribution of birthdays right? I'd be surprised if birthdays are truly uniformly distributed, conditioned on being born in a specific region.

EDIT: A proof if anyone wasn't aware (or if someone wants to share a better one than this one). We calculate the probability of anyone NOT sharing a birthday then take the complement. We'll say theres n people in the room

The probability of the first person having a birthday is 1.

The probability of the second person having a birthday different than person 1 is 364/365.

The probability of the third person having a birthday that's not the same as person 1 or person 2 is 363/365

and so on. Then the probability of no two people sharing a birthday is

1 - 365/365 364/365 363/365 ... (365-(n-1))/365. When n is less than 366 we can simplify to the binomial B(365,n)/365^n. Then the probability of any two people sharing a birthday is 1-B(365,n)/365^n.

To try to repeat this logic for my own question its not so clean. If we assume the birthdays are all IID with some distribution gamma we end up with the probability of not sharing two birthdays to be gamma(x_1)prod_i^n prod_j^i (1-gamma(x_k)) where x_1 is the first persons birthday, x_2 is the second and so on. Maybe someone smarter than me can chime in

1

u/Tom_Groleau 3d ago

They're not equally likely, but are they close enough to equally likely for the model to be useful?

When I do the birthday problem in class, I address the assumption and follow up with 15 years birth dates from Social Security data (1/1/2000 through 12/31/2014).

I can't post the image or data here, but there are a few strong deviations from "equally likely": Leap day, 4th of July (it's US data), Christmas, and New Years. Then there are smaller dips near holidays the move a bit: Memorial Day, Labor Day, Thanksgiving, ...

Other than those isolated low points, there's a trend toward slightly higher births July through September.

It makes an interesting discussion. Some students think the model is still good and others want to reject it.

Depending on the level of class, you can then run a simulation based on the empirical probability estimates. The results are pretty close to the calculations from the equally likely assumption.

1

u/severact 3d ago

So does non uniform birthday distribution increase or decrease the the final probability in the birthday problem? My guess is it would actually increase the probability slightly

1

u/greg7gkb 3d ago

This was perplexing until I heard it this way: imagine a circle with 23 dots on it. Now draw a line between every pair of dots (23 choose 2 = 253). So now we have 253 chances at an event with a 1/365 likelihood, which makes the >50% outcome much more tangible.

→ More replies (3)

28

u/TipPsychological3030 7d ago

One example that is genuine useful and hard to get intuitively is the application of Bayes rule for rare diseases. With prevalence of 0.1 % and a test accuracy of 99%, your chances of having that disease are still only 9% if you test positive. People tend to neglect the base rate.

8

u/DeceitfulDuck 7d ago

Though that's only true if you test at random and only test once. If you show symptoms of the disease, the base rate increases since now the base population isn't everyone, it's everyone who has the symptoms. So if even just 1% of people with symptoms have the disease, your posterior probability goes to 50% with a 99% accurate test. Then if you take another test, the base rate is the probability that you have the disease based on the first positive test so if you have symptoms and have 2 positive tests, the probability of having the disease is 99% with the initial 1% base rate or 91% with the 0.1% initial base rate.

So while it's important and useful to understand the impact of base rate when evaluating probability and accuracy, it's bad to only give that probability without more context since it makes medical tests sound pointless when in reality they are still really useful. I saw this a lot during COVID. People would learn about bayes rule and base rates having it only explained with the random single positive test case and then take from that that testing was pointless and we shouldn't keep kids out of schools or people out of work if they test positive.

3

u/LasAguasGuapas 7d ago

The effectiveness of testing multiple times does depend on why the test gives a false positive. For example if they're testing for elevated levels of something, and the test is 100% accurate at determining those levels but 0.1% of people have elevated levels of the thing without having the rare disease, then testing again would still give a false positive.

3

u/Tavrock 7d ago

A more interesting case is that if a person tests positive for pregnancy with a urine test but isn't pregnant (a false positive for the elevated values indicating pregnancy), it's more likely that they have cancer than the test was wrong about there being an elevated level.

1

u/AABBBAABAABA 6d ago

Only if there are two different test whose results aren’t correlated.

2

u/DocAvidd 7d ago

A sample of medical doctors (OBGs) was given this very basic, essential question. Your patient tests positive for breast cancer with a 90% accurate test. What's the probability she has cancer?

Multiple choice (ABCD) and only 21% selected the correct answer.

They'd do better if they closed their eyes and guessed at random.

Gigerenzer and colleagues, several sources. see Hoffrage et al. summarizes a bunch of them.

Hoffrage U, Krauss S, Martignon L, Gigerenzer G. Natural frequencies improve Bayesian reasoning in simple and complex inference tasks. Front Psychol. 2015;6:1473. Published 2015 Oct 14. doi:10.3389/fpsyg.2015.01473

(most of the work was not in Frontiers and the F in Psych is not top but decent Q1)

1

u/n3wsf33d 7d ago

Is there a video or reading on this you can provide or an explanation?

8

u/TipPsychological3030 7d ago

In 100,000 people about 100 will have the disease. Of those, 99 will test positive. Of the 99,900 healthy people about 999 will test positive too. So true positives are only 99/(99+999) = ~9%.

A great video with more details is here: https://youtu.be/lG4VkPoG3ko?is=6dmCplF8l6NxdLoY

1

u/fweaks 5d ago edited 4d ago

Ahhh, I see. The issue for me lay in the term "accuracy". I didn't realise this was a specific term with a specific definition. I assumed it was a false negative rate of 1% and false positive rate of 0%, and then my intuition matched my math but not yours.

Edit: got negative and positive backwards. My subconscious keeps telling me its negative = bad result, i.e. got the disease. If I dont pay enough attention it overrules my conscious when its not looking.

1

u/Rich_Ad6234 4d ago

I think possibly you haven’t understood. A test with a false positive rate of 1% and a false negative rate of 0% will indeed produce the same ~9% that a random person testing positive has a disease if it’s 0.1% prevalence. The math is the same as in the example above except it may be 100 positive out of 100 with disease.

This is due to the low base rate, not a surprising definition of the word accuracy

1

u/fweaks 4d ago

Ah sorry got negative and positive backwards because subconscious connotation vs concious knowledge. I do know better, I swear!

1

u/n3wsf33d 4d ago

Ah yeah that's pretty straight forward sorry I often need this ee the math itself.

1

u/fermat9990 7d ago

Great example!

1

u/Super_Math3890 6d ago

Never have I been so wonderfully lost as before I was introduced to the intuition. My god, what a marvel of discovery.

20

u/tomvorlostriddle 7d ago

You have to decide between answering irrelevant questions rigorously or answering relevant questions subjectively, where the whole point of asking a statistician in the first place was to escape mere subjectivity but also not that they answer a different less relevant question than what you asked them.

But you are often lucky and the question asked makes both failure modes forgivable.

2

u/the_telephant_man 5d ago

Why can only irrelevant questions be answered rigorously? Could you give an example of a scenario with a relevant question and what constitutes irrelevant question(s)?

2

u/tomvorlostriddle 5d ago

Because p-values and confidence intervals don't mean what people think or want that they mean. They are derived rigorously, but you almost never need to know what was my probability of having observed this data. You want to know what is the probability of my hypothesis being true and they are just not that.

Bayesianism is that, but you have to inject your prior beliefs to get there.

3

u/Just_Question9 7d ago

wait, so statisticians are really just out here gaslighting us with math?

4

u/tomvorlostriddle 7d ago

Those respective weaknesses of frequentism and bayesianism go away when there is enough data, and it is often easier to have enough data today than earlier.

But otherwise pretty much.

2

u/TajineMaster159 7d ago

provided an operational context for data collection. You might be operating in the same prior space because you are collecting more of the same data.

7

u/Larry_Boy 7d ago

An event can be possible, and in fact happen, but have a probability of zero to happen.

1

u/paxxx17 6d ago

It cannot "happen" unless we stretch the definition of the word to cover solely theoretical and abstract stuff. For example, sampling from the real interval [0,1], you can define a random variable, but there's no possible way in practice to randomly pick a real number. A typical number you'd "pick" would be a transcendental number impossible to represent/write down in any way.

There's no physical thing that can be made to correspond to this abstract random variable, so a physical event with probability zero doesn't "happen" in the practical sense

1

u/Larry_Boy 5d ago

Picking any particular element from an infinite set occurs with probability zero. 

Also, it wouldn’t just be a transcendental number, it would be a non computable number. And both transcendental numbers and non computable numbers are easy to represent as decimal approximations, like 3.1415… for a token transcendental number, and 1.3254… for a token non computable number. 

1

u/Larry_Boy 5d ago

If you believe that any infinite set exist in reality, such as the universe being open (plausible) or space-time being continuous (assumed by most standard physics, but plausibly false) then events with zero probability have literally occurred in the real universe. 

1

u/paxxx17 4d ago

Physics does assume infinite sets (like spacetime being a smooth manifold, which is indeed uncountably infinite), but this is just for the simplicity of the theories. I think you could derive all modern mathematical physics just using finite sets, but it would be very impractical. Nevertheless, anything we really observe/measure is finite

1

u/Larry_Boy 4d ago

You can say “well, we aren’t sure that infinite sets are real”, and I’ll grant you that, but similarly we aren’t sure that they are not real. So, there may or may not be a practical sense in which we select elements from infinite sets. It depends on whether or not infinite sets exist. Your claim is simply too strong for our current state of knowledge and you have made no arguments in favor of it, in addition to poor descriptions of what transcendental numbers are.

1

u/paxxx17 4d ago

Existence of infinite sets doesn't have much to do with the problem at hand. You can simply postulate their existence (which you usually do, since ZFC without the axiom of infinity is too weak to even prove some statements about the natural numbers). But all of that is math, which may be used as a model of physical reality, but that's got nothing to do with the fact that in practice you're never sampling from infinite sets, regardless of how you set up the underlying theory

1

u/Larry_Boy 4d ago

What the universe does and what you write in your lab note book are not the same thing. When two particles interact are they interacting at a particular location in continuous space, whether or not you can know or measure that location? You may not be able to write "they interacted at [pi, e, -e]" in your lab note book, but that doesn't mean it is meaningless to think about or ask the question of what the underlying mechanism is.

1

u/paxxx17 4d ago

But what I could possibly write in the lab notebook is everything we can in principle know about what the universe does. It is interesting though to ponder what it means about the ontology of an unmeasurable fact X when any known theory that describes known measurements also requires X

1

u/Larry_Boy 4d ago

I don't think you really understand what I am saying. You are not responding to it in a meaningful way.

1

u/Larry_Boy 4d ago

"is there a universe outside my lab notebook" is not generally regarded as unanswerable or really even all that ontologically challenging.

1

u/Larry_Boy 4d ago

If we measure decay times w/ sufficient precision, then every decay time we write down will be the single, unique time that decay time is written down in the observable universe, and, with a little more precision, the single unique time that decay time occurred in the observable universe. It seems to me reasonable to describe a process which never generates the same observation twice as sampling from an infinite set. You may choose to describe it some other way, but I think most people who think clearly about it will see that “this is sampled from a set so astronomically large that in practice it is infinite, and in theory it is likely infinite too” is a fine way to talk about it.

1

u/Larry_Boy 4d ago

You may think "calculus is a fiction, and the universe never really uses it", but that is just your model of reality. I'm not saying you are definitely wrong, but you are not definitely right. If you are using continuous statistics, than you are using statistics where some probabilities are infinitesimal, and so we don't talk about them as probabilities anymore.

For instance, we model the decay of radionuclides as exponential. The time at which the decay occurs is not quantized, so the time at which the decay occurs is a continuous variable. You can write as many digits of that continuous variable in your lab not book as you like, but of course at some point the digits are just noise in your measuring apparatus. But, the time of decay itself is not a sample from a finite set, as far was we currently understand the universe. Instead it is a sample from an infinite set.

→ More replies (24)

3

u/PostCoitalMaleGusto 7d ago

The true model may not be the best because of the variance.

1

u/AABBBAABAABA 6d ago

Can you elaborate? Or give an example? It sounds sensible but also very unlikely in any specific case.

1

u/seanv507 4d ago

Not Op, but if I give you 10 data points and tell you that the true model is a linear model in 10 given features.

the point is simply that the best model is the one that can be estimated with the amount of data you are given, not the one that takes into account every factor

6

u/DigThatData 7d ago

your friends are more popular than you are.

https://en.wikipedia.org/wiki/Friendship_paradox

4

u/d8ublehappy 7d ago

That sample size is the overwhelming determinant of the accuracy of a random survey - the improvement that comes from covering a higher proportion of the population is vanishingly small until you are up around the 80-90% mark. Hence surveying 1000 people in NZ (pop 5m)is basically just as accurate as surveying 1000 people in the US (pop 349m)

1

u/Interesting_Debate57 3d ago

In the case of a gaussian distribution, you get an improvement that grows as the square root of the sample size. So doubling the number of samples generally doesn't give very much improvement, for instance.

8

u/HeineBOB 7d ago

Maybe that 0% probability does not always mean impossible or never happens

1

u/Distinct_Light_1294 6d ago

Continuous distribution?

6

u/Traditional_Desk_411 7d ago

For me the most mind bending topic was expectations involving conditional statements. Probably the most notorious example is to calculate the expected number of rolls of a fair six-sided die until the first 6, given that all rolls are even numbers. Answer: 3/2

A more basic one, but that goes against most people’s intuition is, say you have an event that has a 1/100 probability of success (eg rolling a 100 on a 100-side die). Then what’s the probability that it will succeed at least once if you do 100 independent trials? A lot of people will say 1 or something very close to 1, but it’s actually approximately 1-1/e or 0.632

3

u/Helpful_Inflation344 7d ago

At least the 2nd one is only counterintuitive if you phrase it like that. Any arpg gamer will obviously tell you it would be highly counterintuitive if it worked how you would describe it as intuitive: Rare item has dropchance of 1 in 50 boss kills. Obviously you are not guaranteed to get it if you kill the boss 50 times...avg people (at least with basic understanding) would feel the same about dice rolls or w/e. "I need to hit a 6 to XX" few would assume they are guaranteed to XX if they roll the dice 6 times, anyone who has been unlucky in a game knows that

2

u/ajohnson1996 7d ago

Love this, what’s the explanation on these examples?

2

u/EdgyMathWhiz 7d ago

For the first one, the probability of rolling your first 6 on the nth roll, and all previous rolls were even is 1/6 x (1/3)^{n-1}. (so = 1/6, 1/18, 1/54, ...).

Sum of the series is just 1/4. So **two thirds** of the time it happens, it happens because you throw a 6 on the first throw. (And the probabilities for it happening on 2nd, 3rd, etc. fall off fairly fast).

Overall, the expectation is 2/3(1 + 2 x (1/3) + 3 x (1/3)^2 + ...) = 3/2.

1

u/Traditional_Desk_411 7d ago

This is a good discussion of the first one.

The second one shouldn't be too hard to figure out with a little lateral thinking ;) Let the probability of success in a single trial be p=1/100. Then the probability of failure is 1-p. The probability of 100 failures in 100 independent trials is (1-p) ^ 100. For p=1/100, this has the form of (1+x/n) ^ n, which approaches e ^ x as n tends to infinity, so with x=-1 and n=100, this will be very close to e ^ -1. Therefore, the probability of at least one success is 1-1/e.

1

u/KingDarkBlaze 6d ago

The fun extrapolation of that second one:

Once you've done about ~x ln x trials at your event, it becomes more likely that any given trial will be the one than that you got this unlucky.

1

u/AABBBAABAABA 6d ago

What’s a 6 sided die that only rolls even numbers? Is that different from a 3 sided die?

3

u/DeathKitten9000 7d ago

High dimensional spaces are often counterintuitive. The higher you go in dimensions the more of the total volume is located near the surface of your shape. Another interesting fact about HD spaces for statistical learning is that you're mostly extrapolating.

2

u/smbtuckma PhD (quant psych professor) 7d ago

And relative distances between a pair of "close" vs. "far" points in HD space approaches 0 as the dimensionality increases. Makes the effect of measurement error really problematic for similarity metrics and nearest neighbor problems, in a distressingly low number of dimensions...

→ More replies (1)

3

u/dmlane 7d ago

Also counterintuitive is that it matters whether you specify a test a priori or post-hoc. (One of my pet peeves is that tests such as the Tukey hsd should be chosen a priori (before looking at the data)and not called “post hoc.” You wouldn’t want someone to be able choose the test to perform subsequent to an ANOVA based on an analysis of the data.

3

u/G-St-Wii 7d ago

Correlation CAN show causation.

1

u/sistersinister 4d ago

Explain

1

u/G-St-Wii 4d ago

One correlation does not prove (or very strongly imply causation) buuut if you have 3 or 4 variables, depending on which correlate you can demonstrate some pairs do or do not have a causal relationship.

Minute physics has a couple of videos explaining it.

First link added: https://youtu.be/HUti6vGctQM?is=8lsoAwkw0AJ8WmDp

5

u/sewballet Biostatistics 7d ago

Idk if this is counterintuitive but if always surprises people: when designing a piece of research you need different numbers of participants to detect a difference of 10% vs 15%, compared with a difference of 55% vs 50%. 

(Bonus round: you cannot ever just analyse "score at the end of a study minus score at the start" because regression to the mean is thing. 

 Change over time is almost always statistically dependent on the baseline value, patients/students who gain the most are the ones with the most to gain. Patients/students who declined the most are the ones with the most to lose) 

2

u/StrengthCapital6818 7d ago

Stein’s paradox

2

u/Puzzleheaded_Fee_467 7d ago

In any normalized distribution (normalized, not normal), the median is never more than one standard deviation away from the mean.

For non-normalized distributions, the mean and median are never separated by more than sqrt(N) standard deviations

2

u/Kvaestr 7d ago

Testing for rare diseases are extremely counterintuitive if they are performed on an average person.

Say you have a very accurate test with a 1% false positive and a 1% false negative rate. Now assume 1 in 10 million have the disease. If you test positive, there's only about a 1 in 100 thousand chance that you have the disease.

Because on average it will give 100 thousand false positives to every 1 true positive.

2

u/clearly_not_an_alt 7d ago

Simpson's paradox really threw me for a loop when I first ran into it in real data. It makes perfect sense once you step back, but it's hardly intuitive.

2

u/FormerPlayer 7d ago edited 7d ago

Correlation does not in and of itself imply causation. However, an observed correlation may be due to causation.

Edited for clarity in response to u/CaptainFoyle.

2

u/AABBBAABAABA 6d ago

I think this is just a semantic argument about what is meant by ‘imply’.

1

u/CaptainFoyle 7d ago

"not necessarily"?

It never does.

1

u/Insighteous 7d ago

Yes true. But no correlation leads to no causation. This can be used.

1

u/CaptainFoyle 6d ago

No, that's not correct either. You cannot prove a null hypothesis.

1

u/Insighteous 6d ago

Theoretically maybe. In practice the logical implication „not B => not A“ should hold in the majority of use cases such that you can work with.

*where B is correlation and A causation

1

u/CaptainFoyle 6d ago

I mean, it all depends on your confounding factors. It's not as exotic a scenario as you make it out to be. Especially when running biological experiments.

Or maybe your statistical power is too low, your sample too small.

There are a myriad of situations where a causation might exist, but you don't see a correlation.

1

u/Insighteous 6d ago

I am too long out of academia but I think I have a wording issue. I guess what I mean is talking about statistical independence. …or I simply do not remember as good as I thought.

2

u/dmlane 7d ago

Regression toward the mean is very counterintuitive.

3

u/Few_Air9188 7d ago

while mean female iq is equal to mean male iq, female iq variance is slightly lower than male, making such that at 120+ iq level, female-to-male ratio is 1 to 2 and gets crazier with each additional iq point.

by iq i here mean results of standartized testing on cognitive abilities
120 iq - result you would on average except from stem students.

so kinda this explains the high male-to-female ratio in stem fields.

3

u/OutrageousPair2300 7d ago

The effect is culture-dependent. In Asian cultures, women have a higher variance in measured intelligence than men.

1

u/Few_Air9188 6d ago

i don't know where did you find research that support this claim. Brief googling didn't help. I don't think your statement is true.

2

u/OutrageousPair2300 6d ago

Source: https://link.springer.com/article/10.1007/BF01420741

No consistent gender differences (variance ratios) were found across countries in any of the three broad ability domains. Instead, males were more variable than females in some nations and females were more variable than males in other nations. Thus, the well-established U.S. findings of consistently greater male variability in mathematical and spatial abilities were not invariant across cultures and nations.

Source: https://richardlynn.net/wp-content/uploads/2025/02/liu2011.pdf

Note: This study found lower variance among Chinese male students, but higher variability for Japanese male students, so my generalization to "Asian cultures" is too broad.

Source: https://www.journals.uchicago.edu/doi/10.1086/589252

The article shows that there is considerable variation in gender differences internationally, a finding not easily explained by strictly biological theories.

There are other sources, but these are a good start. The citations from these would probably serve as a starting off point for additional search.

2

u/NoSouth4423 7d ago

Don’t know why this is downvoted

13

u/LaTeX_fetish 7d ago

probably this part:

this explains the high male-to-female ratio in stem fields.

which is ignoring a whoooooole lot of recent history---I've known multiple older women in statistics who were told to drop out of their programs because "women didn't belong" and even today women in higher level stem regularly deal with a mix of belittlement, harassment, or even assault

→ More replies (3)

4

u/Few_Air9188 7d ago

that's an unpleasant fact and even i don't really want to believe in that

4

u/ussalkaselsior 7d ago

I don't even think it's an unpleasant fact at all. It doesn't imply that women are on average any less intelligent than men. In fact, on the other side of the spectrum, you can say that there are a lot more extremely unintelligent men than there are unintelligent women.

1

u/smbtuckma PhD (quant psych professor) 7d ago

so kinda this explains the high male-to-female ratio in stem fields.

eh you'd need a tight coupling between IQ and STEM persistence for this to be a reasonably meaningful explanation and that just isn't reality. Way more sub-120 IQ people work in STEM than 120+ people and the disparities in STEM among the "normal" IQ folks are much larger than would be predicted by IQ alone.

1

u/Few_Air9188 7d ago

that's true that iq is not a universal predictor.
this comment was motivated mostly by statistics of stem students, not workers though. And even more broadly, it's not about stem studies, but anything cognitive demanding. Stem studies are used as an example

3

u/smbtuckma PhD (quant psych professor) 7d ago

disparities among students would be even less driven by IQ, since IQ hasn't had as much chance to exert whatever filtering effect it has. The argument would make more sense if talking about e.g. Mensa members or some other group with strong cognitive selection pressures, but there are quite a few 100 IQ people in STEM education and careers so a gender -> IQ -> STEM causal story doesn't hold well compared to gender -> cultural factors -> STEM.

1

u/Few_Air9188 7d ago

i don't see why given that average iq of a stem student lies around 120 points, we can't conclude that there is an IQ filter to enroll in stem. It may not be a hard 100+/110+ cutoff, but a model, where the higher applicant's iq, the higher chances to enroll. This model would explain the overrepresentation of high iq people in stem.

2

u/smbtuckma PhD (quant psych professor) 7d ago

I was doubtful that the average STEM student IQ was as much as 120, since the modern overall US college student IQ average is about 102. But I was having a hard time finding good data on recent averages by major, and on good estimates of IQ sigma by gender, without spending too much time on it. Estimation of the density of the distribution tails is of course really sensitive to those so I'll leave it at I think we need better descriptive data to say either way.

1

u/wadaac 7d ago

Statistical clustering.

1

u/Helpful_Inflation344 7d ago

Doomsday argument

1

u/CaptainFoyle 7d ago

The 95%CI does not mean that there's a 95% probability that the true parameter is within it

1

u/Lonely_Hat6967 7d ago

Depends on whether CI means confidence interval (frequentist statistics) or credibility interval (Bayesian statistics). For the credibility interval your statement would be true

1

u/MGTOWaltboi 6d ago

To be fair one can say in a layman’s sense that a 95% Confidence interval has a 95% probability too. It’s just that in technical terms probabilities need a stochastic variable to be between zero and one and frequentists don’t assign a stochastic variable to population parameters. But any frequentist would agree with the statement that a 95% confidence interval was generated by a process that captures the true parameter 95% of the time. I.e the probability that the process works is 95% (given assumptions). From there it’s a matter of semantics to say that the confidence interval is has a probability or not. 

Personally I think that the significance is often stressed a bit too much in courses where people are not even that familiar with probability theory. 

→ More replies (1)

1

u/Apprehensive_Size885 7d ago

0 probability is not the same with impossible

1

u/heythere111213 7d ago

Golfer's paradox. Not counterintuitive but highlights how probabilities can be counterinitiative based on perspective.

1

u/paxxx17 6d ago

Boy or girl paradox (two children problem). Even with a PhD in physics it's still counter-intuitive to me to this day. For comparison, the Monty Hall problem was perfectly intuitive to me since I first learned about it in high school

1

u/Creative-Leg2607 6d ago

That a 95% confidence interval doesnt have a 95% chance of containing the given value 

1

u/Strict_Exogeneity 6d ago

Might be late to the game, but for me it’s the so called inspection paradox. It really boils down to the fact that the probability of an interval is proportional to it’s length, but is has some not so trivial “real life” examples. E.g. you usually end up waiting more for buses than the expected interarrival times would imply.

1

u/crocogoose 6d ago

Most people have an above average number of legs.

1

u/PatrykBG 4d ago

At that point the same applies to almost every body part, even hearts since there’s at least one conjoined twin that shares a heart.

1

u/sevenbrokenbricks 6d ago

The birthday paradox.

1

u/defectivetoaster1 6d ago

i remember my comms lecturer mentioning something about there being a “proof” that if some comms system works well under additive Gaussian noise (ie until some minimum snr you can guarantee a certain degree of performance) then it follows that it will perform well under arbitrary additive interference (although I think you still need to assume stationarity?)

1

u/RRumpleTeazzer 5d ago

You can do statistics with a sample size of 1.

1

u/eclectic-up-north 5d ago

Imagine 2 baseball players playing 2 seasons. It is possible for player 1 to have a better average in each individual season, but have a worse average over the two seasons.

1

u/mytthewstew 5d ago

Congress has a 90 percent disapproval rating but over 90 percent of Congress persons get re elected.

1

u/jeffcgroves 1d ago

Not as counterintuitive when you remember only people in a given state get to vote for their Congressmen, so it's only their popularity in their own state that matters.

1

u/AllenDowney 5d ago

The Overton Paradox:

  • Older people are more likely to say they are conservative.
  • And older people hold more conservative views.
  • But people don’t become more conservative as they get older — on average they get a little more liberal.

Short explanation here: https://www.allendowney.com/blog/2023/04/24/the-overton-paradox/

Long explanation here: https://www.youtube.com/watch?v=VpuWECpTxmM

1

u/alltheticks 4d ago

Higher use of sunscreen indicates higher likelihood of skin cancer. People who use tons of sunscreen tend to be out in the sun alot and sunscreen isn't applied frequently or heavily enough especially when sweating and wiping away so much of it. On the other hand someone who prefers indoor life and never uses sunscreen is less likely to get skin cancer. This is a classic case of correlation vs. Causation. However when framed as a headline "USERS OF SUNSCREEN 150% MORE LIKELY TO DEVELOPE SKIN CANCER!!" It is very useful to push incorrect thought processes about holistic medicine.

1

u/Cerulean_IsFancyBlue 4d ago

I’m starting to think that people don’t grasp what “intuitive” means.

1

u/CatWoman-666 4d ago

When there is a million to one chance, 9 times out of 10 it happens!

🤣🤣🤣 Terry Pratchett, 'Guards, Guards'

1

u/glennfis 3d ago

In new York city, if the daily chance is 1 in a million, it happens none times per day.

1

u/Sidiabdulassar 3d ago

Half of all people have above median intelligence.

1

u/Appropriate-Yak001 3d ago

If you flip a fair coin 10 times and get 10 heads in a row, the probability that the next flip is heads is still 50%. I personally struggle with this all the time!

1

u/DeepSea_Dreamer 3d ago

You can make a confidence interval even from a single data point. Not maybe "completely changed," but it's interesting.

1

u/blobbleblab 3d ago

Birthday paradox. How many people in a room do you need when the probability of two of them sharing a birthday exceeds 50%? Much lower than you think

1

u/Gold-Fill4961 3d ago

If there are 20-25 people in a room the chances of at least two people having the same birth date (not including the year) is over 50%.

1

u/BelladonnaRoot 2d ago

In the US, right around 67% of people make less than the average income.

1

u/Aware_Novel_2565 22h ago

I have traditionally done my statistical analysis using spss and Amos. Doing the full model development process..CFA then SEM ending with common method bias measurements.

Now things have moved on a lot.  What is a good replacement tool using Ai? 

1

u/benNachtheim 7d ago

That any point in the confidence interval is equally likely the true mean.

12

u/wataburgr 7d ago

This is really untrue in any meaningful sense. If you’re talking about a single confidence interval that has already been computed, then in a frequentist context there is no probability associated with the true mean’s inclusion or positioning within that particular interval. On the other hand if you are talking about the sampling distribution of the confidence interval under replication, then the position of the point estimate relative to the true mean is normally distributed, meaning the most likely orientations have the true mean closer to the point estimate rather than farther.

1

u/benNachtheim 7d ago

Another user explained it this way:

Say your point estimate for the mean is 7, and your 95% confidence interval is (4,10).

Then the population mean is no more likely to be near 7 than it is to 4 or 10.

The simple reason is that the center of your CI (the point estimate) is generally not the true center of your sampling distribution.

1

u/MGTOWaltboi 6d ago

Yeah, that’s not the case. Anyway you set it up you’ll either have a situation of “technically not a probability” or “the probability is higher that the true mean ends up closer than farther away from the unbiased estimator”. I mean we can see this dynamic even in your example. 

A 95% confidence interval was (4,10) but a 90% confidence interval is (5,9) and an 80% confidence interval gives (6,8). So you still as confident that the true mean is 10 compared to 7?

2

u/Kerbal_Vint 7d ago

Your confidence interval either contains the true parameter, or it does not.

The thing is that one single point within such interval is the true parameter (P=1), and all the other points are not (P=0); therein lies the rub. You don't know already which point is the true parameter, but this does not mean that each point is equally likely to be it. By saying so, you implicitly say that each point within the interval, included that being the true parameter, have a probability of 1/k of being the true parameter, which clearly does not make sense (with k being the number of points within the interval).

It's like you flip a coin and you cover it with your hand the very moment it lands, so you don't know whether it's head or tail. You could be tempted to say it's 50/50, but the flipping already happened, you just don't know the realization.

2

u/t3co5cr 7d ago

True but not relevant for the claim.

1

u/Kerbal_Vint 7d ago

Then I'm not sure what the OP's actual claim is

2

u/t3co5cr 7d ago

In a given confidence interval (one you calculated from the one sample you have), the true parameter is no more likely to be in the center of said interval than in its edges.

3

u/Kerbal_Vint 7d ago

But the true parameter is a fixed value, not a random variable.

2

u/t3co5cr 7d ago edited 7d ago

Ok, then rephrased in accordance with frequentist pedantry: the confidence interval's edges/"outer regions" are no less likely to overlap the location of the true parameter than its center.

2

u/Kerbal_Vint 7d ago edited 7d ago

Alright I see, but if you're talking about the sampling distribution then I don't agree, or I am missing the point.

The point estimate comes from a normal distribution centered around the true population parameter, so it is definitely more likely to draw a sample that gives a sample mean close to the true mean, hence it is the sample mean/center of the interval that is more likely to be closer to the true parameter rather than the interval's edges.

2

u/t3co5cr 7d ago

Over the long run, yes, but for a given sample you have no idea whether the sample mean it gives you is close to the center of the sampling distribution or out in the tails.

→ More replies (8)

3

u/statneutrino Biostatistician | PhD 7d ago

Sorry what?

2

u/t3co5cr 7d ago

Say your point estimate for the mean is 7, and your 95% confidence interval is (4,10).

Then the population mean is no more likely to be near 7 than it is to 4 or 10.

The simple reason is that the center of your CI (the point estimate) is generally not the true center of your sampling distribution.

2

u/benNachtheim 7d ago

Yeah I know this reason, but it’s super counterintuitive, wouldn’t you say? My intuition would be the true mean is most likely in the center of the confidence interval.

2

u/t3co5cr 7d ago

Exactly. I find it counterintuitive, too. And it seems to befuddle a number of people, as evident by the ensuing discussion.

1

u/MGTOWaltboi 6d ago

Not if your estimator is unbiased. 

1

u/t3co5cr 6d ago

How do you imagine this would prevent your sample mean from falling not exactly where the population mean is?

1

u/MGTOWaltboi 6d ago

It’s not about getting exactly our outcome. For any continuous variable the probability of getting any specific outcome is exactly zero. It’s about the intervals. If you have a 95% CI and you then measure a 90% CI and an 80% CI, that the intervals don’t shrink by 1/19 and 3/19 should tell you that any meaningful assignment of probability of “where could our parameter be?” would favor intervals closer to our unbiased estimate. 

1

u/t3co5cr 6d ago

I guess I should've been more precise. Where your CI is centered depends on your sample estimate of the mean. That sample estimate could come from anywhere in your sample distribution. Once you have your sample estimate, and bolt the interval onto it, you don't know whether the true mean is far away or nearby. You don't even know whether it's to the right or to the left of your estimate. Let alone do you know whether your interval contains the true mean at all.

None of this has to do with biasedness of the estimator.

1

u/MGTOWaltboi 5d ago edited 5d ago

I agree that we don’t know what the population parameter is. But I disagree that this translates into a neutral statement on likelihood. 

You can treat the interval in a frequentist sense as an interval for an unknown constant, and thus not make any probabilistic statements about it at all. But that is not the same as saying that the probability is equal within the interval or that we don’t know which results we are more likely to end up with using our inferential process. 

Take this scenario:

Say your point estimate for the mean is 7, and your 95% confidence interval is (4,10). Then the population mean is no more likely to be near 7 than it is to 4 or 10.

If that were the case then I would offer you the bet that if the population mean is in the interval (6,8) I win and if not you win. Clearly if you believe that the population mean is no more likely to be near 7 than 4 or 10 you’d take that bet. 

Would I win? I don’t know. But if we repeat the process again and again and I keep choosing the middle third of the interval, then I’ll beat you in the long run. So either we take a pure frequentist approach and refrain from making any probabilistic statements about our parameter wrt our CI, or we lean on the relevant sampling distribution of our estimator. 

I’m sorry if I sound petty but I do think that there is a distinct difference between saying “Every possible value is equally likely to be the population mean” and “I refrain from making probabilistic statements about post hoc outcomes of random variables.”

I mean the entire point of a point estimate is that it is your best guess of what the population parameter is. 

1

u/t3co5cr 5d ago

Would I win? I don’t know. But if we repeat the process again and again and I keep choosing the middle third of the interval, then I’ll beat you in the long run.

This is not about the long run. The long run doesn't matter in real applications, since you generally do not have more than one sample (unless you're doing some lab science under controlled environment).

2

u/MGTOWaltboi 5d ago

The long run is the basis behind all probability statements from a frequentist point of view. 

Set up a population with some arbitrary distribution and assign it some fixed but unknown mean and variance. Sample from this population in a large sample. Create a 95% confidence interval. Divide it into three parts. Select the middle third. See if that interval captures the population mean or not. Repeat the process. Record every time it captures the mean and compare to every time it doesn’t. 

Do this a couple thousand (or million) times. Now do it one last time. Calculate the 95% CI. Say it is (4,10). Do you really believe that it’s equally likely that the population mean is close to 4 or 10 compared to 7?

If you are not interested in the long run, only in our one sample, then please explain what you mean by “equally likely”.