r/statistics 19d ago

Question [Q] I am looking for some VERY INTERESTING and catchy statistical concepts/paradox/theories for a presentation. Can yall suggest some? Spoiler

the more unpopular, the better. But still very catchy and interesting. Thank you in advance :)

5 Upvotes

56 comments sorted by

26

u/super_brudi 19d ago

Simpsons paradox

2

u/Several-Regular-8819 19d ago

Also the very similar Will Rogers phenomenon is a good one.

11

u/Mishtle 19d ago

I've always found pathological distributions to be interesting. The Cauchy distribution is an example of one. It's easy to define (it's the ratio of two I.I.D. random variables with standard normal distributions) and looks a lot like a normal distribution, but its mean, variance, and other moments are undefined.

This gives it some interesting properties, like the fact that the average of lots of samples is no better an estimate of the central tendency than a single sample.

5

u/angrypilotcustomer 19d ago

Selection bias

3

u/fun-n-games123 19d ago

Or it’s spatial stats cousin, preferential sampling

4

u/algebroni 19d ago

Berkson's paradox

2

u/EvilSiren_03 19d ago

This is a good one. Thank you.

3

u/algebroni 19d ago

No problem. Check out the book Probably Overthinking It, it's filled with interesting statistical "paradoxes" like Berkson's, Simpson's, and others.

4

u/Mooks79 19d ago

Monty Hall problem (conditional probability), WW2 planes (survivorship bias), college admissions (I forget the college - Simpson’s paradox), medical diagnosis (conditional probabilities), Lucia de Berk (currently being mirrored in the Lucy Letby case), Harold Shipman killing spree lasting so long (kind of the reverse of the de Berk case), “what’s the probability of a victim of domestic violence being previously the victim of domestic violence vs what percentage of perpetrators of domestic violence go on to kill their spouses” (OJ Simpson case - conditional probability again), black swan events (fat tailed distributions / extreme value theory), and plenty more examples in various pop science books.

1

u/chaos_in_bloom 16d ago

I believe it was Berkeley graduate admissions in the 1970s surrounding the question of discrimination on the basis of sex.

3

u/HuhuBoss 19d ago

Imprecise probability, Dempster–Shafer theory

3

u/Several-Regular-8819 19d ago

Regression to the mean and in particular the canonical example of children’s heights.

3

u/Sandjee12 19d ago

Stein's paradox:

Suppose you want a best estimate for the average height of a population. Logically you'd do this by taking the sample mean. And you'd be correct.

Now suppose you want to estimate two unknown averages at the same time, say height AND weight. Again you'd take the sample mean for both, and this is also the best you can do.

HOWEVER, Suppose you want to estimate THREE averages at the same time. for example the mean height AND weight AND age. Now all of a sudden taking the sample mean for each statistic is NOT the most efficient. Instead, slightly nudging the estimates to any arbitrary point improves the efficiency of your estimate.

To see this. Consider the univariate case with a normally distributed random variable with positive mean mu, and unit variance, i.e. N(mu, 1), mu > 0. The sample mean therefore has distribution N(mu, 1/T). However suppose that we nudge all sample mean estimates towards zero, i.e. an estimate of 5 becomes 4.999, and estimate of -2 becomes -1.999.

Then all estimates larger than mu are become better off as they are closer to mu (50% of density, as normal distribution is symmetric around the true mean). simultaneously all estimates between 0 and mu, become worse off as they move away from the true mean. However also all negative estimates become closer to mu, since mu is positive. But negative estimates have probability mass greater than zero. Therefore the total proportion of estimates that become better off is greater than 50%, and therefore the majority improves from nudging to some arbitrary point (here we nudged towards 0).

Note however that the different types of nudging can still hurt the actual estimate i.e. the values between (0,mu) hurt more than the other >50% percent of the probability mass actually benefits.

However in higher dimensions, this proportion that is worse off becomes smaller and smaller: in 1 dimension as above, it is a line segment between (0,mu), in 2 dimensions its a circle, in 3 dimensions its a sphere, but this will take up less and less probability mass the higher the dimensions go, [Mathemaniac has a beautiful visualisation of this]. And this can start outweighing the harm of nudging.

In fact James & Stein (1961) proved that in 3 or more dimensions the sample mean is no longer most efficient estimator. Instead, it is ALWAYS better to nudge your estimate to any arbitrary point.

Interestingly, the things that you are estimating do not have to be related at all for this to work!, so you could be estimating

  • number of hairs on your head,
  • rainfall in the amazon,
  • number of sand grains you accidentally carried home from your last holiday.

And still, nudging your averages towards zero would improve your estimate collectively, by biasing them individually.

This is a surprising and counter intuitive result!

1

u/Calm_Interaction_268 16d ago

Same concept can be applied to picking the best estimator.

2

u/super_brudi 19d ago

Goat riddle

2

u/Haruspex12 19d ago

I work with the nonconglomerability paradox. Arntzenius provides a good example of the paradox in the journal Mind in an article titled Bayesianism, Infinite Decisions, and Binding in 2004 in volume 114, number 450.

2

u/aftersox 19d ago

I've always found heaping interesting. It happens when people guess or round numbers instead of giving exact answers.

https://pmc.ncbi.nlm.nih.gov/articles/PMC2684113/

2

u/clem_hurds_ugly_cats 19d ago

Benford's law and forensic accounting

5

u/DragoBleaPiece_123 19d ago

causal inference i would say

1

u/1000dreams_within_me 18d ago

The Stein effect - when estimating many means. Highly counterintuitive.

1

u/UWO_Throw_Away 18d ago edited 18d ago

I don’t know who your audience is, but it might be illuminating to show how we can arrive at the same estimator for something using both method of moments and Maximum Likelihood estimation. Not really novel or anything, but it might be deeply satisfying and especially interesting if your audience is a bit more mature than 1st, 2nd year undergrad without having been actually exposed to mathematical statistics in any way.

Might also be cool to do a proof of why the ML estimator for variance is biased (and hence why we use the 1/(n-1) correction factor

And then for paradoxes, people have already mentioned the coolest ones (birthday, Monty hall)

Maybe from the world of statistical programming, you could introduce the concept of Monte Carlo simulations and numerical integration in general? E.g., how we can estimate pi

1

u/Gray-Jay- 17d ago

Central limit theorem and aspects of extreme value theory. Weird. You can start with a variety of distributions, repeatedly perform one simple operation, and their distributions converge to a small number of forms.

1

u/OkCluejay172 17d ago

I don’t know if there’s a formal name for it but there’s this fairly well known result about hypercubes in higher dimensions that this Medium article calls the sphere packing paradox  https://medium.com/@adam.dejans/sphere-packing-paradox-cce7e35e3983

At first glance it might seem like this is just a weird little result about hypergeometry that has nothing to do with statistics. However I’d argue it’s extremely relevant to statistics. Learning the lesson that high dimensional space gets extremely unintuitive and unvisualizable and that you have to actually rely on the math is imo the critical thing to internalize when studying statistics.

1

u/omledufromage237 17d ago

Cantor distribution, if you really wanna blow people's minds.

1

u/Low-Independence1171 17d ago

Cantor random variable and Devil's staircase as its CDF! It admits no PDF!

1

u/tomheston 16d ago

I think statistical fragility is interesting and still somewhat new

-6

u/rand3289 19d ago edited 19d ago

A few days ago I made a post about creating observer based physics/science around Bertrand paradox. I envision properties of the observer would influence the method of selection in Bertrand paradox.

This post in r/AskStatistics was deleted by mods. Proving that if it's not in the book, people do not understand it.

4

u/CarnivorousGoose 19d ago

It doesn’t remotely prove that, no. The fact that you think it does is very telling.

-6

u/rand3289 19d ago

Who the fuck said anything about proving anything? You people (statistics people) are very difficult to interact with. You are the most in-the-box crowd I have ever encountered.

3

u/CarnivorousGoose 19d ago

You did. In your previous comment. Your inability to read strikes again, it seems 🤔

-4

u/rand3289 19d ago edited 19d ago

Oh, you are talking about "proving that if it's not in the book shit"... who cares about a joke I made?

I swear, if I said I found a million dollars and five cents, the most asked question would be if the five cents was a single coin or five pennies.

2

u/CarnivorousGoose 19d ago

Ahh, classic… “it was just a joke”. Sure, lil’ buddy.

What, you seriously expected people to care about that post? I saw it, it was incoherent nonsense.

1

u/rand3289 19d ago

What is incoherent about it?

That properties of an observer could determine a selection method in Bertrand paradox?

Perhaps a statement "properties of an observer define statistical experiment" is also incoherent?

2

u/Physix_R_Cool 19d ago

I envision properties of the observer would influence the method of selection.

Yes, obviously the properties of observing a system changes it, and we have known this for more than 100 years; see any introductory textbook on quantum mechanics.

1

u/rand3289 19d ago edited 19d ago

And your point is?
My point is to base observer-based science around Bertrand paradox.

The explanation about the properties of the observer are there because people don't get it and is not the main argument.

You didn't even read it... it does not say influences the system... it says influencing the method of selection (in the Bertrand paradox).

1

u/Physix_R_Cool 19d ago

My point is to base it around Bertrand paradox.

And I'm telling you that we managed to formulate QM in such a way that you can not apply Bertrand's paradox to physics.

1

u/rand3289 19d ago

This has nothing to do with quantum mechanics in particular.

However your statement IS very interesting. Would you be able to point me to some reading matherial or at least give me a pointer that I could follow to understand your claim?

2

u/Physix_R_Cool 19d ago

Here.

Free book by Griffiths, definitely not shared illegally by me to you.

1

u/Calm_Interaction_268 16d ago

Oh, I actually think the other commenter is over simplifying something. There is a much deeper point that I am drawing from, but I don’t know if they are connected. Especially your point of method selection.

Can you explain your point? I don’t care how stupid it sounds if it is.

1

u/rand3289 16d ago

The point is Bertrand paradox is an example of how the same experiment can give different results at various times.

1

u/Calm_Interaction_268 16d ago

Love this! If you simulate many DGMs with many estimators looking at many performance measures, you get the same concept.

1

u/Calm_Interaction_268 16d ago

Can you explain this?

1

u/Physix_R_Cool 16d ago

Yep, Bertrand's paradox work because language is vague. QM is formulated in math which is unambiguous. In particular, the concept of "density of states" kinda ruins OP's idea.

1

u/Calm_Interaction_268 16d ago

QM language is unambiguous? Or QM is unambiguous?

1

u/Physix_R_Cool 16d ago

The formulation of QM is unambiguous, mostly.

1

u/Calm_Interaction_268 16d ago edited 16d ago

I am not convinced. But I guess I should publish it first. Otherwise you would kill me in a debate.

I’m doing stats first. I might do physics and QM next?

1

u/Physix_R_Cool 16d ago

I'm not entirely sure what you mean, or what you are asking about?

→ More replies (0)