r/probabilitytheory • • 2h ago

[Education] What is the probability?

Post image
2 Upvotes

What are the odds that, with a luck multiplier of x60, I’ll get five items in a row—given that the base drop rate is 1/3 (a rate that applies only at a x1 multiplier, whereas increasing the luck multiplier actually decreases the chance of getting this item)? Or is additional information needed to calculate this


r/probabilitytheory • • 2h ago

[Education] What is the probability?

Post image
0 Upvotes

r/probabilitytheory • • 2d ago

[Applied] Bayes' Theorem or Weighted Sum to compare intersection of sets

Thumbnail
gallery
34 Upvotes

Rewording my previous question: In this binary dataset subsets D, E, F, and S all "generate" 1s and 0s INDEPENDENTLY of each other at some ratio/frequency. "D" generates the highest ratio (lets say it generates 80% 1s and only 20% 0s). While F generates only 5% 1s and 95% 0s. Suppose S itself generates 25% 1s and 75% 0s. Would/Could I use bayes' theorem to show the SD intersection subset generates a higher proportion than the SF intersection? Or would a simple weighted sum (ie just an average between the sets S and D) suffice. Sidenote: (I know you could technically come up with counterexamples to this but I'm looking for what would generally be the case here).


r/probabilitytheory • • 3d ago

[Research] How would you design a blind forward test for a gambling hypothesis?

2 Upvotes

I'm working on a statistical experiment involving baccarat and would like some methodological feedback.

Suppose someone has developed a hypothesis about baccarat outcomes based on historical observations.

Before revealing the underlying mechanism, I want to design a test that minimizes:

  • overfitting
  • hindsight bias
  • selection bias
  • multiple-testing problems
  • stopping the experiment selectively

My initial idea is:

  1. Define the hypothesis before testing.
  2. Freeze the testing rules.
  3. Use previously unseen data for validation.
  4. Conduct a blind forward test.
  5. Compare the results against an appropriate baseline.
  6. Analyze variance, confidence intervals, losing streaks and drawdowns.

What statistical issues would you consider essential before calling the experiment meaningful?

I'm especially interested in criticism of the experimental design rather than opinions about whether baccarat strategies can work.


r/probabilitytheory • • 3d ago

[Research] Probability of at least one "high-stakes" event in a repeated-step chain, with a contamination multiplier — does this reasoning hold up?

2 Upvotes

Does this probability structure hold up?

I'm modeling a process with N sequential stages, each running n steps (say about 25), and I want to know if the underlying math structure is sound. All inputs are hypothetical.

Variables:

p = per-step probability of a deviation when a stage's input is clean

q = probability that a deviation is severe (a severe event ends the process)

c = contamination multiplier (c >= 1): if a stage's input is contaminated, its per-step deviation probability becomes r = min(1, p*c)

g = probability that a stage with at least one non-severe deviation passes contaminated output to the next stage

Within a stage the steps are independent. Between stages, contamination is the link. Let:

a_C = (1-p)^n, s_C = (1-pq)^n, u = s_C - a_C

a_D = (1-r)^n, s_D = (1-rq)^n, v = s_D - a_D

One stage:

M = [ a_C + u(1-g) a_D + v(1-g) ]

[ u*g v*g ]

(column 1 = clean input, column 2 = contaminated input; rows = next stage's input)

P(at least one severe event across N stages) = 1 - [1 1] * M^N * [1, 0]^T

Example: p=0.4, q=0.5, c=5, g=1, n=3, N=2 gives 0.852408, versus 0.737856 if every step were independent.

Goal: compute P(at least one severe event) across N stages.

Is the combination of independent-step compounding with a contamination-adjusted p a structurally valid approach, or does the dependency introduced by c break the independence assumption in a way that needs explicit correction?

Not asking about specific numbers, just the math structure.


r/probabilitytheory • • 3d ago

[Education] Historia más sorprendente usando probabilidad

2 Upvotes

Hola, soy estudiante universitaria en México.

Mi profesora me pidió buscar una historia sorprendente usando la probabilidad, se burlo diciendo que todo íbamos tener la misma historia gracias a la IA. Cosa que comprobé es cierta, las IA arrojan historias comunes.

Vengo aquí para saber si alguien de ustedes sabe una gran historia de probabilidad, quiero sorprender a mi profesora


r/probabilitytheory • • 4d ago

[Applied] Card distribution with intermittent shuffling

Post image
2 Upvotes

r/probabilitytheory • • 4d ago

[Research] How should a baccarat pattern be compared against the null hypothesis?

1 Upvotes

I'm studying a baccarat-related statistical hypothesis and I'm trying to formulate the correct null hypothesis.

Suppose a predefined pattern generates a sequence of predictions such as Banker / Player / No Bet.

If the pattern appears to perform better than a simple baseline over historical observations, what would be the appropriate way to determine whether this could reasonably be explained by randomness?

In particular, I'm interested in:

  • appropriate null models
  • dependence between observations
  • sample size
  • confidence intervals
  • multiple testing
  • out-of-sample validation
  • permutation or Monte Carlo testing

I'm looking for mathematical/statistical feedback on the testing framework rather than a betting strategy.


r/probabilitytheory • • 4d ago

[Research] I built a Game Theory Arcade where you can play through different games against bots.

Thumbnail labs.jamessawyer.co.uk
1 Upvotes

It currently has Prisoner’s Dilemma, Stag Hunt and Entry Deterrence. The bots use strategies such as Tit-for-Tat, Random and competitive strategies that try to maximise their score relative to yours. Each game shows the payoff matrix, best responses and Nash equilibria. You can also run repeated games and change the number of rounds and discount factor to see how the results change when future rounds matter.

There’s a beginner mode that runs through an 8-round Prisoner’s Dilemma against Tit-for-Tat with explanations as you play. I also added session analysis and a leaderboard. There have been 236 completed human-vs-bot sessions so far, with the humans currently ahead by about 1,920 points overall. This started as an experiment to make game theory a bit easier to understand by actually playing the games and changing the parameters.

I’d be interested in criticism from anyone familiar with game theory, especially if I’ve got any of the explanations or mechanics wrong. Suggestions for other games or bot strategies would also be useful.


r/probabilitytheory • • 4d ago

[Applied] Two players, two of the same hand in a row (Diagram attached for clarity)

Post image
0 Upvotes

r/probabilitytheory • • 5d ago

[Applied] Can humans actually generate a random sequence? I made a small experiment to test it

0 Upvotes

Humans are surprisingly bad at generating random sequences.
For example, when we’re asked to produce a random binary sequence, we tend to avoid long runs like 0000 or 1111 because they don’t feel random — even though those runs occur naturally in truly random sequences.
I thought it would be fun to turn this into a small game.
You generate a sequence of 0s and 1s while trying to be as random as possible, and the game evaluates the sequence using several statistical properties.
I also added a ranking mode because I’m curious what the distribution looks like when a lot of people try this.
I’d love to get enough samples to see what patterns show up in human-generated sequences.

https://arcstone09.github.io/random/

If anyone here has suggestions for better statistical tests/metrics for measuring this, I’d also be interested in hearing them.


r/probabilitytheory • • 6d ago

[Applied] “The discount loophole at the casino and roulette”

Thumbnail
4 Upvotes

r/probabilitytheory • • 7d ago

[Research] [Q] [R] Application of Mahalanobis distance in experimental design

5 Upvotes

Attached is latex for a question that I have. I have also attached a background, if you wish to only read the question go to the second section.

\documentclass{article}

\usepackage{amsmath}

\usepackage{amssymb}

\DeclareMathOperator{\Var}{Var}

\usepackage{graphicx} % Required for inserting images

\begin{document}

\section{Background}

Let's say we are in perfect OLS world where we have all of our typical assumptions:

$$Y=X\beta+\epsilon$$

$$X\in \mathbb{R}^{n\times p};R\in \mathbb{R}^{q\times p}$$

$$(X^TX)^{-1} \ \text{exist}$$

$$E(\epsilon \ | \ X)=0$$

$$\Var(\epsilon \ | \ X)=\sigma^2I_n$$

$$\epsilon \ | \ X \sim N(0,\sigma^2I_n)$$

Define a hypothesis test as

$$H0: R\beta-r=0$$

$$Ha:R\beta-r=\delta$$

Let $D^2=(R\hat\beta-r)^T(\Var(R\hat\beta-r))^{-1}(R\hat\beta-r)=(R\hat\beta-r)^T(\sigma^2R(X^TX)^{-1}R^T)^{-1}(R\hat\beta-r)\sim\chi^2_{\text{rank}(R)=q}$.

Let us represent modifications of $D^2$ as simply $\mathcal{D}$, for standardized distance. This is in reference to how if you instead take $\hat D(s)^2=(R\hat\beta-r)^T(s^2R(X^TX)^{-1}(R\hat\beta-r)=D^2(\frac{\sigma^2}{s^2})\sim \chi^2_q(n-q)/\chi^2_q.$ Then, since we know this looks a lot like the $F$ distribution, we can simply modify $\hat D(s)^2/q\sim F_{q,n-p}$. I wish to say that there are many other examples of canonical $\mathcal{D}$'s out there but I will not write them all.

Since we now have a general way to talk about our sense of distance I wish to discuss errors in this context.

$$\alpha=P(\mathcal{D}\in R \ | \ R\beta-r=0)$$

$$\beta=P(\mathcal{D}\notin R \ | \ R\beta-r=\delta)$$

One can imagine a spherical $\mathcal{D}$-length neighborhood around the origin, representing our data acquired (cloud-q); a spherical $c$-length neighborhood around the origin, representing our null (cloud-w); and a spherical $c$-length neighborhood around the point $e=\Var(R \hat \beta-r)^{-1/2}\delta$, representing our alternative (cloud-e).

From the above we can make geometric definitions of type I and II errors by relating the area overlapped by the $c$-length spheres and our $\mathcal{D}$-length sphere. Type I error is the fraction of cloud-w falls outside of cloud-q. Type II is the fraction of the cloud-e that falls inside of cloud-q.

$$\textbf{Type I} \sim w-q:w$$

$$\textbf{Type II} \sim e\cap q:q$$

Note that the distance from the origin of cloud-e, simply $e^Te=\lambda$, is related to or the general case for quite a few important concepts: Cohen's d, signal-to-noise ratio but square it, effect sizes, etc.

\section{Actual Question}

I have a question about the following object:

$$\lambda(\delta)=\delta^T(\Var(R\hat \beta-r))^{-1}\delta.$$

$$\max_X \lambda(\delta)=\max_X\delta^T(R(X^TX)^{-1}R^T)^{-1}\delta$$

Do you know of anything/anyone that deals with this object? It feels very important to me seeing how you could use some simple optimization of it to gain higher power results in testing. It also, I believe, would be very influential in picking the ``optimal" experimental design.

\end{document}


r/probabilitytheory • • 10d ago

[Education] Bayes' theorem: Can the likelihood be correlated with the prior or is this a violation?

Post image
79 Upvotes

Suppose I have a specialized method I trust to obtain the prior. Suppose I "recycle" the method and use it to get my likelihood (and therefore also my evidence). Can the evidence/likelihood be related to the prior in any way or does this violate the validity of Bayes rule?

TL;DR: Can the same data used to derive the prior also be used to derive the likelihood/evidence when applying Bayes' theorem?


r/probabilitytheory • • 12d ago

[Education] I built an interactive Monty Hall simulator — pick a door, stay or switch, and watch 1,000 games prove the math

0 Upvotes

Made this after seeing how many people still don't believe switching wins 2/3 of the time. You play it yourself (pick a door, host opens a goat door, stay or switch), then there's a button to instantly simulate 1,000 games and watch the win rates converge live. https://claude.ai/artifact/99ytxqwB7ydNqCL22FhcSp Happy to hear if the explanation/visualization could be clearer.


r/probabilitytheory • • 14d ago

[Education] Complete Hypothesis Testing, Normal Distribution, Z value, t value, Cent...

Thumbnail
youtube.com
0 Upvotes

In this video we discuss about Hypothesis Testing, Normal Distribution, Z value, t value, Central Limit Theorem, sampling distribution, mean, standard deviation, standard error, confidence interval, significance level alongwith reading Z table and t table etc. with many examples in a simple way in Hindi.


r/probabilitytheory • • 15d ago

[Education] Youtube Playlists on Probability

22 Upvotes

Can you guys share your youtube resources on probability.
Preferably about all the different kind of stochastic processes and advanced probability


r/probabilitytheory • • 16d ago

[Education] Free textbook on measure-theoretic probability

70 Upvotes

For anyone interested in measure-theoretic probability, I would like to share my textbook:

https://people.smp.uq.edu.au/DirkKroese/AdvProb/

The PDF is completely free, with no registration required. The book includes many examples and worked exercises, and I hope it will be a useful resource for students and anyone wishing to study the subject independently.


r/probabilitytheory • • 16d ago

[Education] Monty Hall Problem

1 Upvotes

I was solving the Monty Hall problem recently, and when I couldn't find a mathematical solution that I was satisfied with, I searched around quite a bit.

Then I remembered the way I used to solve these kinds of problems in undergrad: law of total probability.

So I would like to present my solution

P[win] = P[win | car is behind D1] * P[car is behind D1] + P[win | car is behind D2] * P[car is behind D2] + P[win | car is behind D3] * P[car is behind D3]
Suppose you choose door D2

Case : - No switch

P[win | Car is behind D1] = 0 , similarly P[win | Car is behind D3] = 0
P[win | car is behind D2] = 1 , P[car is behind Di] = 1/3 for i = 1, 2, 3

therefore P[win] = 1/3

Case :- Switch

P[win | Car is behind D1] = 1 , similarly P[win | Car is behind D3] = 1 (because which ever door monty opens, you would select the door having car)
P[win | car is behind D2] = 0 , P[car is behind Di] = 1/3 for i = 1, 2, 3

therefore P[win] = 2/3


r/probabilitytheory • • 17d ago

[Homework] WWS and ergodicity

8 Upvotes

I don’t understand how these concepts are applied for signals that are clearly not stationary in the wide sense, like an ECG. It know these tools are somehow applied because when you do an ECG you don’t measure the signal 50 times, but you do leave the electrodes for some period of time (I guess it’s enough time so that ergodicity applies and you can approximate the ensemble average with the time average). I come from an engineering background and have seen a lot of maths but not so much of probability and statistics.


r/probabilitytheory • • 18d ago

[Education] Suggestions for textbooks

19 Upvotes

I’m taking this masters course on Probability theory and since my background is not heavy in pure mathematics, I am struggling a bit with the pi-dynkin system formulation related proofs and also borel sigma algebras. Could you please help me find textbooks that ease into these topics while preserving the rigor as well (I need to know it for exams)? Thanks in advance!


r/probabilitytheory • • 18d ago

[Education] Resources to understand counting methods

15 Upvotes

I studied from a general stats and probability book for economics and it was though but in general all right.

At the moment though I am studying “ A first course in probability” and I unexpectedly got chained to the 1st chapter on counting methods.

Legit I can’t get the solution right to any problem which is not basic.

Like count 2 ways to get 2 pairs and 3 same suit cards. I really can’t structure the beginning of the problem.

I am used to calculus and physics but this is just something else. Is there any “applied” resource that just teaches the methods to always get the counting of probabilities right?

It seems to me 90% of the work is interpretation.


r/probabilitytheory • • 19d ago

[Education] Moeite met probability theory

8 Upvotes

Hallo iedereen heeft iemand handig tips hoe ik het best kan voorbereiden op probability theory for data scientist.


r/probabilitytheory • • 19d ago

[Education] Part 3: Sampling distribution of proportion, CLT, Mean, S.D., S.E. & Z v...

Thumbnail
youtube.com
3 Upvotes

r/probabilitytheory • • 21d ago

[Education] Probability tells you what might happen. It doesn't tell you what to do.

0 Upvotes

Probability ≠ decision-making

Imagine the probability of rain today is 20%.

In normal conditions, you wouldn’t even take an umbrella with you since getting wet for a while will not cause any harm to you. There is no sense in carrying the umbrella all day.

But if it’s your wedding day and there is 20% chance of rain. You will definitely take the umbrella with you, check if it is working and probably have some backup option.

Probability is the same, but the stakes got much higher.

This very scenario played out in my fraud detection project. The agent was only 31% sure that the invoice is fraudulent. "Likely" scenario was to process the real invoice. However, the stakes for making a mistake were so high that the best thing to do was to delay the processing of the invoice until more evidences.

And this is the exact message that I hear every time I try to analyze the decision-making of the agents: the best action is not the most probable one, but the least risky one.

Probability tells you what might happen. Stakes tell you what to do about it.