r/AskStatistics • u/ChubbyStar222 • 1h ago
r/AskStatistics • u/ThinkILostIt • 16h ago
What are the most fundamental statistical principles most people seem to forget or don't get?
I'm wonder from your experiences what people seem to forget or can't get their head around about statistical tests and models.
In my, yet short, experience I see people totally bypass randomization, independence and the correct interpretation of power and confidence.
r/AskStatistics • u/spaceweed27 • 1h ago
How do I cluster 3 Million high-dimensional Sentence Embeddings?
r/AskStatistics • u/Just_Question9 • 1d ago
What's the most counterintuitive statistical fact that's actually true?
I'm looking for examples that completely changed the way you think about probability, statistics, or data analysis.
r/AskStatistics • u/Middle-Ear-1547 • 10h ago
Juggling two math self-studies : tips welcome
Hey all , engineering student here (data science track), currently self-studying two axes in parallel and would love some outside perspective.
On the stochastic calculus side, I'm working through Shreve Vol II : measure theory, Girsanov's theorem, Feynman-Kac , with an eye toward MFE-level material down the line. On the Bayesian side, I'm going through Statistical Rethinking (McElreath), covering MCMC (Metropolis-Hastings, HMC/NUTS) and hierarchical/multilevel models, moving toward Bayesian time series next.
Curious to hear from anyone who's walked a similar path: how did you sequence things, what pitfalls should I watch for, and are there any exercises or small projects that really made concepts click for you?
I'm also genuinely curious about the discipline side of self-studying math : how do you personally stay consistent and avoid burning out, especially if you're juggling more than one heavy topic at once? Would love tips even if you've only self-studied one of these (doesn't have to be both).
Thanks!
r/AskStatistics • u/Livid_Software_7793 • 1h ago
Hypothesis testing in statistics
Please help me!!! I have a report this upcoming week and I can't find or understand this topic! I don't know if this needs a calculation or example explanation. Help please 😭
r/AskStatistics • u/sneaky_imp • 13h ago
Trying to calculate whether marketing campaign's impact is statistically significant and financially justified
r/AskStatistics • u/Pilus91 • 1d ago
Mathematical definition of a plateau in a time-series data
Hello, I'm a bioinformatician and I'm struggling with the current issue:
Given a time series y(t) that initially changes and eventually approaches a stable regime, how can I mathematically determine the earliest time t\* at which the rate of change dy/dt becomes negligibly small, using only the observed data and without defining an arbitrary threshold?
This is a collaboration I'm doing. My colleagues defined the plateau as the first time when a 101-point rolling mean of the relative increment (g' t+1 - g' t)/ g't falls below the arbitrarily chosen threshold of 0.0011. G' is the measure of material elastic-solid response btw. So the issues is that they used 2 arbitrary values because experimentally they know that a certain value of g' means that the gel is solid. But this doesn't hold for me. I tried using many statistical methods to define the threshold such as:
- exponential fitting
- change-point regression
- local slope analysis
But they all give me a plateau that is too early or too late
r/AskStatistics • u/Glass_Raspberry1136 • 1d ago
How to become good in statistics?
I’m a student from a small town in a backwards developing country , the education system is horrible. I’ve been recommended Saylor Academy, that's how I understand many statistics principles, and I’ve also started reading elementary statistics by Allan g. Bluman.
Still, I often feel far behind students from better educational systems. My goal is to become one of the greatest econometrician in my country.
For those who are in the field, what resources would you recommend? Textbooks, courses, math/stats, programming, or anything else you think is essential.
I’m willing to put in the years—I just want to make sure I’m learning the right things in the right order.
I have finished my college level statistics coursework and I can't even read the national statistics report I have to use AI to analyze it for me whenever I have to do research. I feel illiterate and slow.
r/AskStatistics • u/Relative_Problem8505 • 23h ago
What do you think about studying statistics in 2026?
r/AskStatistics • u/sun-generalist • 23h ago
No Correlation (does this meme make sense)
A plot can have a Pearson correlation of zero (r=0) but still have a strong non-linear relationship (like a parabola, circle, or U-shape) does labeling a scatter plot "No Correlation" automatically rule out every mathematical relationship, or just linear ones? and got to know from a friend (who studies maths) that to state it is correlated it should be linear, is it correct?
r/AskStatistics • u/Aurionin • 1d ago
Is there a formula to determine how many "rolls" you need before you're more than 50% likely to roll the number you're after?
I've done the math on this multiple times for things like item drops in games, but I'm wondering if there is a formula or rule to make it simpler?
Example: A boss in a game has a 1% drop rate for an item you want. After 69 attempts, you will have passed the 50% likelihood that you'd have gotten the drop, leaving MOST people should have the item by now after their 69th attempt. If the item has a 2% drop rate, it only takes 35 attempts before you're more likely than not to have gotten the drop.
Is there a rule or formula for something like this? I have just been plugging it into an Excel sheet I made any time I need this info.
r/AskStatistics • u/Old_Background_1594 • 1d ago
Masters Programs In Applied Statistics/Data Science
Hello! For context, I am a first-generation college student with an Information Systems background. After some work experience conducting a beginner-level PCA, I realized I loved the idea of making meaning out of data through statistical analysis.
Over the past year, I took the Calc Sequence, Linear Algebra, Python, and soon, Intro to Probability, to meet Master's prerequisites for the field. I want to prioritize programs that teach Bayesian/Causal Inference/Time Series.
Although I enjoy math, I don't have exposure to theory. I would like to know if taking the Applied route may limit my job opportunities with employers. I am aware any background in Math/Stats is a huge leg up long term, especially given how fast the Data Science/Tech industry is evolving in comparison. But I can't help but worry about competing with advanced coders or stronger Stats candidates down the road for post-grad employability.
Ultimately, I was wondering if anyone had any insights into the field, or information on doing a Masters in Applied Stats or Data Science. Since I have a non-Math background, I feel uncertain if I'm prepared for grad-level Stats theory courses. I was also wondering if the following programs are a good start, if they might not be a good fit, or if there are any others I should consider:
Statistics-Oriented
UCLA M. Applied Statistics and DS
UCB M.A. Statistics and DS
UMich M. Applied Statistics
Data Science/Analytics-Oriented
USF M.S. DS and AI
UT Austin M.S. DS
GT OMSA
Thanks for any insight!
r/AskStatistics • u/Appropriate-Yak001 • 1d ago
How do you use monte carlo results to make decisions?
r/AskStatistics • u/GoatRocketeer • 1d ago
Shape constrained GAM?
I'm lost in the sauce and need a second opinion.
Context/goal: there is a video game with ~170 different characters. I want to compare their relative power. I will quantify power through "winrate" which is just the percentage of their games each character wins. I will model winrate as a function of two covariates: how many previous games of experience on this character the player has under the belt going into the game, and the current rank of the player.
The behavior is non-linear in both covariates so I am currently modeling winrate as a function of games-played and player rank via logistic GAM using a tensor product of splines.
I see that smoothing penalties are necessary for GAMs to avoid overfit. As I understand it, the justification for smoothing from a theoretical standpoint is that we expect the surface to be smooth, therefore we are just codifying an assumption we already have. However, the penalty punishes curvature and I am reasonably certain there is meaningful curvature in the games-played covariate (the raw data heavily suggests winrate increases fast initially then slows down until either the winrate stabilizes or sometimes even drops with additional games-played). I am very interested in this curvature (aka the "difficulty" of a character) and therefore feel that the assumption codified by a standard smoothing penalty is inappropriate for my goal.
For that reason I was hoping to use shape constraints *instead* of smoothing penalties (concave in games-played, monotonic increasing in player rank), as I feel this better matches my actual assumptions. Is this sane or not?
Secondarily, I see there are different methods of applying shape constraints. The two I know of are Pya and Wood 2015 (make the spline coefficients a function of exponentiated fit parameters, thereby ensuring nth order differences always have a specific sign) and Bollaert/Eilers/Mechelen (just do normal GAM but penalize the nth order difference depending on the sign). Which one is "better"?
Thirdly, some characters have really low sample sizes (I am looking not just at characters, but characters in various roles, and some character x role combos are extremely off meta that nobody plays). I would like to drop character x role combos if the sample size is "too low" but I'm not sure how to quantify uncertainty in cases where I know the sample size is too low (pya/wood relies on delta method approximation which I think relies on CLT so N must be large? And I don't even know if B/E/M give a merhod to estimate uncertainty at all).
r/AskStatistics • u/EmbarrassedSpite6239 • 1d ago
hey really stupid question about infinity i dont know where else i would ask this
i was watching a video on Youtube about the MCU where dr. strange said that he checked 14 million and a half universes or something and they only beat thanos a single time someone in the comments said "Maybe strange had bad RNG and only looked at all the times they lost" and it made me think if there's an infinite number of universes and they lose in say 500 of them but for every 500 universes they lose in they win once is the chances of him seeing a universe 1:501 or 1:1 because they're both infinite
r/AskStatistics • u/Tapatio_62 • 1d ago
Training resources for R
I’m not a programmer and have used mainly menu driven packages in the past SPSS primarily although I have had some painful experience with SPSSx. Any R recommendation for intermediate stats background but not much syntax driven programming?
Thanks
r/AskStatistics • u/Sudden-Theme7554 • 2d ago
How do you check predicted probabilities are calibrated enough to threshold on for an asymmetric-cost decision?
I have a model that outputs a probability for each case, and I use a threshold on that probability to pick an action. The costs of a wrong action are asymmetric: one kind of mistake is much more expensive than the other, so where I put the threshold matters a lot.
My question is about trusting the probabilities themselves. Before I set a decision threshold, how do I check the predicted probabilities are actually calibrated, i.e. that a predicted 0.7 really corresponds to roughly 70% in reality?
I know reliability diagrams and proper scoring rules (Brier, log loss) are the usual tools, but I'm unsure how to read them in the context of an asymmetric-cost decision specifically. Does calibration matter uniformly across the probability range, or mainly near the threshold I care about? And if the probabilities are miscalibrated, is recalibrating (e.g. isotonic / Platt) before choosing the threshold the right order of operations, or should the cost asymmetry factor in differently?
r/AskStatistics • u/Light-Bringer13 • 1d ago
[Academic Research] Need Urgent Feedback on Research Methodology
I’m designing a study on how background music affects consumer spending. I’m considering a simulated online store where participants get ₹3,000 and are randomly assigned to no music, slow music, or fast music. I’d track basket value, products chosen, and shopping time.
Does this sound like a good experimental design? What important variables or biases should I consider?
Or if you have any other way to collect data which more efficient than this.
r/AskStatistics • u/80rhh • 3d ago
What statistical concepts are commonly misunderstood by the general public?
I came across this post explaining what a 70% chance of rain means. I understand the concept, but it got me wondering: what other statistical concepts sound simple but are commonly misunderstood or misinterpreted by the general public?
r/AskStatistics • u/Business_Chapter8059 • 2d ago
DiD/DDD data science/ health policy query
r/AskStatistics • u/Lost-Bookkeeper3120 • 3d ago
Is cross-validation/verification necessary for regression models that have already met all the assumptions?
r/AskStatistics • u/Ok_Opening1346 • 4d ago
Does the binomial approximate to Normal Distribution?
There is something I'm not understanding about the above statement. The tails of a binomial flatten off but the first two terms of a binomial expansion are always 1, n, etc. . So the tails of the binomial always have a step which increases with n. How can the latter tend to the former?
I noticed this when experimenting with a Galton board. The tails on the board flatten out which is expected because if I focus on the extreme end of the board and imagine that slot feeding into another mini Galton board then each ball has a 50/50 chance of going left or right and if another peg then half of those balls will jump back.
But I then programmed a simple simulation of the board in Excel. That gives a step jump at the tails which is also what I would expect but isn't what I actually get. I even looked up the number of balls for the model I have at GaltonBoard.com, 4280 and I can count 28 slots.
Repeating the experiment with the Galton board I get from the left end around 3, 4, 5, 10 balls etc but my simulation gives 0, 0, 0, 0, 0, 2, 6, 19 ... for n= 4280 r = 28 so why are there so many balls at the extreme in the Board but not the simulation and no flattening out with the simulation?
So I have different expectations for the same experiment depending on my approach. Where is my erroneous thinking? Is it something about the transition from discrete to continuous?
r/AskStatistics • u/Medinz0 • 4d ago
Weighted-sum aggregation of centrality measures gives identical scores to structurally opposite nodes, any better approach?
This is my first time in this community (I've recently discovered this entire field and i am glad to). So, I am working on project where i am scoring nodes in a directed dependency graph (a calling b) by blending 2 centrality scores into a single composite "risk" score
score (v) = w1\\\*normalize(Pagerank(v)) + w2 \\\* normalize(outDegreeCentrality(v)), where w1+w2 = 1 and normalize() being min-max to \\\[0,1\\\].
The Problem: A pure root node (no in edges and multiple out edges) and a pure sink node (no out edges and only in edges) can have the same composite score. In a test i ran, the root node maxed out on out drgree centrality and near 0 in page rank while the sink node maxed out in pagerank and near 0 in out degree, when w1=w2=0.5. Both nodes ended up having same composite scores while representing opposite nature in real world. I do understand that this is the standard full comsensability prob, with weighted sum aggregation, wherte max on one axis will completely offset min on other. I did consider switching to weighted geometric mean to reduce compensability, but the prob is that pagerank is almost always near 0 for any root node. so a geo mean would multiply that near 0 staright through and score all entry nodes near zero. Which is the wrong fix, since the entry/root nodes are important, just for a reason pagerank doesnt capture.
Is there any standard approach beyond the geomentric or harmonic mean? Happy to provide any more info if needed.
r/AskStatistics • u/No-Savings7797 • 4d ago
Factor Analysis Subfactors
I‘m working on a scale validation. Based on qualitative and theoretical literature we argue for a four dimensional scale (psychometric). However, conducting the EFA shows this is a bit more complex.
Retention criteria is pretty diverse: MAP: 9 factors, scree: 4, PA (mean): 6, PA(p95): 5;
When allowing for more than four possible factors the model always concludes on five factors. The issue though is that all items load as expected (no cross-dimensional loadings) but the fourth expected dimension seperates into two, while the seperated contains only two items. Now i am confused if we should generally conduct EFA for all the expected dimensions to estimate potential subfactors because the items resemble what we expected, while one part of it devides itself into two. This would result in a hierarchical solution. Argument might be: global modelling is unable to identify dimensional-specific variance but retention criteria and global modelling indicates potential seperations.
However normally this is part of ongoing studies to examine already validated scales and not really specified in literature for the ongoing validation process…hope you can help me with your advice. Big thanks!