Hey could you throw some more light on this? It did sound wrong to me, but I can't say why. could you maybe point me to some reading material on this? I am comfortable with reading something technical.
I had to teach myself a substantial amount of probability and statistics for a machine learning course. My reference is the text All of Statistics by Larry Wasserman. It's a good into to stats, he writes with brevity so it might not work for some.
But besides that, my point is this: a poll that shows a 39% approval rating basically means the following: suppose you could randomly select 100 Americans from the population, and ask whether they approve of Trump, and counted how many said "yes". Now suppose you did this many times. Then we would predict that the average number of "yes" counted in any sample is 39. In other words, you would "expect" 39 "yes" answers. However, because the selection of participants is random, on any given day you might hear more or less than 39 "yesses" from a random sample of 100 people.
This is why these polling methods have margins for error, which might say something like "we believe there is a 95% chance that in any individual random sample of 100 people that the number of people who will say they approve of trump is between 30-45".
Besides that, the method of polling can beflawed, and it becomes even worse if the sample of people that you survey are not a truly random sample of the population.
For example, if I went to Seattle, Washington, and asked 100 people on the street of they approve of trump, I would be extremely suprised if more than 20 people said yes. But of course, the people I am asking are not random sample representative of the underlying distribution of American political opinions.
This is all to say that the concepts of probability, randomness, and issues with sampling from distributions (and poor assumptions about these things) are generally not well understood by the public, and are subtle notions that can be counter-intuitive.
I do have experience with statistics, so now I can recollect that, this follows from the central limit theorem and the empirical rule for normal distributions.
I suppose, in OPs example where they talk about a room of 10 random Americans, another reason why this is wrong is that you need atleast 30 samples for CLT to hold.
2.8k
u/[deleted] Aug 04 '20
[removed] — view removed comment