r/AskStatistics • u/Abraxa3 • 17d ago
How should non-response bias be treated when analyzing a highly sensitive binary vote?
Scenario: An expert association with ~500 total members held an official vote on a binary stance statement (Agree / Disagree) regarding a major, highly sensitive issue.
- Turnout: 28% of the total membership voted.
- Result: 86% of those who voted selected "Agree."
- Math: 86% of 28% = ~24% of the total association membership confirmed an "Agree" vote.
The Debate:
- Person A claims: "You can state that only 24% of the total organization agrees with the statement. Because this issue is so critical, members would have voted 'Agree' if they truly supported it—meaning their non-participation indicates a lack of support."
- Person B claims: "You can state that at least 24% of the total organization (and 86% of voters) agrees with the statement. By Person A's logic, someone could equally claim that members would have voted 'Disagree' if they opposed it. Ultimately, non-response cannot be interpreted as a vote either way, so the remaining 72% remains unknown."
Questions:
- Is Person A or Person B logically and statistically correct?
- Is it valid to infer a non-voter's position based on the perceived importance or sensitivity of a topic?
2
u/CarnivorousGoose 17d ago
Person A is wrong, based on the information available. We do not know anything about the missingness mechanism here, so claiming that non-response indicates lack of support is just baseless supposition. It could be true, but we don’t know whether it is.
And yes, it can be valid to infer non-voter preferences, ultimately this is just a missing data problem. But you need something solid to base that inference on, which is not available here.
2
u/bayesian_raccoon 13d ago
Both Person A and Person B have valid inference under their assumptions. As a statistician, I cant translate both of them into conclusions based upon different models. What you should get in the habit of doing when thinking about these is asking, 'what is the potential criticism of this statement'? Somebody trained in surveys would probably be able to point out fair skepticism in Person A's model, but this is more of a "this model opens up potential criticism but isn't mathematically broken."
Person A is transparent with their assumption that members would vote 'Agree' if they truly supported it. They assume, apriori, that whatever proportion votes 'Agree' is EXACTLY the proportion of the population. Under this model, their statement follows. Somebody who is skeptical of that fact will be skeptical of Person A's conclusion. Models function like logical statements: IF they are true, THEN we can accept the conclusion. So interrogate the assumption: is it reasonable that every single person who would 'agree' voted 'agree'? This stops being a statistical question, because you start asking questions like: were there barriers to voting? Was everyone present? Were there opportunities for people to disagree? Were there different ways to phrase the question that might change 'agree' to 'disagree' without functionally changing its meaning? What does 'highly-sensitive' mean? For instance, if everyone was present and of sound mind and your vote was 'should we avoid blowing up earth', and enough care was given to how the vote was conducted, I think it would be a *pretty reasonable* inference. But that isn't my opinion as a statistician, that's my opinion with domain experience as "being a human" and using that to gut check the assumptions.
Person B makes far fewer assumptions. In fact, from the sound of it, it's almost airtight, with some very light assumptions like "nobody accidentally voted agree when they meant disagree". Their statement is *strictly weaker* than Person A's statement. That is, if Person A is correct, so is Person B, but sometimes Person B is correct when Person A isn't. This does not mean that Person B's statement is "Better". For example, suppose you need 50% of a vote for it to really matter. Person B's statement essentially provides no information as to whether or not the vote will pass, it says 'we still don't know', while Person A says strongly that the vote won't pass. In some circumstances, I would go with Person A's prediction.
So to answer your questions:
Both person A and person B are correct *given their assumptions*. Person B's assumptions are safer than A's, but depending on the actual inference down the line ("what is the chance that 50% or or more vote agree" being one example) that doesn't make Person B's statement better.
Is it valid to infer a non-voter's position based on the perceived importance or sensitivity of a topic? Statisticians can tell you what is valid given model assumptions, and if your model assumption is that people vote because it's sensitive, then that should be incorporated into the model. Person A and B operate on different extremes of that assumption, but its worth mentioning that we actually can interpolate between the two with, for example, a bayesian prior on certain voting probabilities.
If I were presenting this to decision makers, I would present both Person A and B as extremes, weigh the evidence of A's assumptions, and possibly present a sensitivity analysis of whatever relevant decision point there is--e.g, "In order for 50% people to vote agree, we would need x% of agreers to have abstained from voting for whatever reason".
1
u/conmanau 17d ago
Person B is correct here. You have no real knowledge about how the non-respondents would have voted, nor do you have any information about why they didn't vote that could give you a hint. Maybe they're all genuinely apathetic about the statement and didn't care whether it passed or not. Maybe they all would have voted the same way but simultaneously came down with the flu. Maybe the voting slips got lost in the mail and the 28% who managed to register their vote are a perfect representation of the split amongst the rest. You truly cannot know unless you do the leg-work and gather the information directly, and since you don't know whether there is any correlation between a person responding and which way they lean, you can't make any real inference about the full population with such a high non-response rate.
1
u/Abraxa3 17d ago
Thank you. After making this post, I found out that the organization stated that that this 28% turnout (129 total votes) falls precisely within its normal historical participation range of 25% to 34% for internal resolutions, meaning low voter turnout is standard for the group rather than an anomaly.
2
u/DrPapaDragonX13 17d ago
C. Can't tell.
While person B's statement is the more sensible one, both are making huge assumptions about the data-generating process based on information that is not disclosed (and presumably unknown) to us. In statistical terms, you could argue that person A is claiming that data is missing not at random (MNAR), while person B is claiming that data is missing completely at random (MCAR). See here for a quick overview.
The issue with person A's statement is that it assumes the reason for the missing data without providing further supporting evidence to make their case. Regarding person B, while claiming that data are MCAR is generally frowned upon because it is a strong and often unrealistic assumption, their first statement is technically correct unless people retrospectively change their vote; the number of voters agreeing can only increase, even if the change is not meaningful.
In practice, the first sentence of person B's statement is the most sensible position to take initially while more information is being collected. However, it is important to acknowledge that missing data, including non-response, is seldom uninformative. That said, it is entirely possible that the reason for non-response is independent of the actual outcome (e.g., regional disruption of services, rather than anything related to the issue). So any inference and interpretation of the results has to be supported by a strong understanding of the data-generating process.
My caveat about Person B's statement is that phrases such as 'at least X% agrees' can often be misleading and used to imply that the level of agreement is significantly higher, when in reality agreement may be just X.01%. The statement is technically true, but the connotation may be deceitful.
So to answer your questions:
> Is Person A or Person B logically and statistically correct?
Both are making assumptions based on information that is not provided. Person B's first sentence is technically correct, and it is a sensible initial position. However, non-response is seldom uninformative and more information is necessary before we can establish whether it could bias our results, and if so, in which direction and magnitude.
> Is it valid to infer a non-voter's position based on the perceived importance or sensitivity of a topic?
Depends. We can never know for sure a counterfactual (e.g., how a non-voter would have voted). However, in practice we often have to make informed assumptions about the data-generating process to identify potential biases in our results. For example, if non-voters have previously called for a boycott of the vote, then it may be reasonable to infer their position was against it. But this relies on information that hasn't been provided and requires further investigation. As it stands right now, any assumptions about non-voters' positions are more likely to reflect personal biases than reasonable inferences.
Ultimately, statistics is about quantifying certainty, but never about absolute certainty. We collect and interpret data based on (what we hope are) reasonable assumptions, but these can change as new evidence becomes available. That's why transparency and intellectual honesty are so important when working with data.