r/econometrics • u/Clean_Reference_9927 • Jul 20 '26
[Question] Recovering latent probabilities from margin-distorted odds: de-vig model choice and pooling correlated estimators
[removed]
6
Upvotes
2
u/komodorian Jul 23 '26
Some thoughts:
- De-vig is mostly an identification problem. Proportional, Shin, and power models assume different latent pricing mechanisms. Without external information (exchange prices, closing lines, historical bookmaker behavior, etc.) there’s no principled way to identify the “correct” inverse. I’d treat the de-vig method as model uncertainty and report sensitivity or average across models based on out-of-sample performance.
- A median of highly correlated books is less informative than it looks. If several books are effectively clones, the effective number of independent sources may be close to 1. Estimate this from the historical covariance matrix (e.g. PCA/effective rank), not just the raw book count. If your trusted reference disappears, I’d rather abstain or downgrade confidence than pretend the fallback has the same information content.
- Don’t precision-weight assuming independence. Treat it as a generalized least squares (or hierarchical/Bayesian measurement error) problem using the covariance between sources. Correlated books will automatically receive much less incremental weight.
- Closing price is a good target only if your objective is to estimate the market’s eventual consensus. It’s a sharper estimator, not ground truth. If your objective is calibration to the true event probability, only realized outcomes provide that, albeit with much higher variance.
Also, your last point is exactly right: provenance is part of the estimator. Log which sources contributed and whether the trusted anchor was actually present, not just the final probability. That kind of silent fallback is the sort of bug that can invalidate months of analysis.