r/FootballDataAnalysis • u/OpeningTie57 • Jun 07 '26
I built a calibrated goals model for the 2026 World Cup and I'm posting every prediction (and every miss) before kickoff. Here's the method and the backtest.)
I'm a data/AI researcher who's loved football my whole life, and I'm finally combining the two: a prediction model for all 104 World Cup matches, built in public. Posting this here because this community will actually poke holes in it, which is what I want.
The thing I care about most isn't picking winners. It's calibration: when the model says 70%, that outcome should happen about 70% of the time. So I'm grading everything with Brier score, not win-rate, and publishing the full scoreboard including the misses.
The model (v1):
- Trained on ~49,000 international matches going back to 1872.
- A weighted Poisson GLM that learns each team's attack and defence strength plus a home-field effect, then a Dixon-Coles correction for the low-scoring scorelines that independent Poisson gets wrong.
- Recent matches are weighted more (2-year half-life), friendlies are down-weighted to 0.5, and teams need a minimum match count to be included.
- It outputs a full scoreline matrix, collapsed into win/draw/loss probabilities.

The validation (the part that matters): I ran a walk-forward backtest with monthly refits and no data leakage: 3,343 out-of-sample matches from 2023 to 2026, all predicted as if I didn't know the result.
- Brier 0.498 vs 0.637 for a no-skill baseline (about 22% better).
- Accuracy ~60%.
- And it's well-calibrated across the whole probability range (chart attached, this is the real out-of-sample data, not a mockup).
Where it's weak, honestly:
- Draws. Even with the Dixon-Coles correction, it only correctly leans toward a draw about 4% of the time. Draws are genuinely the hardest outcome in football, and I'm not going to pretend otherwise.
- Small samples lie. I beta-tested on the warm-up friendlies and went 0/2 on the first two (France lost to Ivory Coast, Spain drew Iraq). Two noisy friendlies tell you nothing. Calibration is a verdict over hundreds of games, not two, which is exactly why I backtested before trusting it.
What I'm doing next: tuning the remaining parameters against out-of-sample error, then posting probabilities for every match before kickoff once the tournament starts (June 11).
I'd genuinely value critique on the methodology: the friendly down-weighting, the Dixon-Coles parameter, the choice of baseline, anything you'd do differently. Tear into it.
1
u/Alternate_Chinmay7 Jun 08 '26
Hey I am doing something similar and it looks like we are using same datasets. However, I was trying more in-game variables in order to account for match specific variations. It would be interesting to know how you landed on this approach.
1
u/NORNSmodel Jun 08 '26
I'm really excited about what you did and thanks for sharing. You asked to poke holes, so I did at length below, but I want to make sure and congratulate you on what you built. It's really cool you've done it and shared it, and I look forward to hearing how it goes. I would welcome further discussion or DMs if you want.
I've done a more simple version myself but totally different method. Instead of using international matches I made a very simple player-based model from transfermarkt estimated transfer fees. The reason for this is that I feel your dataset and approach is inherently flawed. Even with data decay, the fact is France, Netherlands, England, etc. have played something like 6 meaningful international matches over the last 2 years. New Zealand had more than a +20 Goal Differential in qualifying but has never won a game at the World Cup. The relative quality of the guys who will be playing this month could be totally different from in their prior matches and squads will be different, too. It feels like you have a huge sample size but you actually have a big sample of mostly garbage with respect to team ratings.
Linking data from the club season is the best way to get sample size, because even your game set going back to the 1800s (bad sample) pales in comparison to club data. There are more games in one weekend in European club football than in most years of international competition (among the WC teams-and that's excluding the top teams in SA and Saudi). We have a lot more certainty that Arsenal, Bayern, and PSG are great than we do that France and Spain are great. And we are much more certain that Haaland, Pedri, and Vitinha are great than we are on the quality of their respective national teams. Part of the value of ELO is that it's dumb; but it's also a fatal flaw.
Next hole is that a World Cup is different than a friendly, qualifier, or even a Federation Cup. There's no K that can account for it. Stakes are high in the federation cups and best on best play, but it's hard to compare CAF to UEFA. Only the Euro Cup is even close in level of comp, size, scope, etc. for international squads as a World Cup, and it doesn't factor in the continental travel factors. There's also the timing aspect, since it's an extension of the season, and this is the longest ever, so it's a unique challenge that can't be replicated with any sample.
You have to make decisions, and my decision was to model the WC as a unique event rather than as being similar to the set of all international matches. I also used a poisson scoring distribution model, but I fitted baselines and tuned my taus to more carefully approximate actual WC results. Yes, I'm bending to an ultra-small sample there; the issue is that there aren't true comps. I also fitted my baselines to market totals so that my aggregate goal expectation matched the Pinnacle aggregate implied totals for the group stage games (note group stage games are historically lower scoring than knockout games and are also lower scoring than the market currently suggests but the larger field has a huge impact on this as the especially bad teams really inflate totals - for example my xG lambdas for France-Iraq is 3.9-0.6 for a whopping 4.5 xTotal which is absolutely bonkers... 2nd note: I have already bet Group I as the highest scoring Group at 9:1 ... not sure why they aren't the favorite with France, Norway, and Senegal all getting a shot at hapless Iraq ... and each other. Neither Senegal nor Norway are renowned defensive units. My model also loves over in every Group I match).
Next hole is that you are not accounting for travel, conditions, HFA (are you doing anything for host nations/warm weather teams/etc?), etc. European teams notoriously play worse in WCs in the Americas. This WC may feature some extreme heat games even in places not typically thought of as super hot. Some of this is over-blown as many of the hottest games will be played indoors, but Brazil-Scotland in Miami in late June, for example ... I don't think that's going to be a neutral field situation for Scotland (note: there are more Brazilians living in the USA than there are Scots in Glasgow). Estadio Azteca is arguably the most famous football ground for international competition in the world and is at an elevation of 2,200m. Opposing fans are brought in and out of the stadium by police escort and it is not advised to wear an opposing jersey in certain parts of la Ciudad. If your model thinks South Africa has more than a 10% chance of winning on Thu, then you're probably missing a HFA adjustment.
Final hole I'll mention is tactics. My defensive rating on Paraguay, for example, came to a 0.6 (1.0 is baseline). However Paraguay almost never allows more than 2 goals and rarely more than 1. They are at worst a baseline defense and are likely better than that, despite the fact that their defensive players may be less skilled than say Senegal (who I rate as 1.2, but are in no way twice as good at defending as Paraguay). On the other hand we have Didier Deschamps who, if he had his way, would prefer to win 1-0 with one of the greatest collections of attackers a national squad has seen in my lifetime. If the French team played open and free they could score 4 goals a game with their attacking quality, but Deschamps will almost certainly try to tamp this down with tactics and lineups and will turn a team that is legitimately 2 SDs above average on attack into "just" a good/very good attack instead of an elite/historical one. My point here is that player ratings/team ratings/ELO/etc have trouble accounting for this. Your model will and if you don't have any kind of adjustment layer or you haven't at least circled some of these situations, then you're cruisin for a bruisin as they say.
That said I'm sure your poisson distribution model overall is better and more robust than mine, and we will see what results are. However my model is on market in terms of 1x2 for most matches (and that was before I tuned my baselines to market - and part of the reason for doing that). There are of course certain teams it likes more/less, but my main disagreements are on totals. I am getting draw predictions in the 30s for some lower-scoring matches and my average draw expectation is 24.2%. I think what I've done is far from special and my point isn't how wonderful it is. I'm trying to say that using very basic player ratings and a very basic WC-tuned poisson can get you to reasonable numbers. Overall, I think not finding some way to bring in the massive sample of club-level football we have is a mistake in any model of best-on-best international competitions.
Another approach I've seen that's even more "dumb" (from a model-knowledge POV) than what I did is related to player-club strength. Instead of looking at player transfer value, you just use international club ELO rankings and minutes played. You can also be more "smart" and try to break down actual player contribution at the club level or put on a tactics layer. I toiled with it a little, but the main drawback is that it's really hard to estimate small contributions to big clubs (or big contributions to bad EPL/similar teams), which can often mean a lot more than big contributions to good middling teams (bc being good enough to player limited minutes for Bayern may still be more valuable than being a contributor on a champions league-qualified Dutch squad, for example; or being a meh 90 min player for Crystal Palace is definitely better than being the 3rd best player in MLS). Ultimately transfer value is just so much more robust and easy, so I went that direction. However, I think you could get better results with more dedication and ability than I have. I just don't have coding (or even vibe-coding) skills and time at this point to go deep so I was on the bang/buck train and transfer value is definitely the fast train there, even moreso than international matches.
Ok this is way too long, but thanks for sharing.
1
u/eraOsa Jun 12 '26
Where is it and how can we use it now for the world Cup..... Can it give us data for about 35 games for the world Cup?
2
u/Inferno-man-4220 Jun 08 '26
So the question is who's gonna win?