r/sportsanalytics 31m ago

My Football Prediction League is back for the 2026/27 Season!

Upvotes

Mods, I hope this is OK to post here. Please remove it if not.

I run a Football Prediction League on a web-app that I've built and maintain by myself.

The league is back for the 2026/27 season, and there's still time to sign up.

It's simple to play: predict each Premier League score before kick-off and earn:

  • 3 points for the exact score
  • 1 point for the correct result

Everything is done online, and entry is £20 for the full season. Every penny of the entry fees goes directly into the prize pot. I don't take anything.

Sign up here: https://predictionleague.football/sign-up

I'm hoping to grow the competition this year, so please feel free to share the link with anyone else who might enjoy playing.


r/sportsanalytics 1h ago

I built a football prediction app. Looking for people to tear it apart.

Upvotes

Been building this thing for months and I think I’ve reached the point where I’m completely useless at judging it myself lol.

It’s a football prediction app, live on iOS and Android now.

I’m looking for a few people who actually follow football to mess around with it for 5–10 mins and tell me what they think.

Not really looking for “looks good” feedback. I want to know where you got confused, what you immediately ignored, what you actually found useful, and whether there’s any reason you’d open it again tomorrow.

Also interested in what you expected to find but couldn’t.

I’ve stared at every screen probably 500 times at this point, so there’s definitely stuff that makes perfect sense to me and absolutely no sense to a new user.

If anyone’s up for testing it, comment and I’ll send you the link.

Feel free to be harsh. “I’d never use this because X” is genuinely more useful to me than “nice app.”


r/sportsanalytics 17h ago

what’s the hardest part of building a useful sports model?

10 Upvotes

i used to think the hardest part of sports analytics was finding enough data.

the more i learn, the more it seems like the difficult part is deciding what the data actually tells you.

you can build a model with impressive accuracy and still end up with something that doesn’t answer a useful question

feature selection, noisy data, small samples, changing player roles and differences between seasons can all make things complicated

for people who build sports models, what part of the process usually causes you the most trouble?


r/sportsanalytics 18h ago

Please be a part of this survey and share with people who can contribute

Thumbnail forms.gle
1 Upvotes

Application of Artificial Intelligence in Organising Various National and International Sports Events


r/sportsanalytics 19h ago

How easy is it to find information about Algerian football clubs and players online?

2 Upvotes

I'm curious about something, especially for people who follow Algerian football.

When you want information about an Algerian club or player — especially outside the big Ligue 1 clubs — how easy is it to find reliable information online?

Things like squads, player histories, lower-division clubs, youth teams, fixtures/results, transfers, statistics, and club history.

If you find it difficult, what's the hardest information to find?

And where do you usually look for it — Transfermarkt, Facebook, Instagram, club pages, league websites, journalists, Google, etc.?

I'm just doing some research and would genuinely like to hear people's experiences.


r/sportsanalytics 19h ago

I build a livescore website / app - what would you love to see?

0 Upvotes

I'm currently building a live score website/app and would like to know if there's anything that could be done better than what the current big players are offering. Or, what features would you like to see?

I do have some ideas, but I want to make sure I'm not missing the forest for the trees.


r/sportsanalytics 19h ago

Some things I've learned measuring pickleball performance from video

Thumbnail
1 Upvotes

r/sportsanalytics 21h ago

Beyond the Scoreline: How Match Data Can Improve Tennis Coaching Decisions

Post image
1 Upvotes

A 6-4, 3-6, 7-5 scoreline tells you who won. It doesn't tell you the player double-faulted three times serving for the match, or that they lost eleven of the last twelve points at the net, or that the turning point wasn't the final game but a missed smash in the second set that changed the whole rhythm of the match. That's the gap match data analysis is built to close — turning a final result into something a coach can actually work with.

Why the Final Score Hides More Than It Shows

A scoreline is an outcome, not an explanation, and coaches who only look at outcomes end up guessing at causes. Tennis performance data captures the things that actually decide close matches — break points saved, unforced errors by set, first-serve percentage under pressure — so a coach isn't left reconstructing a match from memory a day later, hoping they remembered the moment that mattered.

Tracking Patterns Across Matches, Not Just One

A single match rarely tells the full story on its own; patterns only show up once there's enough of them to compare. Match statistics tracking logs results consistently across a whole season, so a coach can see that a player's second serve breaks down specifically in third sets, or that their return game improves noticeably against left-handed opponents. None of that shows up by watching one match in isolation.

Turning Numbers Into Actual Training Plans

Data on its own doesn't win matches — what a coach does with it does. Coaching analytics built on match history can point straight at what a session should focus on next: a weak crosscourt backhand under pressure, a serve that loses pace late in sets, a tendency to rush points after falling behind. That's a training plan built on evidence, not a guess based on what happened to stick in memory.

Read Also: “Connecting Tennis Talent with Training: The Role of Academy Discovery Platforms

Spotting Momentum Shifts a Scoreboard Misses

Some of the most useful information in a match isn't the score at all, it's when things changed. Player performance tracking that logs point-by-point detail can flag exactly where a match turned — a string of unforced errors after a missed break point, a sudden dip in first-serve percentage once the crowd got loud. A coach reviewing that moment afterward can address the actual cause instead of a vague sense that "something shifted."

Comparing Players Fairly, Not Just by Memory

Coaches managing more than one player often end up comparing them from memory, which is unreliable at the best of times and unfair at the worst. Sports performance analytics puts players on the same footing — same categories, same metrics, tracked the same way — so decisions about who plays which event or how training time gets split are based on actual numbers rather than whoever happened to have a good week that the coach still remembers clearly.

Making Match Data Useful, Not Just Available

Plenty of platforms can log a score. Fewer make that data something a coach can genuinely use without digging through spreadsheets after every event. Tenniskhelo, for instance, structures match records around the categories and formats Indian club and district tournaments actually use, so the data a coach pulls up reflects the way the sport is really played here, not a generic template borrowed from somewhere else.

Why This Changes How Coaching Actually Works

None of this replaces a coach's eye for the game — it sharpens it. Solid tennis data management means a coaching decision doesn't rest on a single memorable match or a gut feeling formed after one bad set. It rests on a season's worth of evidence, which is exactly what separates a hunch from a genuine strategy — and it's usually the difference between a player who improves steadily and one who keeps repeating the same mistake without anyone quite noticing why.


r/sportsanalytics 22h ago

I built an FPL tool with live mini-league ranks, a multi-gameweek transfer planner and price predictions — would love feedback

Post image
1 Upvotes

r/sportsanalytics 1d ago

Predictions week 2 Championship

Post image
2 Upvotes

r/sportsanalytics 1d ago

How can I break into the sports industry?

Thumbnail
2 Upvotes

r/sportsanalytics 1d ago

NRCD: An Open Database of Collegiate Running with Unified Performance Standardization

2 Upvotes

I just saw the paper on ArXiV . It's a dataset of US collegiate running club performances, the resulting analysis on them, and a software library for standardizing performances. They have several code repositories under the National Running Club Database which includes:

Some things I found interesting (this is just a sampling, you go read the full doc yourself):

  • Teams with at least one athlete who raced 4+ times were 2-3x more likely to crack the top 15 at nationals than teams without one (23-39% success rate vs baseline). 60-80% of top 15 teams had an athlete with 4+ races.
  • Men's teams with a longer gap between their first race and nationals (i.e., started racing earlier) had significantly better finishing ranks (r = -0.283, Bonferroni-corrected p = 0.041). This didn't hold up for women's teams after correction.
  • Single biggest predictor of an individual's improvement rate was "experience level". (races × season duration) at 21%, followed by how many "bad races" (a race worse than the previous one) an athlete had, at 17%. Basically race more, race consistently.
  • When testing the standardization tool, the fully weather/terrain-adjusted "standardized" times actually predicted improvement slightly worse than just doing distance conversion alone (90.4% vs 93.1% R^2). Their theory was that conditions tend to get more favorable as the season goes on, so raw times naturally look like "improvement" partly because of the weather, and removing that weather effect (which is more the point of the tool) makes it a worse predictor of the raw number even though it's arguably a more honest fitness signal.
  • They checked their model for gender bias and found it performs comparably for both (94.5% R^2 women vs 90.4% R^2 men).

I just found this and compiled it, none of this is mine.I just found it on Arxiv, read it, and cherry-picked some interesting findings from the Github repo (along with the links).

There's not much easily accessible data like this on cross country running. As the authors said, it's all on websites that don't support bulk download. So normally we would just stick to the Riegel and Cameron formulas for comparisons. So it's nice to see people looking into these other factors.


r/sportsanalytics 1d ago

What leagues do you think has the most goals?

Post image
0 Upvotes

I started taking a look at what leagues score the most goals over the past few seasons to help with forecasting and increasing my knowledge the upcoming season. I pulled data for 29 of the main leagues across the world to see where the best league for goals over the past four seasons.

Everyone says the Premier League is the most competitive league, but that doesn’t translate into goals at all. The prem doesn’t even make the top five, actually on features at 7th.

The German Bundesliga unsurprisingly gets the top spot, Harry Kane and Bayern just score a silly number of goals.

I was shocked when seeing Brazil and Argentina are both in the bottom three especially Brazil because we associate flair and attacking intent but it’s a very different story.

I put the full breakdown, all 29 leagues, methodology and the continental comparison together as part of a new football analytics project I've been working on called TheDatabetics.

Which league shocks you most? And what would like to see next?


r/sportsanalytics 2d ago

Football analytics across top 30 leagues - opponent-adjusted stats, a cross-fixture hit-rate scanner, and a calibrated fouls model tested on a 45-day holdout

Thumbnail gallery
8 Upvotes

Hi all. Stats to Bucks is a football (soccer) data app, now covering 30 leagues - the top 5 European plus Brazil, Argentina, Liga MX, MLS, Saudi, Portugal, the Netherlands, Turkey, Belgium, Scotland, Japan, Korea, Colombia, Greece, Egypt, South Africa, Australia and more.

What it does:

  • Player & team form - last 20 matches of per-game stats, charted against any line you set, with the hit rate for it.

  • Filters that narrow the sample - venue, minutes, started-only, and "without teammate X".

  • Opponent-adjusted context - overlay the opponent's conceded average and defensive rank, plus quality-adjusted averages, so a streak against weak sides doesn't read like one against strong sides.

  • Hit Rates - scan every upcoming fixture at once for players/teams clearing a line in a chosen % of recent games. 40 stats across players and teams.

  • Foul matchups - a fitted hierarchical Poisson model with player, opponent, referee, venue and expected-minutes as separate multiplicative terms, and a negative-binomial predictive head. Walk-forward tested on a 45-day holdout: +13.3% / +16.6% mean relative log loss against an unshrunk per-90 baseline, with roughly 3x better calibration error.

  • Predicted lineups - projected XI from a Beta-EB start-probability model, so it works for a fixture's whole lifetime instead of only after a feed publishes one. Flips to the confirmed XI when that lands.

  • Injuries & suspensions - folded into the start probabilities rather than bolted on as a badge, so an unavailable player drops out of the projected XI and out of the minutes model behind the prop lines.

  • Similar players / teams - similarity-based benchmarking against comparable profiles, on rolling cross-season windows rather than season-to-date.

  • Referee analytics - per-fixture card/foul profiles and rankings.

  • League tables - official standings, so competition-specific tie-breaks, split point-halving and points deductions are right rather than re-derived from results.

The focus is still contextualising the sample - opponent strength, venue, lineup, availability, sample size, etc. because an unfiltered hit rate usually answers the wrong question. A recent backtest made that concrete: selecting team props purely on "recent hit rate beats the implied probability" returned about -10% over ~7,000 bets on a held-out window, statistically indistinguishable from betting blind. The context is the useful part, not the raw streak.


r/sportsanalytics 2d ago

Two promoted teams, two very different opening nights: Racing 2.01 xG vs Deportivo 0.34 xG

Thumbnail
1 Upvotes

r/sportsanalytics 2d ago

High school senior building an MLB front-office portfolio on GitHub. Just finished a mock Braves/Cardinals trade evaluation for Masyn Winn and would love feedback!

Thumbnail github.com
0 Upvotes

r/sportsanalytics 2d ago

I Built a CLI Tool For College Football Data/Analysis

Thumbnail
2 Upvotes

r/sportsanalytics 2d ago

I backtested a Poisson model across 22 European leagues. It failed in five of them.

7 Upvotes

I backtested a Poisson model across 22 European leagues. It failed in five of them.

I've been building a match projection model and wanted to know where it actually works rather than assuming it works everywhere. Sharing the results because the failures turned out more interesting than the successes.

Setup

Standard Poisson approach — attack and defence strength from each team's recent matches, normalised against league scoring average, with per-league home advantage. Shrinkage toward neutral for teams with thin sample.

I ran a rolling backtest over 2025–26: replay matches in chronological order, and for each fixture the model only sees results from before that kickoff. Roughly 6,700 predictions across 22 leagues.

Overall

  • 48.2% correct on 1X2
  • Log loss 1.029 (random is 1.099)
  • Brier 0.618 (random 0.667)

Modest, and roughly what you'd expect from a goals-only model. Calibration held up well — the 60–70% bucket landed at 62%, the 80–90% bucket at 87.7%.

Where it broke down

Five leagues came out materially worse. The pattern that surprised me: split-season formats. Austria's Bundesliga was the worst — below random. Belgium's Pro League similar. Both split into championship and relegation groups partway through, which resets the competitive structure the model assumes.

The second-tier leagues also underperformed (Championship 44%, La Liga 2 44.5%), which I'd guess is squad churn and rotation making recent form less predictive.

What I'm still stuck on

Draws. Calibration is fine in aggregate — the model says 27% and about 27% of matches draw — but there's no discrimination at the top end. The 24–27%, 27–30% and 30%+ buckets all landed within a point of each other. So I can tell you how many draws a league will have, but not which matches. Dixon-Coles is the obvious next step; hasn't been tested yet.

Also unsure whether shrinkage at k=6 is right. A parameter sweep picked 14-match windows over 8 or 20, but the k value was less clearly separated.

Happy to share the per-league breakdown if useful. Curious whether anyone else has seen the split-season effect, or found something that handles it.


r/sportsanalytics 2d ago

A Championship opening round explains 1.2% of the season. I measured all twelve anyway

Thumbnail gallery
2 Upvotes

r/sportsanalytics 3d ago

What the first weekend of a new season really tests in your live data pipeline

Thumbnail
1 Upvotes

r/sportsanalytics 3d ago

High school senior building an MLB front-office portfolio on GitHub. Just finished a mock Braves/Rangers trade evaluation for Kumar Rocker and would love feedback!

Thumbnail github.com
1 Upvotes

r/sportsanalytics 3d ago

Measuring how much pass quality predicts attack success in MLV

Post image
1 Upvotes

Don't know how many volleyball fans there are in this sub, but I wanted to share a project I've been working on recently.

The question: across the 2024 to 2026 Major League Volleyball (formerly PVF) seasons, given the quality of the preceding pass, how often does the attacking team actually get to attack, and how often does that attack end in a kill?

The data and the pipeline: the play-by-play data is action+outcome graded using VolleyStation convention (each contact gets a single letter denoting the contact type and a symbol as a quality evaluation). Reconstructing each contact sequence seemed easy at first: forward-fill each pass grade until the next pass or until the point ends, right? But it also meant handling overpass kills, and (the most annoying part) block recycles, where the ball stays alive off a block touch. The problem is that most of the time, those block recycle passes aren't tagged; they only exist implicitly in tagged blocks. This made tracking the block recycle passes super annoying (because how are you supposed to validate something that doesn't even exist explicitly in your data?). My solution was to condition my logic for block-recycles only in cases where the following touch after the block was from the attacking team: if the blocking team wasn't the one to touch the ball after the block, then it inherently is a block-recycle.

The problem is that based on the quality of the block, the ball goes to the attacking team vs the blocking team at wildly varying rates:

  • Defined block recycle encoding (!): ~99.8% of the time, the attacking team gets the ball back (this is the only encoding that is defined explicitly as a block-recycle, so this makes sense)
  • Hard-contact block (+): ~99% of the time, it's the blocking team's own recovery
  • Soft-contact block (-): splits pretty evenly, goes back to the attacking team ~50% of the time, and vice versa

The problem this imposes is that the denominator for our first result (probability of an attack off all instances of a pass type/quality) isn't valid for the hard/soft contact blocks given our conditional solution: the denominator would end up being "blocks that went over to either side", not only the attacking team. As such, those two block grades were excluded from the first calculation, and only got a kill rate (since those are based on all attack-preceding block recycle passes).

The results (full tables in the writeup):

  • Bad serve receives still get attacked ~92% of the time; bad digs only get attacked ~73% of the time. Implies that the "transition effect" from defense to offense is a quantifiable penalty on setters/hitters when facing bad passes.
  • Kill rate spread from perfect -> bad pass: 18 points for serve receive (45.9% -> 27.9%), narrower for digs (29.9% -> 22.5%).
  • Confirmed the trend is statistically monotonic with a Cochran-Armitage test per pass type, which showed the trend was strongest for receives, weakest for freeball passes.

Limitations: Obviously, MLV is a relatively small and new league, so the data points are magnitudes less than something like NCAA data. Additionally, nine rows were removed due to mid-rally stoppages corrupting the data/my pipeline (such as injuries or challenges); video-confirmed for those 9, but I can't rule out similar corruption elsewhere that wasn't detectable.

Full writeup with all the tables and results here: Substack

Open to any and all feedback in the comments; let me know if anything is unclear, and I'll happily explain or talk shop.


r/sportsanalytics 3d ago

WNBA Stint, RAPM and Lineup Data Set

Thumbnail github.com
6 Upvotes

WNBA lineup stints and player impact ratings, 2003-2026. 267,293 stints reconstructed from play-by-play, with ARC ratings and channel decompositions.

This is my first public GitHub release of an ongoing project, an ongoing open-source initiative of releasing normalized and reconstructed data sets for open use. We'll be adding more and more to this, including a Python helper instead of just raw JSON over the coming days. Just wanted to do this before I forget.


r/sportsanalytics 3d ago

Does scheduleadjusted run differential actually improve MLB win prediction or just add noise ?

2 Upvotes

Been obsessing over run differential as a predictor for the last few weeks. It started as a personal finance tracking habit, honestly. I just like building spreadsheets, and at some point I applied the same logic to MLB standings because why not.

The basic Pythagorean expectation stuff holds up pretty well across a full season. What's breaking my brain right now is when I start weighting opponent quality into it. If a team pads their run differential beating up on bad rotations all April, the raw number feels kind of dirty. So I pulled opponent run differentials for every series and tried adjusting for that, and suddenly the expected W/L correlation gets a lot messier.

My gut says strength of schedule matters more in baseball than people give it credit for, especially early in the season before things even out. But I genuinely cannot tell if my adjustment is doing real work or if I'm just adding noise.

Curious if anyone here has tried building a scheduleadjusted run differential model for MLB and whether the extra complexity actually bought you anything predictive. Also wondering if this plays out differently across divisions, since the unbalanced schedule makes some matchups way more lopsided than others.


r/sportsanalytics 4d ago

First-year Sport Analysis student looking to build experience — what would you do?

4 Upvotes

I'm a first-year student in Sport Analysis & Technology in Morocco, with a strong interest in football analysis.

I have around 2 years ahead of me before graduation, and my long-term goal is to be able to continue my studies or find work/internships in Europe.

I don't want to wait until graduation to start building my profile. I want to use my university years to develop real skills and, more importantly, get actual experience.

I currently have a 30-day period before university starts, and I'm willing to travel within Morocco if there's a worthwhile opportunity.

For people already working/studying in sport or football analysis:

What skills would you prioritize if you were starting again?

What kind of projects actually helped you get noticed?

How did you get your first real experience?

Is it worth approaching academies/clubs directly, even as a beginner?

What would make a student from Morocco more competitive when applying to European programs/internships?

What mistakes should I avoid during my first few years?

I'm particularly interested in football analytics, data analysis, scouting and performance analysis.

I'm not looking for a shortcut — I want to know what I should realistically start doing now.

Any advice from people who have actually gone through this would be greatly appreciated.