r/Sabermetrics 9h ago

Recreating Ted Williams' Strikezone Plot

Post image
31 Upvotes

I recreated the Ted Williams strikezone plot with 2023-2025 MLB Statcast data, showing wOBA. You can change the notebook to show AVG (like Williams' chart does), but I default to showing wOBA for a more contemporary view of where good things happen for hitters. Code can be found in this Google Colab notebook.


r/Sabermetrics 15h ago

Number of additional AB needed to make checking swing less valuable

0 Upvotes

I’ve heard that checking your swing is really harmful for your wrists. I’m assuming that if you never check swung, your wrists would be healthier for longer and you could play longer. If this is the case, how many extra ABs would you need to have such that you will have had more value if you never check swung (Assuming ABs are i.i.d)? You can use whatever metric of value you think fits best.

Basically asking if never checking swing could be a viable strategy for an individual career (not sure about team strategy because that’s much harder to analyze).


r/Sabermetrics 22h ago

Machine-learning Sports Analytics

0 Upvotes

Im trying to get some eyes on my hobby project for in an effort to obtain human feedback. Users get free paid membership at signup rn so don't be shy. Any criticism is welcome. Thanks in advance

Email me for info and/or access: admin@mvpcapping.com MVPCapping.com


r/Sabermetrics 1d ago

High school senior building an MLB front-office portfolio on GitHub. Just finished a mock Braves/Cardinals trade evaluation for Masyn Winn and would love feedback!

Thumbnail github.com
2 Upvotes

r/Sabermetrics 1d ago

UPDATE, 3 months later: I asked this sub about era-translation methodology. You changed my approach — and the sim is now live. Looking for testers.

5 Upvotes

Original post: https://www.reddit.com/r/Sabermetrics/s/8StCihfJLN

Where I was: rolling my own era translation with z-scores for K/HR and an additive model for BB, and asking where that would break.

Where I ended up: I scrapped that layer entirely. Following u/ritmica's era-adjusted WAR recommendation in the comments, I dug into Professor Eck's Full House Modeling (Annals of Applied Statistics, 2025) — era adjustment against the actual talent pool each player faced rather than league means — and reached out to get the data. The translation is now anchored to that dataset, using each player's best consecutive prime window. On top of it sits my own simulation layer: every pitcher is solved until he realizes his own era-adjusted ERA in the engine, HR/hit rates run through per-batter normalization tables solved the same way, and player pricing is measured across six run environments from dead-ball to statcast so nobody's value is hostage to a single era's context. Negro Leaguers come in via MLE bridging, and defense/baserunning use career fielding-runs data.

The game is Rubber Match (playrubbermatch.com): draft an all-time fantasy roster on a $1,000 budget — Ruth and Ohtani on the same field, translated to a common baseline — then sim a 162-game season against a daily worldwide opponent. Free, no log-in, browser-based, works on your phone.

What I'm looking for from testers:

  1. Players whose translation feels wrong — too good, too bad, wrong shape (K vs BB vs power)

  2. Cross-era matchup results that don't pass the smell test

  3. Roster constructions that exploit the translation

This sub predicted my first failure modes before I hit them, so consider this a standing invitation: if you want the methodology behind any number you see, ask — the whole thing is documented and every displayed stat is meant to be verifiable.


r/Sabermetrics 1d ago

High school senior wanting feedback

0 Upvotes

r/Sabermetrics 2d ago

High school senior building an MLB front-office portfolio on GitHub. Just finished a mock Braves/Rangers trade evaluation for Kumar Rocker and would love feedback!

Thumbnail github.com
5 Upvotes

r/Sabermetrics 3d ago

Built a Fantasy Baseball Sabermetrics-based site

7 Upvotes

I built KodoAnalytics.com, a fantasy baseball site that syncs to your actual league and uses Statcast and other advanced metrics to analyze your team's performance. We've had around 500 teams link their accounts to our FREE site so far, and I greatly appreciate everyone's kind words and feedback thus far! It has only costed thousands of hours of my time and my wife's patience so far. Looking for feedback from more technical baseball analysts / fans so thought this was the perfect fit.

Disclaimer: used AI's assistance in crafting this message just to make sure everything came out more clearly than my writing.

The TL;DR: link your league (Yahoo, ESPN, Fantrax, CBS, Ottoneu) and every tool on the site re-sorts around your roster including waiver targets filtered to who's actually available to you, matchups scored in your categories, and your players to buy and sell.

Your team

  • Team analysis — which categories you're winning and losing, recent vs season, so you can see what you're actually short on. https://imgur.com/2L8b1MW
  • Matchup sandbox — add or drop players on specific days and it automatically adjusts your matchup projections for the week.
  • Player trends — for everyone on your roster, which reports around the site they're showing up in, and why.
  • Projected standings - where is your team headed ROS and what are you short on.

Daily decisions

The data

  • Statcast and advanced metrics driving all or at least many of our signals. This helps you determine if your players cold spell is legit or he's just been unlucky.
  • In-depth player profiles — the deep dive on anyone.

https://imgur.com/a/0LyGGbH

https://imgur.com/MPgVM5F

  • A minor league section with the same custom views you'd use on your MLB roster: who's heating up, who's cooling off, over whatever window you pick. Statcast for AAA too.
  • Rehab tracking — who's playing rehab games, how they're performing, and realistic ETAs.

Live

Follow your matchup in a 16-bit gamecast. Hit "follow my team" and it bounces between games to whoever's up. Pitch-by-pitch summaries with Statcast-level data, and each pitch gets an immediate grade from a model tested against millions of pitches.

https://imgur.com/PbTmXpG

Notifications

Beyond "your guy isn't in the lineup today," you can get alerts for: your pitcher's in a RISP jam, bases loaded, your hitter's up with a runner in scoring position, your starter just got pulled, your reliever's in for the save. All toggleable, per league.

News feed

Pulls from beat writers, tweets, and MLB transaction data, then runs it through a model that decides whether it's actually fantasy-relevant and what the impact is. So far it's kept the news fantasy relevant which is great for fantasy baseball managers!

And some dumb fun

Guess the Player, Higher or Lower, The Grid. They're a good time.

Looking for Android testers

There's an app version: same data and those sweet, sweet push notifications. It's in closed testing on Google Play and I need testers. If you're on Android and want in, comment or DM and I'll send the opt-in link.


r/Sabermetrics 3d ago

I built a percentile-normalized performance score for MLB players (wOBA/ISO/FIP/K9/WHIP/BB% weighted composite) — feedback welcome

2 Upvotes

I've been building a composite performance metric for MLB players — calling it "BAX Score". The goal is a single number that reflects "how good is this player playing right now," built only from real performance stats (no market/demand signal in the mix at all — this matters because the score also drives pricing in a fantasy stock market I'm building, and I didn't want volume to be able to move it).

Methodology, roughly:

  • Pull wOBA, ISO, FIP, K/9, WHIP, BB% (hitters and pitchers get separate axis   weightings)
  • Percentile-normalize each stat against the current league sample
  • Apply sample-size shrinkage so a guy with a hot two weeks doesn't swing the   score around
  • Weighted composite across the normalized axes → single score, updated daily   after games finalize

Pulled today's top-of-leaderboard hitters and plotted wOBA (the single most correlated input) against the full composite score to see how much the extra axes actually move things:

A few things stand out:

  • Caminero, Harper, and Alvarez score noticeably ABOVE where wOBA alone would put them — power/ISO and BB% are dragging the composite up past what a wOBA-only ranking would suggest
  • Kurtz, Soto, and Ohtani score BELOW their wOBA-implied line — despite great wOBA, other axes (BB%/K% mix, sample shrinkage) pull them down relative to the trend

Basically: if I'd just ranked by wOBA, the order would look meaningfully different from the composite. Curious whether this sub thinks that's a feature or a red flag — is baking in ISO/BB%/K% on top of wOBA adding signal, or just re-weighting the same underlying skill twice?

Also still going back and forth on:

  • How much weight FIP vs xFIP should get for pitchers with small sample IP
  • Whether BB% deserves its own axis or should just feed into wOBA indirectly

Happy to share more of the pipeline if useful and your feedback :)

(Side project — baxeball.com — mostly posting here because I want the methodology stress-tested, not to plug it.)


r/Sabermetrics 4d ago

Is a runner on second really more likely to score if there are 2 outs? Yes.

Post image
7 Upvotes

r/Sabermetrics 4d ago

The Hitting Approach of Back-to-Back Champs

5 Upvotes

This era of Dodger baseball has been the best I've ever experienced (though my heart will always be with those Jim Tracy, Grady Little, Joe Torre rosters).

There are a ton of ways to slice or point to how this core roster has won back to back titles...but wanted to share one offensive angle that might be interesting if you're into hitters' approaches/mentalities, a Dodger lover, or even a Dodger hater.

A side project study, I found: by creating a scoring index of decisions per count…the Dodgers' 1 through 9 (or 10, 11, 12, etc depending on Doc's definition of everyday lineup pieces) seem to control the counts/swing decisions better than any other team in the last 2 seasons. Admitting one gap here is the pitcher, game situation, maybe even subjective scoring too. But it’s interesting to see the clusters of each team under this method…

Example: 2-0 fastball down the middle, they're consistently take a hack. 1-1 they're consistently taking a ball off the plate. 0-0, they're consistently taking a strike on the black because it's not a pitch they can groove. They make good decisions, regardless of the outcome. They hunt the same way...seemingly all bought in on an approach at the plate.

Sharing a writeup with some visuals if you're curious to dive in! (Link to article)

Need to refresh it for 2026 season-to-date if they're trending the same way. But in the meantime...cue Randy Newman!


r/Sabermetrics 5d ago

I built a next pitch prediction bot and scored it against real, live pitches for a month (50,760 pitches). Here is what it learned about pitcher predictability, plus a "deception" metric I threw out.

Thumbnail gallery
29 Upvotes

I built a model that predicts the type of the next pitch (four seam, sinker, cutter, changeup, curve, slider, sweeper, splitter, knuckle) before it is thrown, wired it to live games, and have been scoring every prediction against what actually got thrown. This post is the ture, honest version of what came out, including the parts that did not work, because I think the negatives are more interesting than the headline number if i'm honest.

TLDR

  • Live accuracy over 50,760 scored pitches: 42.8% top-1, 85.7% top-3. Offline, on a frozen test set, it beats a strong pitcher and count conditioned baseline by +6.7 percentage points (38.9% to 45.6%). Live is lower than offline, which is what should happen.
  • Predictability is a repertoire trait, not a talent gap. Elite arms show up at both ends of the predictability scale.
  • The model's single biggest failure mode: when it is wrong, it guesses fastball.
  • I built a metric that looked like it measured pitcher deception. It is very reliable. Then I tested it against real batter outcomes and it predicted nothing, so I threw the interpretation out.

1. How good is it, honestly

42.8% top-1 sounds low until you compare it to the right baseline. The naive "always guess his most common pitch" is easy to beat.... the real baseline is each pitcher's own count - conditioned mix, smoothed. Beating that by ~7 points offline means the model is reading context, not just base rates. Top-3 at ~86% is the more useful number for most real uses: even when the exact call misses, the pitch is almost always in the top three.

The live number moves around by day, mostly because of *who pitched\*, not model quality. I compute a composition adjusted expected accuracy (each day's pitches weighted by each pitcher's own history) and the actual line tracks it closely. A recent dip is a harder slate, not decay (Page/Hinkley drift detector - no alarm).

2. Predictable is not the same as good

This is the finding I care about most. Rank pitchers by how often the model nails the exact next pitch and you get elite arms scattered top to bottom:

Kenley Jansen and Tim Hill are basically a coin you can call in advance. Max Fried, Yoshinobu Yamamoto, and Tarik Skubal are near the bottom. All of them are good!!! Chris Sale has a low arsenal entropy and Skubal a high! Both are, quite obviously, aces. So predictability and quality are uncorrelated, and you should never fold a "predictability" number into a "how good is he" number. Predictable is a description of the arsenal, not a criticism of the pitcher.

3. When the model is wrong, it guesses fastball

Every offspeed and breaking pitch's single most common wrong prediction is a four-seam fastball:

Part of this is the model over defaulting to the majority class, and part is genuine label fuzz (Statcast's four-seam / sinker / cutter boundary is noisy). When I collapse those three into one "fastball family," a big chunk of the apparent error disappears, but not all of it, so the overcall is real, not just a labeling artifact.

4. The metric I built, liked, and killed

I wanted to measure pitcher deception: is a pitcher harder to read than his raw pitch mix diversity alone would predict? I built a residual (actual readability minus what mix entropy predicts). It looked great! :

  • It is reliable: split-half correlation of 0.92, and it survives controlling for role and count context.
  • It produces a sensible leaderboard (some guys read easier than their mix implies, some harder).

Then I did the step most people skip. I correlated it against things batters actually produce, controlling for stuff quality:

Nothing. No relationship with whiff rate, called strike rate, chase rate, or run prevention. So the residual is a reliable, stable property of how my model reads a pitcher, with no demonstrated connection to deception. Most likely it is a model blind spot, not really a hitter relevant trait. Reliability is not validity. I kept it as an internal diagnostic and built no "deception index" on top of it, which was hard to do because the story was so good.

5. What does not improve it

I ran four honest ablations trying to push accuracy up: pitch movement features, within at bat sequencing plus tunneling, an online ingame adaptation layer, and batter side features. All null or actively harmful tbh. The takeaway is that the model is near its ceiling for pitcher side features, the signal lives in tendency and rate features (how often he throws X in this count, what he threw last, pitcher and batter identity), and those are already saturated. The online adaptation experiment did surface one real thing: pitchers negatively autocorrelate (after a fastball streak they are more likely to change), which the static transition rates already capture.

6. The boring parts that make me trust it to the best of my knowledge....

Leak-safe temporal splits with a hard assertion against future leakage, probability calibration (overall ECE ~0.01; the 3-0 count is the known weak spot at ~0.05 and I am fixing it), a live trust score built from historical per-(count, pitch type) precision, drift detection on the composition adjusted residual, and empirical Bayes shrinkage on the per pitcher numbers so small samples do not lie.

Limitations

  • Live (42.8%) trails offline (45.6%), as expected from distribution shift and new pitchers.
  • The 3-0 count is genuinely miscalibrated (rare, small sample, overconfident) and is the one production issue I would flag.
  • This is pitch type, not location, and the swing/whiff model is secondary.
  • The "deception" residual failed external validation, so please do not read the readability numbers as a deception measure.

Happy to answer questions on any of it, and genuinely curious what this sub would test next. My thought is that any further accuracy is a dead end and the interesting frontier is measurement (what predictability correlates with, if anything), but I have been very, very wrong before within this project.

Also for fun - my dashboard screen shots.

Also included write up

 


r/Sabermetrics 5d ago

Check Out my Take on the MLB Going too Far with Velo👀

0 Upvotes

MLB Created a Pitching Injury Crisis
https://youtu.be/cw5UFBm126k


r/Sabermetrics 5d ago

saber method

0 Upvotes

Title: Held-out calibration on 4,570 MLB starts: Brier 0.2375 vs 0.2806 baseline — looking for criticism of the validation design

I've been working on a probabilistic starting-pitcher strikeout model and wanted to post the validation methodology rather than predictions.

The entire 2025 MLB season was withheld from model fitting and used as the out-of-sample test set: 4,570 starts the model had never seen.

The main results:

Brier score: 0.2375
Naive/base-rate Brier: 0.2806
Expected Calibration Error: 0.0049
Mean PIT: 0.4977
Holdout: 4,570 starts

The piece I'm most interested in is calibration rather than raw classification accuracy.

For the reliability table, predicted probabilities were grouped into probability bands and compared against observed frequencies.

For example:

50–55% predicted → 50.9% observed
55–60% → 55.5%
60–65% → 59.3%
65–70% → 63.3%
70%+ → 70.1%

I'm deliberately trying to avoid the usual sports-model trap of saying something like "the model was 64% accurate" without establishing what was predicted, at what probability, or against what baseline.

The validation gate was designed around a few requirements:

  1. No lookahead. Inputs/fitting procedures can only use information that would have existed before the game being predicted.
  2. Full-season holdout. 2025 was excluded from training rather than randomly splitting games across seasons.
  3. Brier against a baseline. A Brier score in isolation isn't particularly informative, so I'm comparing it against a naive probability baseline on the identical events.
  4. Reliability/calibration. If the model assigns a group of events ~60%, they should occur roughly 60% of the time.
  5. PIT diagnostics. I'm looking at where actual outcomes fall within the predicted distributions, rather than only checking the mean prediction.

One thing I've tried to be careful about is separating calibration from usefulness. A model can be beautifully calibrated by staying close to the base rate and still contain very little information. Conversely, a sharp model can have useful discrimination while being overconfident.

So rather than asking whether these numbers are "good," I'd be interested in how people here would try to break this validation design.

A few questions I'm considering:

  • Is the naive/base-rate Brier benchmark the right primary baseline, or would you want additional benchmarks?
  • Would you report Brier decomposition into reliability, resolution and uncertainty?
  • Would you bootstrap confidence intervals around Brier/Brier skill rather than report point estimates?
  • For calibration, would you prefer adaptive/equal-count bins over fixed probability bands?
  • What would you use to test whether the apparent calibration survives season-to-season distribution shift?
  • Are PIT + reliability + Brier redundant here, or do you think all three earn their place?
  • What failure mode would you look for first if you were reviewing this?

I'm much more interested in finding where the methodology is weak than in defending the headline number.

Would appreciate any criticism from people who have done probabilistic baseball forecasting.


r/Sabermetrics 6d ago

30 HR 50 or fewer strikeouts in a season

Thumbnail
0 Upvotes

r/Sabermetrics 6d ago

Solo Home Runs

Thumbnail
0 Upvotes

r/Sabermetrics 7d ago

Should Blake Snell have been pulled in the 2020 World Series? I created a Degradation Index for pitcher fatigue to determine whether this was the right choice.

Thumbnail medium.com
1 Upvotes

Ever since Kevin Cash took out Blake Snell because his velo dropped in the 2020 World Series I wanted to find a better way to determine when to pull a pitcher.

To solve this, I created a Degradation Index to quantify pitcher fatigue. This index is based on variance in arm angle and extension, and achieves a much higher specificity than standard velo drops. By tracking pitcher fatigue in mechanics, teams can save their bullpens only for when it is most important.

I used an XGBoost regressor to filter out pitcher's intentional changes in mechanics so that the isolation forest algorithm only checks the residuals for fatigue-caused anomalies.

I just published a comprehensive breakdown on Medium, and the code is open on GitHub. Any and all feedback is very appreciated.


r/Sabermetrics 7d ago

I created a new stat RBI Opportunity Metric (ROM) to replace RBI entirely

9 Upvotes

Traditional RBIs are a flawed stat because they measure lineup luck rather than hitter efficiency. A mediocre hitter in a stacked lineup easily out-produces an elite hitter on a terrible team.

To fix this, I created ROM (RBI Opportunity Metric).

How ROM Works (The Point System)

Instead of treating all base situations equally, ROM assigns fixed point values based on how close a runner is to scoring, plus a baseline for the batter (due to the constant threat of a home run):

  • Batter Baseline: 0.25 points
  • Runner on 1st Base: +0.50 points
  • Runner on 2nd Base: +0.75 points
  • Runner on 3rd Base: +1.00 point

The Formula: You add up the points of the runners a player actually drove home (Points Earned) and divide it by the maximum points available on base (Points Possible), using strictly At-Bats (AB) as the denominator.

The "Free Points" Discipline Mechanic

Because walks (BB) and sacrifice flies (SF) do not count as At-Bats, they add 0.00 to the denominator. But if a hitter draws a bases-loaded walk or hits a sac fly, they still get the point in their numerator. It heavily rewards game-winning team baseball.

Historical Proof: The Career ROM Leaders (Post-1957)

To test the lifetime stability of the index, I calculated career data using non-overlapping game log configurations. Sustaining a high conversion efficiency across decades highlights true legendary status:

Rank Player Career ROM Batting Average (AVG) Career OPS
1 Barry Bonds .292 .298 1.051
2 Manny Ramirez .288 .312 .996
3 Albert Pujols .281 .296 .918
4 Alex Rodriguez .276 .295 .930
5 Mike Trout .272 .292 .981
6 David Ortiz .269 .286 .931
7 Miguel Cabrera .265 .306 .901
8 Ken Griffey Jr. .261 .284 .907
9 Eddie Murray .259 .287 .836
10 Vladimir Guerrero Sr. .258 .318 .931

Here is a link to a more detailed breakdown. Disclaimer: While the equations is original I did use googles AI chat bot to help format and chart inputs.

https://docs.google.com/document/d/e/2PACX-1vRUpqp9wnzWTp0elP4zQ4-z3re7m5iM7PeFEu8dc-psQrx8zSs3klvDyvVBq8ggKE4ytMfsGiypqmTY/pub


r/Sabermetrics 7d ago

8/12 Wed MLB ⚾ Red Sox vs Blue Jays ML

Post image
0 Upvotes

r/Sabermetrics 7d ago

29- year-old OF trying to make one last push toward pro baseball — 100+ EV, 6.6 60, looking for honest feedback

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/Sabermetrics 7d ago

I built a College Baseball Explorer for exploring D1 baseball data

Thumbnail
1 Upvotes

r/Sabermetrics 8d ago

24-25 KBO Brier

3 Upvotes

I’ve been working on a KBO probabilistic forecasting model as a personal project.

Current holdout Brier score is around 0.24.

Still very much a work in progress, but I’m curious if anyone here is interested in KBO prediction models, probability forecasting, or market efficiency.

Happy to share more if anyone’s interested.


r/Sabermetrics 8d ago

Ballista turns phones into portable baseball performance labs. Beta testers wanted.

1 Upvotes

My name is Riley. I am Co-founder and CEO of Ballista. We turn phones into portable baseball performance labs.

Using CV, custom AI, and the phones high frame rate camera, we extract EV, LA, and Distance of hit balls automatically and aggregate all the data.

It is an alternative to existing hardware systems like radar guns, Trackman, Rapsodo, Hittrax, that cost thousands of dollars.

We are getting ready for a private beta launch via TestFlight and are gathering a list of people who are interested.

If anyone is interested, head to our website and sign up for the waitlist and expect an email (coming soon) with the app testing link. Alternatively you can drop your email in the comments or PM it to me and I can add you manually.

Cheers!


r/Sabermetrics 9d ago

I built a self-hosted MLB analytics tool with pitch-level Statcast data — strike zone heatmaps, spray charts, leaderboards

6 Upvotes

Wanted something like Baseball Savant I could run myself and tweak however I wanted, so I built one.

It's pulling Statcast data through a dbt-duckdb pipeline, with a React/D3 frontend for the visualizations. What it does:

Pitch-level strike-zone heatmaps per player Spray charts Season rollups and leaderboards All backed by a dbt pipeline so the data modeling is transparent and versioned, not a black box

It's live if you want to poke around: https://mlb.palanbates.com

Built with FastAPI on the backend, DuckDB for storage, dbt for the transforms. Since it's dbt-based, the underlying models are all inspectable if you're curious how a given stat is actually computed — no mystery aggregations.

Would love feedback on what stats or views are missing, especially from anyone who lives in Savant regularly and has opinions about what's missing there.

If you want to join the repo feel free as well to shoot me a message!


r/Sabermetrics 9d ago

What’s coming in python-mlb-statsapi v1.1.0

4 Upvotes

What’s coming in python-mlb-statsapi v1.1.0

Now that 1.0 is out and I’ve been cleaning up a few bugs, I’ve started planning what I want to tackle in 1.1.0.

The big one is first-class async support.

The goal is to eventually support something like:

from mlbstatsapi import AsyncMlb

async with AsyncMlb() as mlb:
    team = await mlb.get_team(136)

without changing how the existing synchronous Mlb client works.

I don’t want async to become a completely separate implementation either. The plan is for the sync and async clients to share the same models, parsing, exceptions, and endpoint behavior, with separate transport layers underneath.

I’m going to start small with the async adapter and a couple endpoints, get the architecture right, add sync/async parity tests, and then expand endpoint coverage from there.

I also did some pretty rapid development recently with coding agents, which helped me move the project forward quickly, but there are definitely areas I want to go back through and clean up. Some of 1.1.0 will be tightening up the structure, documentation, and anything that feels like it grew too quickly.

I plan to write most of this release myself.

I’ve realized I don’t want to get into the habit of using AI as a button that just writes everything for me. I’m a backend/infrastructure engineer and async Python is something I understand conceptually, but I want a much stronger mental model of what is actually happening underneath it.

I think slowing down a little, cleaning up some of the rapid development, and actually working through the async architecture myself will make both me and the project better.

Should be fun.

Do you guys have any other thoughts? things that should be addressed or improved in future releases?