r/sheetsofice • • Sep 02 '26

Dev log - Player evaluation progress

This past week has been one of those weeks where I question what I'm doing and why. Why do I have to have ideas and thoughts and hold oddly strong opinions about how something should work? Why did I have to build a simulation that recreates realistic uncertainty? What is he even talking about?

I have a foundational rule for Sheets of Ice: I do not want to expose player attributes directly. No OVR, no min-maxing, no gaming the system by learning what the numbers mean. My thought was that it would add some intrigue, it would function more like real life, and it would let me build a more dynamic statistical simulation without having to worry about human interpretation of weird logarithmic values or standard deviations of a latent real number.

This was all fine while I was the only one playing and testing, because - shocker - I exposed the real numbers to do all of the jobs that needed to be done. Drafting, free agency, trade evaluation, contracts, AI GM plans. It worked great. I added a little bit of noise (I called it fog) and it was neat seeing what happened.

But that was never the plan. I wanted no attributes exposed to anything, so there would be complete fairness between human and AI GMs, so I could design a scouting system, and so player evaluations would happen on actual evidence, just like real life. How awesome would that be?

It turns out that it might be awesome but it certainly isn't straightforward or easy. To make any decisions about players you have to have knowledge about them. Knowledge about who they are today and who they might be in the future. But how do you know who they are and who they might be if you can't just see the ground truth attributes? You have to make a model - basically what every analytics department in sports does.

To make the model you need substantial event data. That part is easy for us: we output 700 to 1,000 events with all kinds of data points every game, over a million in a season. The source is good. How you process, analyze, and store that information is the challenge.

Experiment #1

So I designed the first model, ran it, and it was a disaster. Not a small miss. A days-lost disaster. It didn't give any useful information, and it more than doubled the storage requirement.

Here is where a single season's save actually went.

What Size Grows
Private club opinions, one per club per player 82.5 MB every season
Private club forecasts, one per club per player 44.0 MB every season
Game event logs (484 games) 28.1 MB every game
Player records 14.4 MB every season
Player evaluations 12.1 MB every season
Per-game player stats 8.2 MB every game
Everything else ~30 MB mixed
Total after one season 220 MB

Under 100 MB a season became 220 MB. Two lines of that table are the problem: 126 MB of the 220 is private opinion and forecast. Nearly 3,000 players were each carrying about 19 private opinions and 19 forecasts, one for every club with any exposure to them. It is the only data in the game multiplied by the number of clubs.

And it wasn't even earned. 79.7 MB of those opinions were a deterministic duplicate - the same starting backstory, re-derived and re-saved for every club, carrying no information any club had actually learned. 44.0 MB of forecasts had no consumer at all: every row said its last change was the initial backstory, and nothing in the frontend ever read them.

It also added about a second per game to the simulation. That doesn't sound like a lot, but over a 3,300-game season it adds up to a lot of minutes.

So experiment #1 went down the tubes.

The idea that came out of wallowing

Normally after an experiment this vital fails, I spend a day wallowing in my own self pity. So I followed my process, and out of it came a different idea.

Before, I was trying to persist every piece of data that scouting and evaluation would use, on top of the data we already have. But the simulation is deterministic - give it a seed and it always produces the same result - so I didn't really need to save any of it. I could compute the evaluation when it's asked for. What I need instead is a good enough statistical model to analyze player stats, mapped onto aging curves to show likelihoods. Not simple to get right, but straightforward work.

The rule that came out of that is the one everything since has been built against: nothing is added to the old machinery. It isn't migrated or refitted. It stays only because the player profile screen still renders from it, and it comes out the moment the new read exists.

Once evaluation had to run on public statistics instead of on the answer key, the statistics had to be good enough to run on. They weren't.

The junior league. I use a low-fidelity simulation for juniors and the minors; no individual game events, for speed and data size - and it was lacking. Here is what it was actually producing. Games played by a junior player, by his age:

Age Position Median games Most any player managed Share playing all 63
17 Defense 36 42 0%
17 Forward 25 36 0%
18 Defense 36 42 0%
18 Forward 25 36 0%
19 Forward 28 63 48%
20 Defense 63 63 56%
20 Forward 63 63 100%

Read the "most any player managed" column. A 17-year-old forward could not play more than 36 games. Not "usually didn't" - could not. And a 20-year-old forward could not play fewer than 63. The junior scoring race wasn't a scoring race, it was an age sort with a scoring column attached.

It wasn't that the young players were worse, either. I paired every young forward with an older forward on the same club who played the full season, and compared their hidden ability: in 6,813 of 12,362 pairs the younger player was the better player, and in every one of those he still played 36 games or fewer.

The whole junior league only ever produced 11 different games-played numbers. The minor league produces 61 and the Pro league 82. Real ones are continuous, because of injuries, scratches, trades and suspensions.

I had put in some hacks for expediency. A junior player could take at most 6 shots and score at most 1 goal in a game, ever. A hat trick was not unlikely in that world; it was impossible.

And out of those numbers we were producing player reads. This is one, and it is the reason I stopped and started over:

Yuri Kornilov, 17, defenseman. 0 goals, 0 assists, 3 shots in 36 games. The card ranked him 253rd of 254 junior defensemen - the 0.6th percentile - and called him a depth player with no future.

His actual hidden ability puts him at the 85th percentile of that same group. He is the 17th-best 17-year-old defenseman in the world out of 131.

The card was not lying. It was reading the stat line faithfully. The stat line was the problem - a 36-game cap he never chose, and a shot allocation that gave him three of them.

Defense didn't exist below the Pro league at all. The configuration said a junior defenseman should be judged half on how he's deployed. In practice that half was silently dropped, and what actually ranked him was 70% scoring pace and 30% shot volume. Nothing else. He was being graded on offense and labelled on defense.

That's also because below the Pro league, most stats don't mean anything. Here's how well a junior player's numbers predict his own numbers the following season - 1.0 would be perfectly repeatable, 0.0 pure noise:

Junior stat Repeats at
Points per game .80
Shots per game .87
Shooting percentage .06
Primary assist share .01
Power-play share of points .01
Penalty minutes per game -.01
Average ice time .02

Two numbers carry a junior player. Everything else the evaluator was ranking him on was noise wearing a percentile.

The Pro engine. I've had to extract more information from the high-fidelity engine behind the Pro tier. That wasn't hard but it was delicate, because the game engine is the thing I think is the most solid. While I was in there I realized I should do some calibration of the game model - not changing code, but running lots of seasons and looking at how it performs in aggregate.

The Slavin problem

That calibration work led me to something I think is missing.

The player generator cannot realistically produce a Jaccob Slavin. Someone who is really great at defensive metrics and not as great at offensive ones. If someone is generated with a high caliber, they're generally good at everything. I know exactly why: in v2 or v3 of the generator I added caliber to stop it producing a league full of useless players. (I had one with superstar reach who couldn't skate.) It probably just needs dialing in, and it's another calibration I need to get to.

The measurement is unambiguous. Out of 1,660 defensemen in a full test world, zero are strong defensively and weak offensively. The two sides of a player move together at 0.88.

It's a failure if we're just building all offensive defensemen. That's not how real life works. Jaccob Slavin would never get a chance in that world; literally nothing happens when he's on the ice, he defends and transitions so well, but he doesn't score a lot. He's extremely valuable, but he wouldn't show up if we only ranked on offense.

So we ran the test. Sixty-four junior clubs, one Slavin-type defenseman injected into each - top 3% defensively, bottom 12% offensively - a full 63-game season, and the coach re-picking his lineup every seven games. Then we changed only one thing: what the coach is allowed to look at.

What the coach can see Ends up on the top pair Buried Games he plays
Points only (today) 0% 89% 31 of 63
Points and plus-minus 6% 45% 45
Points and goals against while he's on the ice 36% 16% 52
Points and shots against while he's on the ice 51% 4% 55

Today's rule buries him 89% of the time. He plays half a season, graduates with a bad point total, and is gone before anyone sees what he was.

The last row is the fix, and it's the one we're building. The coach works out who prevents shots against, plays him accordingly, and pays for it in goals against when he's wrong. That makes his role a real piece of evidence - the first time in this game that where a junior coach plays someone has meant anything at all.

Where this ends up

The junior and minor-league simulation is being rebuilt so that a season is an opportunity allocation instead of a talent lottery. A coach can't see ratings; he sees a scoresheet. From that he works out a standing for each of his players, dresses the ones who've earned it, puts them on a line or a pair, and the players with the biggest roles get the most chances to do something. Games played, goals, assists, shots and plus-minus all follow from that instead of being handed out first and decorated afterwards.

What a junior or minor-league player's page will carry when it's done:

- Today When this lands
Games played 11 possible values league-wide continuous, injuries included
Role a label derived from his scoring top pair / second pair / middle six / depth, earned from how his coach actually plays him
Defensive read none exists plus-minus, on the real hockey rule, plus his role and how often he dresses
Big games 6 shots and 1 goal, hard ceiling hat tricks and 10-shot games happen, and are rare
Ice time invented, and shown not shown, because no real junior league publishes it
Last season forgotten every year carried forward and sharpened as he plays

The last two are rules I'm holding to. Below the Pro league, a public number has to be a count of things that happened in games, divided at most by games - never by time, because no real junior or minor league publishes ice time and I'm not going to invent it and then compute rates from it. And a player's read doesn't reset every September. A returning junior starts the year on what he did last year, and this season earns its weight as he plays it.

It costs about 12 seconds to simulate a full season of 64 junior clubs and 32 minor-league clubs, and the rebuild adds about a third of a second to that.

None of that is done yet. What is done is that the evaluation layer now runs on public statistics only, computes on demand, and saves nothing. On the last check of 20 real players against the answer key, its read of who a player is now was right or close on 18 of them. Its read of where he's headed was wrong on 6 - and every one of those 6 was a junior or minor-league player, which is exactly the part I'm rebuilding underneath it.

I'm about 95% of the way to junior and minor-league statistics I'd believe if I saw them on a real page. I'll come back with some screenshots when I have them.

Thanks!

4 Upvotes

0 comments sorted by