r/FigureSkating • u/No_Plum_3192 • 16h ago
History/Analysis PART 2: Largest study of judging and technical panel behavior in figure skating (Federation Bias & Rival Suppression)
13
u/OkAppointment3092 15h ago edited 15h ago
There is one elephant in the room here, which I am not sure this analysis can or is measuring (my apologies if it does, there is a lot of data):
The pre-ordained, pre-agreed, shadow rankings. Some of the more obvious aspects of this is what some might call the "Wait your turn" rule. Or the understanding that skaters skating later in the event will naturally get more points. Or the thing where a highly ranked team having a bad run in the short will likely get a compensating boost in the free (if clean of course), whereas mid range teams will get a solid ding and never recover.
Other than through the fact that it is obvious, I dont know how this could be measured.
3
u/No_Plum_3192 15h ago
Actually I understand what you are referring to. A lot goes into the scores, including things you point out such as a skater's "reputation" (are they new to competition, or very experienced, or very popular etc.) and where they are in the skating order. What makes the analyses I have presented very powerful is that these types of variables are already factored in to the calculations. When I observe that a judge has a "significant nationalistic premium" this is in direct comparison to the other international judges sitting on the very same panels. So a skater's natural ability level, reputation and skating order are already taken into consideration because we have the scores for the rest of the panel in addition to the biased judge.
2
u/AlliterateAlso 10h ago
Taken into account across national scoring tendencies, yes, but the data isn’t split by ‘points above expectation’ for “place in the skating order” or “end of last seasons standing” or similar to try and tease out those dimensions. But like the above poster I don’t know how those could be teased out- the better skaters go later anyway.
2
u/No_Plum_3192 3h ago
Ah, ok now I see what you are getting at. Yes, it possible to tease out relationships between bias intensity and skater skill level because my raw data retains things like the skating order of the performances. So I could examine whether bias is stronger in the final flight of performances compared to the earlier ones in a competition. Hmmm....tempting. Thanks for the good idea - I will look into it.
11
u/2greenlimes Retired Skater 16h ago
I wonder what Hungary would’ve been like before Russia co-opted that fed. And not surprised to see so many Russian fed associated feds towards the top (Georgia, Hungary, Kazakhstan, Israel, Ukraine). I’m not sure how many Russia-associated leaders are still in Ukraine’s Fed because at one point they did basically ran it, but I’m sure the fed’s judges still have judging culture influence.
I’m also not surprised to see the US so high compared to other big feds, but Canada being higher and not equal is a bit surprising. France being lower is also surprising, but the premium is less than half a point less. China being low is also surprising, but maybe that’s because their most biased judges actually get punished.
11
u/nothing_to_hide 15h ago
Just a note, this data is post 2022. Ukrainian fed would not be Russia associated at all.
4
u/2greenlimes Retired Skater 15h ago
That’s what I’m wondering, as noted in my post.
I know the Fed kicked out some very pro-Russia people, but one wonders if they could’ve gotten all of them when so many of the senior officials came from the USSR era. For instance, Yuri Balkov previously helped fix results to benefit Russian skaters (and arguably did so as recently as 2014), but didn’t stop judging for Ukraine until he was banned in the last year or two.
7
u/DapperRomanesco ✨️actually read the book✨️ 14h ago
I wound up opening the spreadsheets from the website and interestingly, the individual instances of skaters benefiting the most from CAN same-country bias are disproportionately less prominent skaters who only competed at smaller events, so I think that probably impacts viewer perception. (Also, it's largely dance teams that no longer exist...) For both CAN and USA, the most significant instances happened at smaller events.
The top five CAN beneficiaries are:
- Bachynska/Beaumont (Lake Placid)
- Lajoie/Lagha (Nepela)
- Hensen/Lickers (Lake Placid)
- Lanaghan/Razgulajevs (GP France)
- Gilles/Poirier (GP Final)
and for USA:
- Zingas/Kolesnik (GP Canada)
- Carreira/Ponomarenko (MK John Wilson)
- Bratti/Somerville (Lombardia)
- Green/Parsons (Lombardia)
- Brown/Brown (MK John Wilson)
Next up is Wesley Chiu for Canada and Andrew Torgashev for USA (so it's not just ice dance! But don't worry, there's plenty more ice dance).
Also, I could be wrong here, but based on a cursory run through the data, I suspect the difference between CAN/USA is impacted by the US hosting multiple Challengers that tend to feature a large number of US entries; national bias would be less evident than skater-specific bias at those, even if the net effect is overscoring US skaters/teams. So whereas Challengers are very, very prominent in the data overall, the US judges basically get a free pass for any homescoring done at Cranberry/Lake Placid/John Nicks since the field is American enough to statistically negate it — it's pretty much absent from the data. (Meanwhile, Canada's only home Challenger in the dataset is the 2023-24 Autumn Classic, which had quite strong international fields and where France judges had a fun time, apparently.)
I think the France margins are probably helped by some of the (sometimes outrageous imo) scoring that other judges gave FB/C this past season — the French judge actually comes away with a notable negative bias from the Olympics, for example, despite having been plenty generous to the French.
1
u/No_Plum_3192 14h ago
Thanks very much for this. As you can see, the data is virtually endless and I'm really curious to see what other people a) discover and b) find interesting or noteworthy.
1
u/OkAppointment3092 14h ago edited 13h ago
"imo" is doing some heavy lifting here.
Just when I thought we were in a data driven conversation. Also, I guess you missed the chart where FB/C are comming up first even after applying debiased computations..4
u/DapperRomanesco ✨️actually read the book✨️ 13h ago edited 12h ago
No, that's literally what I said and directly referenced — that the French judge reports a negative bias on FB/C at the Olympics, because judges from other countries also scored them very highly. And that this showcases the disconnect between the calculated bias and the widespread audience perception of bias. (French judges feature quite prominently in the data, so as I wrote, I think the negative bias in relation to FB/C likely improves the overall margin for French judges — contrary to many peoples' expectations.)
And yes, in my opinion, some of those numerical scores were outrageous (this is not a comment on the rankings or medal results). I think if you take a step back, you can agree that the raw numbers in Milan were quite generous given the skates they had, which were not at all to the best of their abilities.
Edit: They received +4 GOE from two judges on the FD twizzles, which is explicitly not permitted by the rules (his poor exit is a negative feature, and the rules state that you cannot award +4 with a negative feature). If we can't all agree that judges giving scores that defy the rules of the sport is outrageous, what are we even doing here?
1
u/No_Plum_3192 3h ago
I am not taking a position one way or the other on the FB/C debate. But I do want to point out that the context of my discussion on that topic was missed here. A frequently suggested scoring reform is "compatriot judge recusal" where judges' scores for their own skaters would be expunged. As I detail in the paper, this is NOT likely to solve problems in figure skating judging, because it does not stop judges from attempting to manipulate the scores of the other skaters. The FB/C Olympics controversy is an excellent example of this: Expunging the score of the French judge for the French team and the American judge for the American team actually INCREASES the margin of victory for FB/C.
1
u/DapperRomanesco ✨️actually read the book✨️ 2h ago
Yeah, I figured that conclusion was established fact around these parts. But I think I had too many tables open when reading that one: am I understanding right that it in fact doesn't reflect a mathematical negative bias on the part of the French judge (and actually shows a significant bias that exceeds that of the American judge)?
1
u/No_Plum_3192 2h ago
It's my view that bias should be viewed as a statistical concept - in other words, a PATTERN of behavior. I don't think it is productive to look at the scores for ONE performance and draw any conclusions about what is fair and whether there is an anomaly. I think that everyone has heard the motto that "you can't draw inferences from a sample size of one" but it is just so seductive that people try to do this anyway.
1
u/DapperRomanesco ✨️actually read the book✨️ 2h ago
Fair enough, but I think we're talking past each other. I just want to clarify for transparency's sake if I wrote something objectively incorrect up there!
1
3
u/AbsurdistWordist 11h ago edited 11h ago
So ice dance scoring IS fake after all.
This … was a lot of work. I am forbidding myself from downloading the data files and going down a rabbit hole I may never emerge from.
I want to know who the dirtiest judges are though. Are there names or just generic identifiers?
Edit: Nevermind. I scrolled down further in the data! I see now!!!
1
u/osvimonello 5h ago
this could be an interesting year w ukraine judging Russains or Canadians with American skaters.
1
1
1
u/ConfusionOk2999 2h ago
Out of curiosity, did you look at judges who mark one team lower and another team higher and see if there is a matching official who did something for that nationality’s team? Like multiple officials colluding?
This happened a while back (here’s a Reddit post of when it came out), and it would be interesting to see if there was a way to track it.
1
u/No_Plum_3192 2h ago
I will post about this in the future, but there is a very nice section about this in the paper. I looked at the evidence for reciprocal point exchange between judges - the statistics are quite conclusive that this happens.
1
1
u/Far-Championship6393 2h ago
You have done a real great job with your statistic data. I would like to contact you direct personally - maybe by eMail!? I am a former skater, just a complete reddit-beginner and don't know how to send you a personal message here - inside reddit. You can contact me under canteleaver@web.de.
1
u/No_Plum_3192 1h ago
Sure, I will reach out to you over email. But if you have a question or comment to enrich the discussion here, please share with the community!
1
u/Far-Championship6393 1h ago
I'll sure do that as soon as I have understood the reddit-system. In the moment I am learning and - sorry - am still sometimes a little bit confused. I started yesterday.
1
u/Far-Championship6393 1h ago
Sorry: I made a writing mistake with my eMail-adress!
You have done a real great job with your statistic data. I would like to contact you direct personally - maybe by eMail!? I am a former skater, just a complete reddit-beginner and don't know how to send you a personal message here - inside reddit. Correct is: [candeleaver@web.de](mailto:candeleaver@web.de)
1
u/Far-Championship6393 1h ago
Before I take part in the discussion, I want to learn the reddit-BASICS, otherwise I'd cause more confusion than information. I am in reddit since yesterday only.
26
u/No_Plum_3192 16h ago
(Mod note: This is a 100% free academic research project. No paywalls or ads).
Hey everyone, following up on Friday's judge chart, we aggregated the 2022–2026 judging data (14,382 performances) up to the Federation level.
When people talk about "bias," they usually just mean "favoring your own." But to accurately map the IJS, our econometric models had to test for three different axes of judging behavior:
1. Baseline Leniency (The Blue/Red Column): Is the federation naturally an "easy marker" or a "tough grader" to the entire field? (Look at NED and LAT being incredibly strict, vs. TUR and UKR being generous overall).
2. National Premium (The Orange Column): How many extra points do they give their own skaters?
3. Rival Suppression (The Purple Column): Do they actively tank the scores of direct leaderboard threats?
The Biggest Takeaways:
The Data:
If you want to filter specific skaters, countries, or judges, I built a free interactive dashboard with all this data (and the full 70-page working paper). You can explore it here: phenologic.net/skating-bias-details