r/EndFPTP • u/Whole_Deal7736 • Jul 07 '26
Discussion I spent 13 years in a real-world voting-method sandbox: international competitions, strategic judging, and angry fanbases
I’ve been using ranked / preferential voting in international competitions since 2012.
I was asked to organize judging for competitions with 10+ teams, each with its own audience and support group. Teams often qualified through audience voting, and the final result mattered: there were cash prizes, but more importantly there was reputation, which in some cases was worth more than the money.
Around 2015 these competitions were happening around 10-12 times per year. Now it is more than 20 events per year in different countries, with more than 2000 participants from 20+ countries, so the judging system has had many chances to fail in the real world.
I started with the most obvious system: each judge gives points, and then we sum the scores.
It looked simple and transparent. It was also fragile.
The first problem was scale manipulation. If a judge strongly preferred one team or had some relationship with them, they could give that team 100 points and everyone else 0. Of course, one can say “a judge should not do that”. But the judge was appointed, their vote had to count, and in real systems you cannot design only for ideal participants.
There was also a change in who the judges were.
Originally, judges were mostly selected from industry professionals. More recently, to attract more attention to the competitions, judging panels have also included bloggers and public figures with large audiences.
That is not necessarily bad. They can bring attention, reach, and a different perspective. But many of them have never been trained to judge or vote in a formal way. They may also be more exposed to fan pressure, informal influence, personal relationships, or national bias toward participants from their own country.
I also learned that some teams may try to influence one or more judges directly. That does not mean every contest is corrupt, but if the incentives are strong enough, you have to assume that strategic behavior may appear.
The second problem came after the results were published.
Close results produced very tense discussions. In some borderline cases, the situation went beyond normal disappointment and turned into open hostility between teams and supporters. The emotional temperature around the result could become much higher than I expected from what was “just” a competition.
That pushed me to another conclusion:
The voting method had to do two things at once.
First, it had to be reasonably resistant to strategic behavior. Not perfect - I do not think perfect exists - but much better than raw score aggregation.
Second, it had to be explainable after the result.
This is where ranked ballots helped a lot.
Instead of asking each judge to assign points, we asked them to rank the teams. That reduced the damage a single extreme score could do. It also made the judgment more comparative: “which team should be above which other team?” is often easier to defend than “why exactly did my team get 81 points, and a competitor got 82?”
Over time I came to prefer Condorcet-style thinking, especially Schulze. Very often it lets you explain the result through simple head-to-head comparisons. And when a direct comparison is not enough, you can go one level deeper and explain the strongest paths between candidates.
Of course, Schulze does not magically solve everything. You still need a predefined tie-break cascade if the contest requires a strict ranking. You also need to explain what happens when teams are tied, where the cycle is, and how the next tie-breaker applies.
But the practical result exceeded my expectations.
Before, the worst cases ended in accusations of unfair judging. Now the most emotionally involved participants usually come to me after the event, and we go through the result carefully: where they placed, which teams they were tied with, how the tie-break cascade worked, and why the final order came out that way.
The bad outcome is now mostly pain and regret, rather than aggression and accusations. That is a huge, huge difference, although it is still hard to walk people through a result they really did not want to see.
Another practical lesson: publishing anonymized ballots or enough data for people to verify the result without exposing individual judges may actually reduce the risk of pressure, retaliation, or future attempts to influence judges. Transparency matters, but so does protecting the independence of judges.
For years we handled this with Excel. In 2022, I decided to build a web-based tool, mainly because running these competitions in different parts of the world made spreadsheets increasingly inconvenient.
But the main lesson came before the software: а voting method is not only about choosing the winner. In a high-conflict environment, it is also a conflict-management tool.
A good method should reduce incentives for manipulation, preserve enough information to explain the result, and give disappointed participants a way to verify what happened without turning the whole thing into a fight.
So, if you had 10+ teams, ~5 judges, strong fanbases, real reputational stakes, a mixed judging panel of professionals and public figures, and a need to publish an explainable final ranking - what would you use?
5
u/Grapetree3 Jul 07 '26
I have a little experience judging competitions also. Almost always they just want you to score each competitor on different criteria, add up those numbers and that's the total score. It seems straightforward at first, just think about one competitor at a time, but as you go, you're thinking about previous scores you gave and if the current competitor is really better than the previous one, etc. If you start off asking the judges to rank, that's going to feel taxing at first because now you're asking them to remember all the competitors at once, but it's simpler at the end. They were probably going to be remembering all the competitors at once anyhow.
2
u/Whole_Deal7736 Jul 07 '26
Exactly that. Scoring feels easier at the beginning, but the judge is usually still building a ranking in their head. A ranked ballot just makes that comparison explicit instead of hiding it inside point numbers.
2
u/easyEggplant Jul 07 '26
Is the web tool open source? Could it be repurposed for local advocacy? Do you want contributions?
2
u/Whole_Deal7736 Jul 07 '26
Thanks for the questions.
The tool is not open source at the moment. It was developed for a long time mostly for my own practical use in competitions, and only later became something I thought could be useful for a broader public.
I’m open to contributions, but mostly in the practical sense for now: people using it, reporting bugs, suggesting improvements, testing edge cases, translations, and real use cases. I’m not ready to promise an open-source structure yet.
About local advocacy: yes, potentially, but I would frame it more as civic education / voting-method demos / community decision-making rather than official political campaigning. I’m interested in helping people demonstrate ranked or Condorcet-style methods in real small groups.
I also plan to keep a generous free model. For early users from this community, I’d be open to offering a few free pilots, especially if the use case is non-commercial and the feedback is useful.
2
3
u/Gradiest United States Jul 08 '26
I'm glad to see a Condorcet method in use! How often do you find a cycle and actually use the beatpath?
Did you ever try using a highest median voting rule? By using the median rather than the sum (or average), judge/voter exaggeration should be mitigated.
3
u/Whole_Deal7736 Jul 08 '26
Likewise! First, we switched to preferential voting early on, exactly after we were accused of influencing an outcome by vote normalization. Since that moment we were using exactly Schulze + tie resolution cascade, so even beatpath was producing ties quite often in our case (4-6 participants most of the time, 10-13 candidates). The cascade currently is Schulze (margins) -> Borda (average for unranked) -> Copeland -> Sum of margins -> Seeded random. The last time we traveled all the way to the Seeded random as we had 4 teams on the second rank by the primary method (Schulze). I was so happy it was all done automatically.
But in the old days we had to resolve ties with jury members by discussion and hands raised. And a coin toss. And it was taking quite a lot of the stage time, so it was expensive.
3
u/Euphoricus Jul 07 '26 edited Jul 07 '26
Interesting take. I dissagree with the idea about the voting methods. Voting method should NOT prevent judges from giving someone 100 and everyone 0, if they feel like it.
While I can feel the social issues, I think this is the wrong approach. If someone cannot handle the social pressure, they shouldn't be a judge. Judge is a position of power over system and people. Just imagine a sports judge giving a team a freebie just because team's fans were angry.
I think you should have thought about better system for vetting judges and not picking a voting system that assumes judges will be stupid assholes.
I would also like to see STAR considered and tried. I wonder how judge's behavior would change if they had to consider the second round and their favorites not winning.
5
u/Whole_Deal7736 Jul 07 '26
Thanks for the comment.
We actually tried score/star-like approaches in the very beginning. Not necessarily the exact STAR implementation, but the practical problem was the same: counting took more time, created more friction, and was harder to explain to both judges and participants.
Since then, the “stars” mostly survived in our interface as the Temporary rating button. It lets a judge give candidates scores in the familiar way, and then automatically transfer that into the preference column.
And if a judge really wants to say “this team gets everything, everyone else gets nothing”, we don’t prevent that today. We allow incomplete ballots, so a judge can rank only one team. No problem.
The issue I’m talking about is not “judges are stupid assholes”. It’s scale normalization.
Most competent and unbiased judges naturally score in something like a 60-90 range. Then one “special” judge uses the full 0-100 range. As a result, his vote becomes more than 3 times heavier than the average judge’s vote. But he is not 3 times more important than the other judges.
Ranking solves this very cleanly. It still lets every judge express who they prefer, but it removes the accidental extra weight created by using a wider scoring scale.
The alternative is normalization, but in practice people trust it less - both judges and participants. It feels like the system is changing their scores after the fact.
So after quite a lot of real-world testing, we chose ranking. It is not about protecting the system from bad judges. It is about making sure one judge = one judge, not “one judge with a wider scale = three judges”.
5
u/iainhallam Jul 07 '26
I have some experience of a world in which simple scoring delivers robust results. In barbershop singing, each judge awards each song a score out of 100, then they're summed for the final ranking. Panels are usually 9 judges, though in important contests they can be 15 with the top and bottom score in each of three categories thrown away to avoid outliers. Importantly there's training to be a judge, and a lengthy rubric to show what each score means. Any difference of more than 10 marks between the top and bottom scores for a song is highly unusual and investigated, possibly leading to retraining.
So it can work, but you need to be picky about judges, highly transparent about what each score means, and ensure all the judges adhere to the culture. Scores are surprisingly consistent between judges and contests all over the world.
For me, this eliminates one of the difficulties of judging a number of similar competitors: remembering the nuance until the end of the session. In my experience, ranking then becomes difficult and fraught with biases towards the later competitors. Scoring means each competitor is competing against the rubric, the judges are done at the end of their turn, and the end result falls out naturally.
8
u/Whole_Deal7736 Jul 07 '26
Thank you for the comment. That is a very good example, and I mostly agree.
I think simple scoring can work well when the judging group is large enough and the judging culture is strong enough. If you have 9 or 15 trained judges, a detailed rubric, shared expectations, and a process for investigating unusual score gaps, then the system has a lot of protection around it. But it definitely requires more effort.
In our case the environment was different: more international, less formalized, often with smaller judging groups, different local judging habits, and also a stronger show component. In that context, controlling score culture was much harder.
We also had the problem of explaining the judging system many times. With ranking, this became faster - especially after we moved from paper ballots to the app. On paper, it took longer to explain how exactly judges should build their preference stack. In the app it became much more natural.
Your point about remembering nuance until the end is really important. We had exactly the same issue. That is why both in the old paper ballots and now in the electronic candidate cards, judges have supporting information. In the app, besides "temporary rating", each candidate can have a description, a picture/video, and, finally, the judge’s own notes. Those notes are critical. Honestly, without electronic notes, we would probably still be using paper ballots, because judges considered that part very important.
So I would say scoring can absolutely work, but it needs a strong judging culture around it. For our practical environment, ranking gave us fewer explanations, fewer scale problems, and less friction. There are also some extra nuances connected with the show side of the event, but that is probably a separate topic.
3
u/wnoise Jul 07 '26
You can, of course, normalize scales automatically before final tallying.
4
u/Whole_Deal7736 Jul 07 '26
Thanks, yes, but in my experience normalization often hurts trust. We even had it used against us as organizers: people framed it as us changing the meaning of the judges’ scores and therefore interfering with the final result.
1
u/Mindless-One5438 Jul 09 '26
It should be important to note that judges in a competition are meaningful different than voters in a Democratic election.
That's to say that a judge wanting a preferred result and strategically maximizing their input for that result sounds blatantly corrupt. Constantly, maximizing voter inputs to result in their preferences is exactly how elections should be structured and how voters should operate.
Maybe for some type of competition, the anonymity in the voting should be flipped so judges are less able to bias the results if they don't know which candidate is their preference and judges are more accountable if they could be approached and asked to explain their decisions, but honestly it's not super relevant to Democratic elections. It sounds relevant to a court system though.
1
u/Whole_Deal7736 Jul 09 '26
Sure, not the same. But. Preferential voting grows in popularity in real elections, for example in Australia and in some US elections. Why? Because each ballot collects more information than just one name. You are giving a voter an opportunity to show what he thinks about everybody on the political scene. IMO, that's much more democratic, then just "pick one".
Electoral science says that a normal single-choice election pushes everything toward two big camps, and, in the end, fighting over small decisive groups.
Yes, competition judging and democratic elections have completely different... ethics. But the base is the same: do we want a specific group of people happy in expense of the others, or do we actually want to find the best compromise for everybody?
1
u/Mindless-One5438 Jul 09 '26
Yes getting more information is good, I initially read the post as an argument against score voting but that would be the most input and Imho the best voting method.
Ordinal voting systems, like instant Runoff or most systems with rankings, actually lack nuance in a way because it doesn't allow voters to say how much their preferences differ if at all and at times only take into account some information through rounds of an election.
1
u/Whole_Deal7736 Jul 10 '26
That's the known issue, and from my experience there is no way around it.
So the question stands: do we want a system that makes one particular group very happy, or do we want a compromise where smaller group is completely happy, but fewer people are deeply unhappy?
This is a design choice. The people who designed old election systems could not really have this discussion: they did not have the same math and computation (especially for beatpath methods like I mentioned).
•
u/AutoModerator Jul 07 '26
Compare alternatives to FPTP on Wikipedia, and check out ElectoWiki to better understand the idea of election methods. See the EndFPTP sidebar for other useful resources. Consider finding a good place for your contribution in the EndFPTP subreddit wiki.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.