Judges

Rankings of judges, read carefully

A judge ranking measures activity and client satisfaction over a window, not who is 'best', and reading it as the latter leads to the wrong commission.

By Updated 4 min readJudges

Guides on Judges: Boundaries belong to the judge, The case against the price list, What the skill in judging actually is

A leaderboard of judges measures activity and satisfaction over a recent window - volume, on-time delivery, feedback and recency - not talent. It puts names in order, and an order looks like a verdict even when it is not one, which is why reading it as "who is best" is its most common misuse.

What actually feeds a ranking

Most leaderboards are built from a handful of measurable signals, and none of them is a direct read on talent.

Volume over a window. How many commissions a judge completed recently, which rewards judges with capacity and a steady queue, not necessarily the most skilled register work.

Completion and delivery. Whether clips landed on time and as quoted, which is a real quality signal but a narrow one - it measures reliability, not how good the clip itself was.

Client satisfaction after delivery. Tips, repeat bookings, and any rating left after a clip, aggregated over the same window. This is the closest thing to a quality signal a leaderboard has, and it is still filtered through whoever bothered to leave feedback.

Recency. Most rankings weight recent activity over all-time activity, so a judge who paused for a few weeks drops even if nothing about their work changed.

Rate Cock runs a leaderboard built roughly this way, as one example of the genre - useful for finding an active judge quickly, and worth treating as exactly that rather than a verdict on craft.

What it cannot capture

A ranking has no way to see register fit. The judge at the top of a leaderboard might specialise in worship, and if what you want is a clinical, detached read, their position on the list tells you nothing about whether they are right for you - the default tone a judge falls back to matters more to your outcome than where they sit in an aggregate ranking.

It cannot see boundaries either. A highly ranked judge may simply decline the specific thing you want, and a lower-ranked judge who does exactly your register well is a better commission for you regardless of position. Judges set their own limits independent of how busy or popular they are, and a leaderboard position says nothing about where those limits sit.

It also cannot see honesty about capacity. A judge climbing fast on volume alone may be taking on more than they can sustain, and judges who pace themselves deliberately sometimes rank lower for it while doing more careful work per clip.

It cannot see how a judge treats a decline, either, and that is one of the more useful things about them. A judge who turns down work near their limits cleanly, rather than taking it reluctantly and delivering something flat, is doing something a ranking has no field for, because a decline does not generate a data point the leaderboard counts.

Activity is not the same as quality

The clearest way to see the gap: a judge who takes twice as many commissions in a window will usually outrank one who takes half as many and spends longer on each, even if the slower judge's clips are more specific and better crafted. A ranking built on throughput will always reward throughput, because that is the thing it is actually counting. Visible rankings also feed themselves. In Salganik, Dodds and Watts's 2006 experiment in Science, 14,341 participants chose unfamiliar songs with or without seeing earlier listeners' choices, and stronger social influence increased both the inequality and the unpredictability of success. As the authors put it, "success was also only partly determined by quality." None of this makes the ranking dishonest - it is measuring exactly what it says it measures, which is activity and satisfaction over a period, not craft in the abstract.

How to actually use one

Treat a leaderboard as a shortlist generator, not a decision. Use it to find judges who are currently active and reliably deliver what they quote, then narrow from there on register, stated limits and portfolio - reputation signals that actually mean something go beyond a rank number, and they are what should decide the commission.

If a consistent, ranked number against a wider population is genuinely what you are after rather than a person, that is closer to what an automated scoring tool is built to give you, and understanding how that kind of system arrives at a score is worth doing before you lean on it the same way people over-lean on a leaderboard position. And if the comparison you actually want is a measurement against other people rather than either a judge's opinion or an algorithm's score, that is its own separate question with its own method, unrelated to how any judge ranks this month.

Read next

Full archive