Judges
Why a judge is not trying to be repeatable
A judge gives a different clip to a different brief on the same photo, and that variability is the product, not a flaw to be engineered out.
Guides on Judges: Boundaries belong to the judge, The case against the price list, What the skill in judging actually is
A human judge trades consistency for responsiveness on purpose: send the same photo twice with two different briefs and you get two different clips, because a person answers what was asked rather than repeating an output. That is not an inconsistency to explain; it is what a person is for.
Two things that trade off
Any rating system, human or automated, sits somewhere between two properties that pull against each other.
Repeatability means the same input reliably produces close to the same output, regardless of context. It is what makes a result comparable - across your own photos, across time, across different people's submissions.
Responsiveness means the output actually answers what was specifically asked, adapting to a particular brief rather than running the same process regardless of it.
A system cannot maximise both at once, because responding fully to a specific, changing brief is exactly what breaks repeatability, and locking down repeatability is exactly what limits how much a system can respond to what is actually in front of it. Every rating method sits somewhere on that line, and the honest thing to do is pick a spot rather than pretend a system does both. Researchers who rely on human raters treat the gap as something to measure: Hallgren's 2012 tutorial on inter-rater reliability describes quantifying "the degree of agreement between two or more coders who make independent ratings", and treats disagreement between them as measurement error.
Where a human judge sits, and why
A judge sits deliberately toward the responsive end. The brief changes what the clip covers, how it is delivered, what register it takes, and often what conclusion it reaches emphasis-wise, because the entire value of paying a person rather than running a fixed process is that a person can actually answer what was asked rather than applying the same template to everything that arrives. Asking a judge for a consistent, repeatable read on the same material regardless of brief is asking them to stop doing the thing you are paying for.
This is not the judge being unreliable in the way that word usually means. Reliability, for a human review, is about whether the clip matches what was quoted and requested, not whether it matches some other clip made under a different brief. A judge who delivers the register you asked for, addressing what you specified, has done a reliable job even if a friend who briefed differently got a completely different clip from the same person on the same photo. Whether a judge can be objective at all is a related question worth reading on its own.
Where the other end sits
An automated system sits at the opposite end on purpose, and that is a legitimate design choice rather than a lesser one. AI Penis explains why its consistency exists and where it runs out, and the short version is that a fixed process trades responsiveness for exactly the repeatability a judge is not offering. That trade is what makes an algorithmic score usable for comparing your own photos against each other over time - the same instrument, applied the same way, every time. It is a genuinely different kind of usefulness from what a person provides, not a worse attempt at the same thing.
What this means for a buyer
Do not expect a judge to behave like an instrument, and do not read variability between two commissions - your own, or two different people's - as evidence that one judge is more accurate than another. Accuracy in the measurement sense is not what is being sold. What is being sold is a response to a specific brief, and the fact that a different brief gets a different response is the mechanism working correctly, not a defect in it.
Where this goes wrong in practice
The mistake usually shows up as a comparison, not a complaint. A buyer commissions two judges with the same photo and a similar brief, gets two different clips back, and reads the gap as one judge being better or worse at the job. Sometimes that is true - skill varies - but the gap is just as often two people answering the same words differently because the words left room to, and a judge who fills that room with their own read is doing exactly what a person is for. The same trap catches repeat buyers: commission the same judge twice with briefs that only look similar, and a different clip is not the judge being inconsistent with themselves, it is the second brief asking for something the first one did not.
Reading a brief back before sending it helps here more than any amount of judge comparison does. If the brief leaves the register, the pacing or the emphasis open, expect the judge to make a call on it, and expect that call to differ from whatever a different judge - or the same judge on a different day, with a different unstated brief in mind - would have made. Tightening the brief tightens the range of reasonable clips it could produce; it does not, and should not, force every judge toward the same one.
What to ask for instead of consistency
If what you actually want is repeatability - the same read applied the same way, so you can compare outcomes across photos or over time - a judge is not the tool for that regardless of how carefully the brief is written, because responsiveness to the brief is the entire mechanism, not a setting that can be dialled down for one commission. The fuller comparison between the two approaches covers the rest of what changes when a person does the rating; this is the specific piece of that difference worth naming on its own, because it is the one buyers most often mistake for a fault. Wanting both a judge's responsiveness and an instrument's repeatability from the same commission is asking for two properties that cannot coexist in one system, no matter how good the judge is. A platform like Rate Cock makes that trade explicit up front, rather than letting a buyer discover it mid-commission. An automated tool, the kind measuring a photo directly, aims for the opposite property on purpose - the same input should always produce the same output. So does a scoring model: the same photo submitted twice gets the same number back, which is the whole point of using one instead of a person. This is closely related to why a judge's number is not a score in the first place - a different concept doing different work.