Judges
What changes when a person does the rating
Not a question of which is better. They answer different questions, and most disappointment comes from asking one of them for the other one's answer.
Guides on Judges: Boundaries belong to the judge, The case against the price list, What the skill in judging actually is
Human rating differs from an algorithm in what it answers: an algorithm gives a consistent, comparable number, while a person responds to you specifically, in a register you chose. Given the same photograph, the two produce things that look similar - an assessment, some commentary, often a number - and are not the same product.
Being clear about the difference before you commission anything is the single best predictor of being happy with what arrives.
What the algorithm is good at
Consistency. The same input returns close to the same output. It does not get tired, generous at the end of a session, or influenced by the previous submission. AI Penis covers why, and where the consistency runs out.
Impersonality. It does not know who you are and cannot be embarrassed. For a lot of people that is the entire appeal.
Speed and price. Minutes, and a price on the card before you commit.
People are not always fair to it, either. Dietvorst, Simmons and Massey (2015) found people lose confidence in an algorithmic forecaster faster than in a human one after seeing both make the same mistake.
Comparability. Because it is consistent, scoring several of your own photos on one system produces a meaningful ranking. That is a real analytical tool - provided you read the scores as positions in a distribution rather than as verdicts.
What a person is good at
Responding to you specifically. This is the whole difference and it is not a small one. You write what you want; what comes back answers it. An algorithm generates from a template and cannot do otherwise, however well the template is tuned. Consumer researchers have a name for the doubt people bring to automated judgement: Longoni, Bonezzi and Morewedge (2019) call it "uniqueness neglect", the belief that an automated system cannot account for a person's individual characteristics.
Register. Encouraging, blunt, degrading, clinical, affectionate - a person who does one of these well does it in a way a tone setting approximates rather than achieves. Most people commissioning a human review are buying the register, not the assessment.
Judgement about what you actually asked for. Requests are frequently a bit off from what the person wanted, and a good judge reads through to the intent. Nothing automated does this.
Presence. Someone looked. For most buyers this is the point, and it is not a thing an algorithm can supply at any price, because its absence is definitional rather than a limitation.
What neither is good at
Telling you something true about yourself. Neither is a measurement, either; if a length is what you were after, a ruler and the standard method is the answer and it involves nobody's opinion. One is a model's read of an image, the other is one person's opinion. Both are entertainment, and treating either as a verdict is a way to have a worse time with it than it was designed to give you.
Which to ask for
Want a number you can compare against other numbers? That is the algorithm's job, and it does it better.
Want a specific person to respond to a specific thing in a specific register? No amount of model quality substitutes.
Worth being honest with yourself at this point, because what people are actually buying is usually not the assessment - and the ones who work that out before ordering are reliably the ones who come away happy.
If you are weighing the two as a purchase - price, wait, what you get - Rate Cock has written that comparison directly, which is a more useful place to start than anything general.
If you have decided on a person, the next questions are practical: what actually arrives, how it gets priced, and how to ask for it well. The last one changes the result more than most people expect.