Reviews
No exchange rate between judges
Different registers, different personas and different briefs mean two clips about the same photos are two performances, not two data points.
Guides on Reviews: SPH as a genre with conventions, Five things a review is mistaken for, The words people use, defined, What happens in a two-minute clip
Send the same photos to two judges and you will get two clips that do not agree, and the disagreement is not a sign either one is wrong. There is no shared scale behind commissioned reviews for two clips to agree on in the first place. Expecting one is the mistake, not the mismatch.
The three things that move between judges
A review is the product of a judge, a register and a brief, and all three change every time you switch who you commission.
The judge. Every judge has their own frame of reference, built from whatever they have seen and whatever tone they naturally sit in. One judge's "pretty average" and another's are not calibrated against each other, because nothing forced them to be.
The register. Worship reads warmer than honest by design, and honest reads more clinical than reassurance by design. Comparing a worship clip to a clinical one is comparing two genres, not two assessments of the same thing.
The brief. A judge responds to what was sent and what was asked, and two briefs for the "same" submission are rarely identical once you look at the wording. A different specific named, a different length requested, a different thing to focus on: each of these steers the clip somewhere the other brief did not.
Change any one of the three and the clip changes with it, and a real comparison would need all three held constant, which never happens between two independent commissions.
Why this is not the same problem as inconsistency
It is worth being precise about what is being claimed here, because it is easy to hear "not comparable" as "unreliable," and that is a different question with a different answer. A single judge can be a thoughtful, consistent professional and still produce a clip that cannot be set next to a different judge's clip on a shared line. Consistency is about whether one judge is dependable with you over time. Comparability is about whether two judges' outputs share a scale, and they do not, for reasons that have nothing to do with either judge's skill. Where agreement between raters matters, researchers measure it rather than assume it; the guideline by Koo and Li (2016) treats an intraclass correlation below 0.5 as poor reliability and above 0.90 as excellent. Two competent professionals with different training, different tools and different aims will not converge on one number, and expecting them to misunderstands what each of them was doing. Human judgement is inconsistent on purpose covers the single-judge version of this in full; the comparability problem here is the multi-judge extension of the same idea, not a separate flaw.
What "comparable" actually requires
A meaningful comparison needs one process applied the same way to every input, so that the only thing that varies is the input itself. That is a specific, narrow condition, and a commissioned review does not meet it on purpose: it is built to respond to you, specifically, which is the opposite instinct from holding everything else fixed. Even trained raters only converge when they work at it: in Watari and colleagues' 2022 study, two examiners scoring recorded clinical exams against a shared rubric saw their agreement (weighted kappa) rise from 0.49 on the first ten videos to 0.82 on the last ten, after repeatedly discussing why each gave the scores they did. What changes when a person does the rating covers this trade in full: responsiveness is the entire value of a human review, and it is also exactly what breaks comparability. You cannot have a clip that answers your specific brief in a specific voice and also have it sit on a ruler with someone else's clip from a different judge answering a different brief. Wanting both from the same instrument is wanting two products in one purchase.
What to use if comparability is actually the goal
If the real goal is a ranking, a percentile, or "how do I compare across submissions of my own," that is a job for a system built to hold everything else constant, which a commissioned review was never trying to do. An automated score run on a consistent method gives you exactly that: one process, applied the same way each time, producing numbers that can be set against each other honestly. It will not give you a register, a specific reply to what you wrote, or the sense that someone looked, which is what a judge is for. The two are not competing answers to one question; they are answers to two different questions, and the comparability one belongs to the automated tool by design, the way AI Penis explains its own consistency in more detail than a review will ever offer. If a stated figure matters to what you are asking either instrument, take it with a proper method first rather than asking either one to estimate it for you.
Two clips, two answers, not two scores
Commission two judges if you want two considered responses, in two registers, from two people who took your brief seriously. That is a reasonable thing to want and a reasonable thing to pay for twice. Just do not average the results, rank them against each other, or treat the gap between them as an error somewhere. Rate Cock's comparison of the two approaches is the place to start if you are still deciding which question you actually have, before you commission anything at all.