Reviews
Recognising a poor clip
Generic phrasing, register drift, ignoring the brief and visible discomfort are the marks of a bad clip, and most of them trace to a mismatch, not malice.
Guides on Reviews: SPH as a genre with conventions, Five things a review is mistaken for, The words people use, defined, What happens in a two-minute clip
A bad review clip shows one of four marks: generic phrasing, register drift, an ignored brief, or visible discomfort. It is rarely a judge failing at their craft, and knowing which mark you are looking at tells you where the problem sits - with the judge, with the brief, or with a mismatch that was nobody's fault.
Generic phrasing
The clearest sign: nothing in the clip could not apply to any buyer. No detail from your material gets mentioned, no line from your brief gets answered, the register is competent but interchangeable. This is almost always a brief problem rather than a judge problem - a judge can only respond to what was offered, and a request with nothing specific in it produces a clip about nobody in particular by construction, not by laziness. The fix lives on the buyer's side: one concrete detail in the next brief changes this more than choosing a different judge would.
Register drift
This one sits with the judge. A clip that opens in one register and slides into another - warm at the start, flatly clinical by the end, or an SPH request that softens into reassurance halfway through - has lost the thread of what was agreed. Some drift is a register outside the judge's actual range being attempted anyway rather than declined, which is on them; a good judge turns down a register they cannot hold rather than perform an unconvincing version of it. Occasional drift is fatigue on a long queue day, which is still worth naming as feedback, since it is the kind of thing a judge can act on.
Ignoring the brief
Different from generic phrasing, and worse. A clip that ignores a stated specific, addresses the wrong length, or skips something the buyer asked to have covered is a delivery that missed what was agreed, not just a clip with nothing extra in it. This is squarely the judge's error if the brief was clear, and it is the kind of miss a revision exists for - though that is a separate conversation from diagnosing why it happened in the first place. Sometimes the honest cause is a brief that buried the actual ask in a lot of surrounding text; a judge working from a long, unstructured message can genuinely miss the one line that mattered, which is why a short, four-line brief outperforms a detailed one more often than people expect. Writers also overrate how clearly their meaning comes through on the page: Kruger, Epley and colleagues (2005) traced overconfidence about email tone to egocentrism, "the inherent difficulty of detaching oneself from one's own perspective".
Visible discomfort
The hardest one to fix after the fact, and the one that most often means fit rather than fault. A judge who looks like they are enduring the request rather than performing it has taken on something outside their actual range, willingly or under queue pressure, and it shows. This is not a moral failing on either side - it is what happens when a request lands with a judge who should have declined it and did not, and the honest fix is upstream: reading a judge's stated boundaries before sending the brief prevents this far more reliably than any amount of care in how the request is phrased.
Sorting the four by cause
Two of these four - generic phrasing, and often ignoring the brief - trace back to the brief itself. Two - register drift and visible discomfort - trace to a mismatch between the request and the judge, sometimes compounded by workload. None of the four is really about a judge's underlying skill; what separates a genuinely great clip from a competent one is a different set of qualities entirely, and a judge capable of a great clip in the right circumstances can still produce a bad one in the wrong ones.
This is not a symmetrical failure mode with an automated tool, worth saying once. An AI judge does not drift or grow uncomfortable - its failure modes are about model consistency, a different axis entirely. A measurement can be simply wrong, taken badly rather than badly received, which is a data problem, not a fit problem. And a bad result on a photo-scoring tool traces almost entirely to bad input, since there is no register to miss and nothing to be uncomfortable about. Only a human review can fail in the specific, personal ways described here, and that is the cost of a format built around a person actually responding to you - worth weighing against what an algorithm gets you instead if reliability matters more to you than presence does.