Full review defeats the purpose

Checking every output means the work is still being done by a person, just differently — and it is the state in which teams conclude the automation isn’t saving anything. It is correct at the start and unsustainable as a permanent posture.

Sample randomly, not conveniently

Reviewing whatever is at the top of the queue over-samples one time of day and one kind of request. A genuine random sample across hours and categories is the only version that tells you about overall quality — and out-of-hours output is the least reviewed.

Reviewing what escalated tells you what it caught. Only sampling what it resolved tells you what it missed.

Review the resolved, not just the escalated

Escalations are self-reporting; the expensive failures are the confident wrong answers that never escalated. Deliberately sampling successful-looking interactions is the only way those surface.

Score against your existing standard

Use the same criteria you would apply to a person doing the job, not a stricter one. The comparison that matters is against what you would otherwise do, and teams frequently hold automation to a standard their own team doesn’t meet.

Convert every miss into a permanent test

A problem found in review should end up in the evaluation set so it cannot recur silently. That loop is what makes review compounding rather than repetitive — the same mechanism as turning an incident into a regression case.

Let the rate follow the evidence

Start high, reduce as findings become rare, and increase again after any change to prompts, knowledge, or the model. The review rate is a dial, and the person who owns it should be the one turning it.