Short answer
Because “a person looked at it” is not the same as an independent, capable and adequately resourced verification. Reviewers can be anchored by fluent AI prose, lack the source material or subject expertise, inspect only the surface under time pressure, or assume a system that is usually right is right this time. Effective review requires deliberate workflow design: identify the claims that matter, expose independent evidence, give reviewers genuine authority to reject the output, and test the review layer with sampling, seeded errors and incident data.
Human review is not one control; it is a set of conditions
Many risk registers add “human in the loop” as though a confirmation button completes governance. The useful questions are more specific: Who reviews? What must they check? Against which evidence? How much time do they have? What competence is required? Can they overturn the system? What happens after they find a problem?
If a reviewer sees only an AI summary and not the contract, clinical record, dataset or meeting recording, they can assess whether the prose resembles an answer but not whether it is true. If they must clear hundreds of cases and performance is measured by speed, rapid acceptance may be the predictable result. The human stage then becomes responsibility theatre rather than an independent defence.
NIST's human–AI interaction material identifies automation bias, over-reliance and unclear allocation of functions as risks that need active management. Review quality must therefore be designed and measured; it cannot be inferred from the presence of a final click. NIST AI RMF: Human-AI Interaction
Fluent writing makes an answer easy to read, not easy to verify
Generative AI is good at producing complete structure and confident tone. Headings, numbered arguments and professional vocabulary reduce reading friction, but they can also lower vigilance. A false detail embedded in accurate background is harder to detect than obvious nonsense. A nonexistent legal authority that follows familiar naming conventions can feel plausible before anyone opens a source.
Reviewers often answer the wrong question: “Does this sound reasonable?” The necessary question is: “Does independent evidence support every claim that could change the decision?” A plausibility read may catch contradictions and absurdity. It will not reliably catch an invented date, a transposed number, a real source cited for the wrong proposition, or an omitted exception.
Break the output into verifiable objects: facts, calculations, citations, inferences and recommendations. Verify decision-changing facts and numbers first; check whether sources support the neighbouring propositions; then assess whether advice fits the actual context. A single top-to-bottom read is not a substitute for these distinct tests.
The first answer anchors the judgement that follows
When a person sees the AI conclusion before examining evidence, their search can become an effort to confirm or slightly refine it. They notice supporting material and treat conflicting material as an exception. This is not simply laziness. The sequence of work has created a cognitive anchor.
For consequential judgements, ask the reviewer to record a short preliminary view from source material before revealing the AI recommendation. Alternatively, let one person prepare an AI-assisted draft and another reviewer assess critical claims without seeing the model's confidence language. An experimental study in recruitment found that erroneous algorithmic advice could still affect human judgement, demonstrating that keeping a nominal human role does not remove automation bias. Automation bias in AI-assisted recruitment study
Blind review need not apply to every sentence. Use it for high-impact conclusions, disputed cases and quality samples. Its purpose is to preserve at least one reasoning path that is not governed by the same initial answer.
The reviewer may not possess the knowledge the model lacks
Giving specialist output to a non-specialist reviewer merely transfers the unknown from the model to the person. A general employee cannot reliably validate a complex tax conclusion; an engineer may not know whether a privacy notice satisfies a legal obligation. An AI explanation can even create an illusion of understanding for someone unfamiliar with the field.
Define minimum reviewer competence for each task. A person trained in the process may check formatting, spelling and presence of required fields. Judging medical risk, legal duty, engineering limits or statistical method calls for relevant expertise. If no qualified reviewer is available, narrow the AI task instead of enlarging a generalist's nominal responsibility.
Experts still need tools, current material and time. Professional status does not automatically provide the correct policy version, complete input data or a reproducible calculation environment. A sound review design supplies capability, evidence and workable conditions together.
Review remains superficial when source material is hidden
AI summaries often remove qualifications, minority views, footnotes and date limits. If the interface displays only the summary, a reviewer does not know what disappeared. A citation link without a precise location can require searching a long PDF; merely opening it is not verification.
A better interface connects each critical claim to a source passage, data row or recording timestamp and reveals attachments the model did not read or failed to parse. For calculations, retain inputs, formulas and units. For translation, align source sentences and proposed text. For document comparison, display exact additions and deletions rather than only a change summary.
Evidence must be independent of the output under examination. One model generating an answer, summarising its supposed sources and then judging its own correctness is not three layers of assurance. It is one possible error expressed three ways.
Time pressure turns review into an acceptance-rate exercise
Suppose genuine verification takes five minutes per item while a system produces one hundred items an hour. No amount of reviewer training can satisfy that design. The response is not another reminder to be careful; production and verification capacity are mathematically incompatible.
Measure realistic review time before deployment, including opening evidence, resolving conflicts, recording a reason and escalating a difficult case. If throughput exceeds capacity, reduce generation, automate only unambiguous low-risk cases, prioritise by risk, add qualified reviewers or slow the service. Do not redefine impossible work as human oversight.
Watch unusually high acceptance rates, too. They may indicate an excellent system, but they may also reveal fatigue, missing override authority or pressure not to dissent. Interpret acceptance alongside seeded-error detection, amendment types, reviewer disagreement and later incidents.
Reviewing AI can cause a person to abandon a correct answer
A human–AI combination is not automatically better than either participant. When the human is right and AI is wrong, the person may surrender their judgement to confident output. When both rely on the same flawed record, they can agree and remain wrong. When an interface visually privileges the AI recommendation, alternative possibilities receive less attention.
Research on human-in-the-loop decision-making has found conditions in which adding human involvement did not improve model decisions and erroneous AI advice could reduce overall accuracy. The relevant evaluation unit is the complete workflow, not separate claims that “the model is accurate” and “a human reviews it.” When human-in-the-loop can decrease accuracy
During testing, insert known but unobtrusive errors and measure whether reviewers find them. A pass rate calculated only from ordinary production outputs does not reveal whether oversight works at the moment it is needed.
Ambiguous responsibility weakens the depth of review
When a model provider, system team, business operator and final approver each believe somebody else owns the result, review tends to become procedural. Approvers must understand which judgement they own. System owners should be accountable for data, versions, permissions and monitoring. The organisation must provide challenge and remedy. It is unfair to place every failure on the final clicker, but “the AI suggested it” cannot excuse the decision owner either.
For each task class, name a decision owner, a technical owner and an incident owner. The review record should capture the material reason for accepting, changing or rejecting an output without demanding so much low-value prose that everyone pastes a template.
How to design a stronger review process
First, tier the work by risk. Sample low-risk formatting suggestions; verify every high-impact fact and irreversible action. Second, write checks as precise questions—“Does the invoice support this amount?” or “Does the source directly support this sentence?”—rather than “Is this correct?” Third, expose source material, applicable rules and version information.
Fourth, create a usable rejection and escalation path, including authority to pause automation. Fifth, preserve model version, input scope, output, edits and final approver. Sixth, test review performance periodically with samples containing known errors and investigate why errors escape. Seventh, monitor complaints, reversals, rework and loss, because internal approval is not proof of external correctness.
The Australian Government's AI adoption implementation guidance emphasises use-case-specific testing, continuing monitoring, assigned accountability and effective human control. Together, these practices treat review as an operational mechanism that can be tested, not a sentence in a policy. Australian Government: Guidance for AI Adoption—Implementation Practices
My assessment: audit the review system as well as the model
Organisations often count model errors but rarely measure how many errors human review misses. If the human layer is the final safeguard, it has its own accuracy, coverage, capacity and failure modes. Useful measures include detection of seeded errors, verification of critical fields, disagreement between reviewers, realistic handling time, escalation rates and corrections after release.
Do not turn those measures into mechanical punishment. That would encourage reviewers to hide uncertainty. Use them to find inadequate training, poor interfaces, excessive workload and badly chosen automation boundaries. Sometimes the most valuable signal is not a high approval rate, but a reviewer who can safely say, “The evidence is insufficient; I cannot approve this.”
Review checklist
- Does the reviewer have competence to judge this content, not merely operate the tool?
- Can they see source material, complete context, applicable rules and document versions?
- Do they identify critical claims and test them against independent evidence?
- Does throughput allow real checking, recording and escalation?
- Can the reviewer reject AI without unreasonable performance pressure?
- Have you tested their ability to detect hidden errors, omissions and false citations?
- Do records preserve edits, disagreements, incidents and remedies—not only a click?
- Does someone periodically decide whether a task should be less automated or stopped?
Conclusion
Human review still misses AI errors because people can be anchored by the answer, lack expertise or original evidence, face impossible volume, or lack genuine veto power. Reliable oversight needs an independent judgement path, explicit verification targets, sufficient capability and time, traceable evidence, and tests of the review layer itself. The decisive question is not whether a human appears in the diagram. It is whether that person has the evidence, conditions and authority to stop an error before it becomes an action.
Related questions
- How Can You Verify an AI Answer Without Repeating All the Work?
- When Is the Cost of Error Too High for AI Automation?
Continue reading: All articles in How Far Should You Trust AI?
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.