Short answer
Do not redo an answer line by line. Identify the load-bearing claims that could change the decision, rank them by consequence, and choose an independent verification method for each type. Return dates, figures and citations to original records. Use deterministic tools for calculations. Compare summaries with decisive passages. Test recommendations against their assumptions and counterexamples. Verification is not proof that AI never fails; it is a time-bounded process for catching errors capable of changing the outcome.
Start with how the answer will be used
The same paragraph requires different review when it is private brainstorming and when it is advice to a client. If the output remains in a draft and any mistake is reversible, a quick check may be proportionate. If it affects payment, legal rights, health, safety or public reputation, it needs stronger evidence, qualified review and records.
Verification cost should follow consequence, not word count. A two-thousand-word creative draft may need only a directional read. A single incorrect bank account number must be checked character by character. Decide the use before allocating attention.
Write the consequence in concrete terms. “Could be inaccurate” is too vague. “Could cause us to miss the filing deadline”, “could disclose a client's diagnosis”, or “could overpay by $20,000” identifies what the review must prevent.
Build a claim list instead of reading from the first sentence to the last
Mark content that can be true or false: names, dates, figures, events, laws, product functions, quotations and causal claims. Then circle claims that would alter the conclusion or action if wrong. In a comparison of two services, interface colour is not load-bearing; price, cancellation terms, retention and required capability probably are.
Treat the rest as supporting context or expression. Checking the decisive five claims can be more effective than evenly sampling twenty sentences. This is not permission to publish known minor falsehoods. It is a method for deciding where verification begins and when an answer should be rejected altogether.
You can ask AI to produce a claim table containing type, source, supporting location, consequence if wrong and verification state. The table is itself a draft. It organises the review; it does not perform it.
Match the verifier to the claim
Return factual claims to primary sources: legislation, government pages, official product documentation, original research, contracts, meeting invitations and database records. A second blog repeating the first claim is not independent evidence. For changing information, check publication date, effective date and version.
For a number, locate the source and examine denominator, unit, period and definition. A twenty per cent rise may be from five to six. “Ninety-five per cent accuracy” may come from a test that does not represent the deployment. Recalculate decisive derived figures with a calculator or spreadsheet rather than asking the same language model to calculate again.
For a summary, return to the passages that matter to the decision instead of rereading every page. Check whether the summary preserved exceptions, negation, thresholds, responsible actors and time conditions. If a summary says a policy “allows” something, search the source for “must”, “may”, “except” and “unless”.
A recommendation has no single source to look up. Inspect its objective, assumptions and trade-offs. Rewrite it as a conditional: “If the goal is X, the constraint is Y, and risk tolerance is Z, choose A.” If the conditions do not hold, the recommendation cannot be adopted directly.
Change the evidence path instead of repeating the prompt
Asking “Are you sure?” in a new chat may produce a longer version of the same answer. Even two models can rely on overlapping public material and make correlated mistakes. Independent verification changes the path: open the original document, query an official register, run a test, recalculate, or ask a competent person.
Lateral reading is a useful discipline. When professional fact-checkers encounter an unfamiliar website, they leave it and investigate what other sources say about its publisher, author and claim. A cluster-randomised classroom study by Stanford researchers found that teaching strategies based on lateral reading significantly improved students' ability to judge online credibility. Stanford Civic Online Reasoning: Lateral Reading on the Open Internet
Apply the same method to AI. Do not remain inside its narrative and references. Start a separate path from the relevant agency, original paper, legislation database or product change log.
Use a two-minute, ten-minute and thirty-minute ladder
A two-minute check suits low-risk output. Identify load-bearing claims, click every citation, confirm that pages exist and titles and dates make sense, search one distinctive name or figure, and check whether fact and advice have been blurred. Any broken link or title mismatch raises the review level immediately.
A ten-minute check suits ordinary content leaving the organisation. Read the directly relevant source passages, recalculate key figures, compare one independent authoritative source, check version and jurisdiction, propose a counterexample that would defeat the conclusion, and confirm that no sensitive information has been included.
Thirty minutes or more is appropriate for consequential work. Obtain the complete record, use a qualified reviewer, test boundaries and known failures, preserve evidence and the system version, and add a second approver where needed. These times are not universal standards. The ladder prevents every draft from receiving an audit while ensuring high-impact output does not receive only a glance.
Escalation should be automatic where possible. If a citation does not resolve, a required field is absent, a total differs, or the model expresses unsupported certainty, the workflow should stop rather than rely on a hurried user to notice.
Sample intelligently
Random sampling estimates ordinary error but may miss a systematic weakness. Combine random items with the highest-value or highest-impact cases, rare cases, incomplete inputs, different languages or groups, and cases the system found uncertain.
Also sample outputs that look most professional, because fluency suppresses suspicion. When a load-bearing claim is wrong, do not patch only that sentence. Check whether the same generation process affected related items and reduce confidence in the batch until you know the scope.
For repeated work, compare errors by category: invented fact, wrong source, omitted exception, calculation, misclassification, privacy exposure or unsuitable advice. Trend data tells you which control should change. A single overall “accuracy” percentage does not.
Verify the transformation, not merely the final prose
Many AI tasks transform A into B: document to summary, recording to minutes, requirements to code, or data to chart. Create traceability between them. Each important summary conclusion points to a source location. Each decision in minutes points to a recording timestamp. Each code change maps to a requirement and test. Each plotted value maps to a source cell.
This mapping is stronger than “check your work” because it shortens the route back to evidence. The model can still provide an incorrect page or timestamp, so sample the mapping. Once the structure exists, however, review no longer requires consuming the complete source again.
Ask for an omissions list as well as a summary. The model should identify attachments it could not read, illegible sections, conflicting statements and decisions that were implied but not spoken. An explicit gap is safer than false completeness.
Use deterministic tools for deterministic questions
A language model can explain a calculation but should not be the only verifier of critical arithmetic. Totals, tax rates, date differences, file hashes, format constraints, duplicates and program behaviour are better checked with calculators, spreadsheet formulas, database queries, test suites or rules.
Likewise, a paper's existence can be checked through DOI resolution, Crossref, the publisher or a scholarly index. A URL can be opened. A current policy can be checked on an official version page. Automating what can be determined allows human attention to remain on meaning, exceptions and values.
Deterministic checks should fail visibly. A quiet warning beside a polished answer is easily ignored. Prevent the next action when an amount does not balance, an identifier fails validation or a required citation is absent.
Verify that the answer addressed the right question
Every fact can be correct while the answer is irrelevant. AI may interpret “cheapest” as purchase price without maintenance, cite United States law for an Australian question, or answer account deletion with instructions for uninstalling an app. At the beginning of review, compare the original question with the output's objective, jurisdiction, date, subject and constraints.
Check missing inputs. If the recommendation requires age, budget, contract term or risk tolerance and none was provided, the result is only a conditional draft. A good answer identifies what it needs. A suspiciously complete one may conceal assumptions it supplied itself.
Restate the question in a one-line acceptance test: “This answer must compare total three-year cost for an Australian customer as at this date.” If the output cannot be judged against that line, the task was not defined well enough.
Preserve verification so the next check is cheaper
For repeated work, maintain an approved source set, check rules and a library of past failures. Record which official page establishes each type of fact, which fields must be recalculated and which prompts cause omissions. When a source or model changes, retest affected parts rather than rebuilding everything.
This determines whether AI truly saves time. If every output requires research from zero, the tool may not reduce total cost. When claims are traceable, deterministic elements are checked automatically and established sources can be reused, human verification can become materially faster than performing the original task.
Do not let reuse turn into stale authority. Give every source set an owner and review trigger. A superseded law or product document can efficiently produce consistently wrong checks.
My assessment: verification searches for failure, not the appearance of diligence
Effective verification is adversarial: where is this most likely to fail; which failure changes the decision; what evidence could disprove it fastest? This is more reliable than reading gently and waiting for something to feel wrong. The objective is to maximise the chance of finding consequential error with limited resources.
If output cannot be decomposed into decisive claims, has no traceable basis and lacks a defined use, the problem is not merely high verification cost. It is not ready to enter a decision. Rejecting it may be more efficient than attempting repair.
Rapid verification checklist
- What is the intended use and concrete consequence of error?
- Which three to five claims can change the decision?
- Is each a fact, calculation, summary, inference or recommendation?
- Did the check use an independent evidence path rather than a repeated AI answer?
- Do citations exist, remain current and directly support the sentence?
- Are units, denominators, jurisdiction and period consistent?
- Did the summary preserve exceptions, negation, thresholds and responsible parties?
- Which counterexample or boundary condition would defeat the conclusion?
- Can this verification be recorded and reused?
Conclusion
Verifying AI does not require repeating the whole task. Identify the use and load-bearing claims, then apply primary sources, deterministic tools, counterexamples and expert judgment according to risk. Independence matters more than repetition; critical positions matter more than uniform sampling. A good AI workflow does not merely make belief faster. It makes important claims easier to disprove, confirm and trace.
Related questions
- Why Can AI Sound Confident and Still Be Completely Wrong?
- How Can You Separate Facts, Inferences and Advice in an AI Answer?
Continue reading: All articles in How Far Should You Trust AI?
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.