After AI Enters the Workflow · Season Two: “From Personal Tool to Organisational Capability” · Article 10
A department produces thousands of AI summaries, classifications and recommendations each day. Management knows that they cannot all be trusted, but no reviewer can read every item. The department samples one per cent. Several weeks pass without a severe error, and the system is described as “subject to human oversight”.
Later, the organisation discovers that a rare class of customer has been repeatedly misclassified. Because the cases represent a small share of total volume, random sampling seldom selected them. For the affected people, however, the error was systematic.
Auditing cannot mean glancing at a few items from a large pile. It must be designed around risk, groups, failure modes and change. It should demonstrate that important problems can be discovered, not merely common ones.
Reviewing every item is neither possible nor necessarily effective
Human inspection of every artefact would remove much of AI’s scale advantage. People reviewing repetitive queues also fatigue and approve mechanically. Effective audit does not add an unlimited workforce at the end. It places different controls at different stages.
Objective properties such as format, required fields, numerical ranges and permissions can be checked automatically. High-consequence decisions receive qualified review before adoption. Low-risk, high-volume work is monitored through stratified samples. Complaints, anomalies and system changes trigger targeted examination.
The audit object is also larger than the text. It includes the input source, configuration, tool calls, approvals and actual outcome.
Risk tiers determine audit intensity
Use cases should first be classified by consequence, reversibility, visibility of error and affected population.
Low-risk internal drafting can rely mainly on automatic controls and periodic samples. Formal customer communications and outputs affecting money, employment or rights require higher coverage, professional review and complaint records. Irreversible or safety-critical action cannot rely on after-the-event sampling instead of preventive control.
The NIST AI RMF calls for the level of risk-management activity to reflect organisational risk tolerance and for governance, measurement, monitoring and feedback to continue across the lifecycle. NIST AI RMF Core
When audit capacity is scarce, it should first protect the most consequential work rather than the most numerous work.
Combine random and targeted sampling
Random samples help estimate aggregate error, but can miss small groups and rare failures. Stratification is also needed by task type, language, region, user group, monetary value, model confidence condition and degree of human editing.
A failure-oriented stream should include complaints, overturned decisions, repeated generations, tool timeouts, permission denials, unusually long inputs and refusals. Every major incident should produce a new audit rule or test case.
The US Government Accountability Office’s AI Accountability Framework organises oversight around governance, data, performance and monitoring and provides questions for managers, auditors and third-party assessors. US GAO, “Artificial Intelligence Accountability Framework” Its value lies in examining goals, data quality, continuous monitoring and responsibility rather than accuracy alone.
Records must be sufficient for audit
If the system retains only the final prose, an auditor cannot determine whether failure originated in a source, retrieval, model, tool or human edit.
Consequential use cases should preserve a proportionate task identifier, time, configuration version, source references, tool actions, approval and final outcome. Logs may themselves contain sensitive material and therefore require minimisation, restricted access and retention limits.
The record also needs a route to remedy. When a pattern is discovered, the organisation should be able to locate similar cases, suspend the configuration and reprocess affected work—not merely edit the sampled item.
Australia’s government AI policy requires covered agencies to maintain internal use-case registers, assign use-case accountability and conduct impact assessments. These requirements provide the minimum inventory from which audit can begin. Digital Transformation Agency, “Policy for the responsible use of AI in government”
An organisation that does not know which AI uses are operating cannot audit their outputs coherently.
Independence should be proportionate to risk
Developers understand the system but may be too familiar with their assumptions. Business staff understand the process but may mistake a longstanding practice for the correct standard.
A low-risk use can be reviewed by another team member. Higher-risk systems require independent security, legal, professional or internal-audit participation. Particularly consequential uses may need external assessment or feedback from affected groups.
Independence does not require ignorance of the system. It means that audit criteria and conclusions are not controlled solely by the team being assessed and that findings have a direct escalation route.
ISO/IEC 42001 incorporates performance evaluation, internal audit and continual improvement within an AI management system. ISO, “ISO/IEC 42001” Audit is therefore not a pre-release stamp. It is part of a management cycle.
Audit whether human oversight is real
A field stating “human approved” does not demonstrate meaningful oversight. Auditors should compare approval times, modification rates, workload, whether sources were viewed and whether different reviewers detect seeded errors.
If a person approves dozens of complex recommendations each minute, oversight may exist only as metadata. The organisation should reduce the volume requiring judgement, improve the interface or move the automation boundary. A human name should not be used as a certificate for an unreviewable process.
Audit should also examine whether affected people can complain, how quickly they receive a response and whether errors are reversed. Without remediation data, the organisation sees only internal procedure and not real consequence.
The audit system must itself be evaluated
If automatic controls generate large volumes of irrelevant alerts, teams learn to ignore them. If sampling rules remain unchanged for years, new failure modes escape. The organisation should periodically seed known problems and measure whether controls find them, including the time from anomaly to detection, escalation and correction.
Audits also create false positives. If normal cases from one group are repeatedly flagged, those people may experience additional investigation or delay. Audit design therefore needs its own analysis of group effects and proportionality. “Risk control” should not become a second unfair classification system.
The assurance function needs outcome measures: detection of serious issues, unresolved backlog, repeated incidents, completed remedies and whether findings actually changed the system. Counting reviews merely repeats the production layer’s mistake of confusing volume with effect.
Reports to use-case owners and leadership should also state the residual unknowns: domains without enough samples, missing logs and supplier changes that cannot be independently verified. An honest audit explains what remains invisible.
That disclosure is operationally useful. It tells leaders where automation must remain constrained, where extra evidence should be collected and which assurances should not yet be repeated to customers or regulators.
Audit findings also require remediation capacity. If a team can accumulate defects but lacks authority to pause the process, notify affected people or arrange reprocessing, more audit merely creates a better record of institutional powerlessness. The organisation should reserve people and funding for serious findings, impose correction deadlines and trace whether affected outcomes were actually restored. A failure that is discoverable but not remediable has not yet been governed.
Conclusion: audit makes important failure visible
Comprehensive human inspection cannot scale with AI output. The answer is not to remove oversight, but to combine layered controls with purposeful sampling.
The final principle is:
Automate checks with objective criteria, preserve professional judgement before high-consequence action, monitor both the whole and minority groups through random and stratified samples, and let complaints, anomalies and changes continually reshape the audit scope.
Auditability is not measured by log volume. It is measured by whether the organisation can use records to discover systemic harm, identify affected cases, stop the failure and prove that the correction worked.
Primary sources and further reading
- US GAO: Artificial Intelligence Accountability Framework
- NIST AI Resource Center: AI RMF Core
- Australian Government: Policy for the responsible use of AI in government
- ISO: ISO/IEC 42001—AI management systems
Continue reading: Explore the After AI Enters the Workflow series.
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.