Short answer
Compare the net result of the complete task, not model generation speed. Count preparation, waiting, verification, editing, rework, incidents and maintenance while measuring quality, throughput and employee burden. AI is efficient when it reduces total resources for equal or better quality, or produces a defined quality gain at an acceptable cost. The strongest practical method establishes a baseline and comparison over representative work, reports medians and failure tails, and repeats the measurement after several weeks.
“Generated in ten seconds” is not “completed in ten seconds”
AI demonstrations turn a blank page into a full draft and look fast. Real work includes locating material, cleaning files, writing instructions, checking figures and citations, adjusting tone, obtaining approval, copying the result into another system and dealing with follow-up. If generation takes ten seconds and verification forty minutes, while manual completion took thirty, the speed of generation produced no efficiency.
Define task time from the start of preparation until the result reaches a usable or approved state. Include active human time, elapsed waiting and downstream rework. Distinguishing active from elapsed time still helps capacity planning: AI may reduce writing labour while increasing approval queues, or shorten elapsed delivery without reducing staff effort.
Do not exclude time on the ground that “review was required anyway.” Compare the complete process before and after AI, not an AI step with all manual work.
Decide which form of efficiency matters
Efficiency can mean faster, cheaper, higher throughput, better quality, lower backlog, more consistent service, or more expert time for valuable judgement. A process that saves minutes and adds errors is not necessarily efficient. A process that takes the same time and produces materially more complete analysis may be valuable.
Choose a primary measure and safeguards for each use. Customer correspondence might use time to qualified resolution, guarded by accuracy and complaints. Code work might use features passing tests, guarded by defects and maintenance burden. Research summary might use coverage of material claims and source support, with reading time as cost.
Without predetermined outcomes, teams select flattering measures such as generated items, AI uses or acceptance rate. These measure activity, not value.
Establish a non-AI baseline instead of relying on memory
Select tasks representing the real distribution: routine, complex, incomplete and boundary cases. Have suitable people complete them through the existing process and record time, quality, rework and outcome. Then use AI assistance on comparable or randomly allocated work.
Account for experience. Novices and experts may gain differently; people familiar with AI differ from newly trained users. Letting only enthusiastic volunteers use AI creates selection bias. Where possible, randomise or cross over so the same participants use both workflows on different tasks.
A small team does not need a publication-scale experiment. Twenty or thirty real tasks in a category can expose a large difference. But do not choose only cases the model handles well or base long-term procurement on one afternoon's impression.
Measure total labour and elapsed cycle time
Break the workflow into source preparation, prompting and upload, generation wait, reading, fact verification, editing, approval, publication, rework and incident response. Human minutes can be costed, but different roles carry different opportunity costs. Specialist review cannot be priced as though it were interchangeable with junior time.
Elapsed time includes queues and dependencies. AI may finish drafts earlier while concentrating every task on one reviewer, lengthening the queue. Average task labour falls while customers wait longer.
Include system maintenance: instruction templates, connectors, permissions, evaluation, model upgrades, provider management and training. Pilots routinely omit these shared costs; they become visible at scale.
Review cost depends on consequence and verifiability
An internal invitation is quick to check. A legal research draft containing fifty citations may take longer to verify than to write manually. The more AI generates plausible complete content, the larger the potential inspection surface. Word count is not a uniform unit of value.
Tier claims into formatting and style, ordinary facts, decision-critical facts, professional judgement and irreversible action. Record verification time and escaped-error risk for each. Good AI candidates tend to have clear input, cheap verification and reversible errors. Bad candidates may be fast to generate and slow to prove.
Sampling can reduce inspection, but sample strength should come from risk and observed error. Removing review to manufacture savings merely postpones cost until complaint or incident.
Compare quality together with time
Slightly slower AI-assisted work may be worthwhile if quality improves; faster work that causes repeated customer clarification is false economy. Build a rubric covering correctness, completeness, relevance, actionability, tone, sources and compliance.
Where possible, blind reviewers to workflow to reduce expectations about AI. Use tests, reconciliation or real outcomes for objective work, and multiple reviewers with explicit criteria for open work. The model that generated an answer should not be its only judge.
Business outcome matters more than an internal prose score: Was the customer issue solved? Was the proposal approved? Did defects fall? Did people make better decisions? Outcomes can lag, requiring continuing observation.
Averages hide the most expensive failures
AI might save ten minutes on eighty per cent of simple tasks and create two hours of rework on five per cent. An average alone conceals the tail and operational volatility. Report median, 90th or 95th percentile, failure rate and credible maximum loss.
Distinguish correlated failure. A person's mistake generally affects one item; a faulty automated template can affect every output. Do not discard a mass correction as an outlier when it is a real property of scale.
If a system delivers small stable savings with occasional extreme loss, reduce its automation or add caps and approval. Efficiency cannot be separated from risk tolerance.
Published studies cannot replace local measurement
One study of customer-support workers found an average productivity increase of roughly fourteen per cent from generative AI assistance in that setting, with larger gains among less experienced or lower-skilled workers. This is important evidence about particular workers, tasks, tool and organisation. NBER: Generative AI at Work
Another field experiment across 7,137 knowledge workers at 66 firms found that generative-AI users spent about two fewer hours on email each week, but detected no significant change in task quantity or composition. A local time change does not necessarily become an organisational output change. NBER: Shifting Work Patterns with Generative AI
The results need not conflict. Benefits depend on work type, implementation, people and whether the organisation can redeploy saved time. A procurement demonstration or external average cannot replace a local baseline.
Novices may gain more and still find errors less readily
AI can make experienced patterns available as suggestions and narrow some performance gaps. Less experienced people can also be less able to detect invention, wrong terminology and inapplicable advice. Faster completion does not guarantee long-term skill formation.
Measure time, quality, review requests and learning separately by experience. If novice output improves but dependence rises, use independent first attempts, reason explanations, worked comparisons and mentor samples. AI can be scaffolding rather than a permanent replacement for foundational judgement.
Watch whether expert time merely moves. Ten juniors each save an hour, while one expert gains fifteen hours of review and becomes a bottleneck; the total system is worse.
Acceptance rate is not a reliable proxy for savings
High acceptance may indicate strong suggestions, or fatigue, time pressure, default UI choices and inability to spot errors. Low acceptance does not necessarily mean failure; an AI draft may help a person reach the final version faster even when much wording changes.
Better measures include time to approved state, number and type of material edits, escaped errors, rework cycles and task outcome. Separate stylistic changes from corrections to facts, responsibility or conclusion. The latter indicate higher verification burden and risk.
Count output that is generated and discarded. It still consumes time, compute and attention and cannot disappear from the efficiency denominator.
Where did the saved time go?
If goals, capacity and scheduling do not change, employees may spend saved time on more communication, additional polishing, waiting or new work. There is no automatic conversion from personal minutes to organisational profit. Before a pilot, decide the intended use: faster customer service, less overtime, backlog reduction, deeper analysis or protected learning time.
This does not require surveillance of every employee minute. At process level, examine whether the bottleneck, output, quality or workload changed. Once drafting stops being the constraint, approval, data access or customer response may become the new bottleneck.
Some benefits are resilience: maintaining service during peaks, supporting cross-language work or reducing blank-page friction. Include them where relevant, but define and measure them rather than labelling every intangible benefit “innovation.”
Calculate total cost of ownership
Beyond subscription or API price are integration, permissions, data governance, training, evaluation, monitoring, review, provider due diligence, model migration, incidents and exit. A free product consumes employee time and may add data risk.
Allocate shared cost monthly or quarterly and divide by qualified tasks actually completed. Low usage can make unit cost far higher than the licence price; at high scale, review and incident cost can dominate model calls.
Include the opportunity cost of not adopting. A fully manual alternative may leave backlog and response times worsening. Compare AI with credible alternatives, not idealised zero-cost manual work.
Use staged measures to avoid novelty effects
The first week may be slow because people learn, or unusually strong because of enthusiasm and extra support. After four to eight weeks, prompting, allocation and review habits settle. Months later, model updates and maintenance become visible.
Measure pilot, stable use and scaled operation. Record training time and whether benefits remain concentrated among a few power users. Mature templates and workflows can become standard operations; constant individual repair means the system is not yet productised.
Define stopping criteria: quality below baseline, review backlog, a sensitive incident, cost beyond budget or increased employee burden. A valid pilot can discover that scaling is inappropriate; it need not exist only to justify adoption.
My assessment: measure “completed and trustworthy”
Generated words, summaries and suggestions reward volume without asking whether it can be used. A stronger unit is an approved customer response, report, code change or decision that meets its quality threshold, with total time and cost per qualified unit.
For material knowledge work, also measure the treatment of uncertainty. A system that asks when evidence is insufficient may seem slower than one that answers immediately while preventing downstream error. A reliable stop can be genuine efficiency.
Efficiency checklist
- Is there a measured non-AI baseline rather than subjective memory?
- Does timing include preparation, prompting, verification, editing, approval, rework and maintenance?
- Are quality, business outcome, escaped error and risk compared at the same time?
- Are tasks and users grouped by complexity and experience?
- Are median, tail failures and correlated mass errors reported?
- Did savings simply move work to experts or downstream teams?
- Are model, integration, training, governance, migration and incidents in total cost?
- Will measurement recur after stable use, with explicit scale and stop thresholds?
Conclusion
To know whether AI improves efficiency, move from “how fast did it generate?” to “how well did a trustworthy result reach completion?” Establish a non-AI baseline, measure the complete workflow, quality and tail risk, and observe whether saved time actually improves throughput, waiting or workload. AI sometimes delivers substantial gains and sometimes exchanges writing time for review time. Only local comparison and continuing measurement distinguish them.
Related questions
- Which Tasks Should You Give to AI, and Which Should You Keep?
- Why Does Human Review Still Miss AI Errors?
Continue reading: All articles in How Far Should You Trust AI?
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.