After AI Enters the Workflow · Season Two: “From Personal Tool to Organisational Capability” · Article 5
An organisation chooses a model that ranks highly on public benchmarks and uses it to handle internal technical requests. The test team prepares fifty standard questions, and the model achieves a strong score. In production, employees discover that the system often fails on the requests that are genuinely difficult: information is dispersed, terms differ among departments, users omit important context, and a correct result may require several systems.
The benchmark was not fraudulent, and the model did not suddenly become less capable. The benchmark answered, “How does the system perform under specified inputs and scoring?” Real work asks, “Can this complete workflow succeed amid untidy records, actual users, existing permissions and consequences?” The missing element is not another general score. It is context.
Why benchmarks still matter
Benchmarks give different systems the same tasks and compare them through a common method. They are useful for tracking capability, screening candidate models, identifying obvious weaknesses and subjecting vendor claims to some external scrutiny.
Without standardised tests, organisations would depend on demonstrations and impressions. The mistake is not using a benchmark. It is treating a benchmark result as an adoption decision.
A leading result in mathematics, coding or question answering does not establish that a system can use an organisation’s records, enforce its permissions, detect unstated exceptions or preserve an audit trail. A benchmark measures the property it was designed to measure, not reliability in the abstract.
Real work introduces four additional layers
The first is input variation. Employee questions are ambiguous, spelling differs, documents are obsolete and critical information may be absent.
The second is system variation. A deployed product includes retrieval, instructions, tools, permissions, caches and interfaces. The model is only one layer.
The third is interaction variation. Users may over-trust, misunderstand, retry or apply AI to work that evaluators did not anticipate.
The fourth is consequence variation. A benchmark error loses a point. A production error may disclose information, delay service or enter an official decision.
NIST’s ARIA program explicitly combines model testing, red-teaming and field testing, including observation of human–AI interaction and impact in realistic settings. NIST, “Assessing Risks and Impacts of AI” Its pilot report describes scenarios, dialogue annotation and measurement trees used to connect evaluation to context. NIST, “ARIA Pilot Evaluation Report”
Test how work fails, not only how it succeeds
An evaluation built entirely from ordinary inputs overestimates performance at the edges. Real tests should include obsolete documents, conflicting rules, missing material, unauthorised access, malicious content, tool timeouts and impossible requests.
They should also ask whether the system knows when to stop. A model that answers every question may have a higher apparent completion rate than one that says, “The evidence is insufficient; obtain human confirmation.” The first can nevertheless create more organisational risk.
Scoring must distinguish types of error. A punctuation mistake in a customer letter cannot be averaged against an incorrect statement of legal rights. High-consequence failure should be evaluated separately and may need to block release regardless of the overall score.
Why expert work differs from standard tasks
In 2025, METR studied experienced open-source developers working on large repositories they knew well. With the tools then available, AI use made participants slower on average. The researchers noted that common coding benchmarks often use self-contained tasks for scalable evaluation, while real work requires extensive prior context. METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”
By 2026, tools and work habits had changed. METR altered its subsequent experimental design and discussed biases created when developers who depended on AI were reluctant to participate in a no-AI condition. METR, “We are Changing our Developer Productivity Experiment Design”
Together, these materials show that capability changes and that evaluation is affected by the environment it tries to measure. An organisation cannot preserve one historical score indefinitely as the reason for procurement or permission.
Move from model selection to use-case validation
A useful organisational evaluation should sample real work, remove unnecessary personal information and retain genuine difficulty and exceptions. The object under test is not a bare model. It is the deployed configuration: system instructions, retrieval collection, tools, permissions, interface and human stages.
At least four groups of measures are needed. Was the task completed correctly? Were important constraints observed? Were evidence and records preserved? When uncertainty or failure occurred, did the system escalate appropriately?
Results should be broken down by user experience, task type and affected groups. An average can hide repeated failure on a small class of complex cases.
NIST’s work on AI test, evaluation, validation and verification emphasises that how a component is measured can change according to the context in which the AI system operates. NIST, “AI TEVV” Evaluation is therefore not an examination completed once before procurement. It is an operating function of the use case.
Use shadow operation before granting consequences
Deployment should not jump directly from a test environment to automated decisions. A system can first run in shadow mode: it processes live tasks but does not change the existing workflow. The organisation compares AI suggestions, original human decisions and final outcomes, recording differences and their causes.
AI may then take on drafting and low-risk classification while explicit approval remains. Permissions should expand only when error rates, error types, monitoring and remediation satisfy agreed criteria.
After release, sampled evaluation continues. Material changes to the model, instructions, knowledge base or user population trigger renewed testing. Without change records, a decline in performance cannot be attributed or repaired.
The evaluation set also requires governance
A use-case test is closer to work than a public benchmark, but it also becomes stale. Employees learn common cases, policy and customer populations change, and new failure modes may never enter the sample. Repeated optimisation against one fixed set can improve the score without improving reality.
The set therefore needs provenance, version, ownership and update rules. Some cases should remain independent of development. Complaints, human corrections and operational anomalies should supply new examples, with sensitive information removed or controlled. Adding a serious failure to the test does not itself solve it; the organisation must show that the workflow and remedy have changed.
Ground truth also needs a source. If qualified experts reasonably disagree, evaluators should not force a false single answer. They should test whether the system recognises the dispute, preserves the competing evidence and escalates to the authorised role. Evaluation quality depends on the governance of answers as much as the selection of questions.
A real-work evaluation should also define a failure budget in advance. Not every error has the same consequence. An awkward phrase in an internal draft may be corrected cheaply; wrongful denial of eligibility or disclosure of protected information cannot be offset by a high average score. The organisation should set tolerances for distinct failure types, immediate stopping conditions and recovery steps. Evaluation then becomes an authorisation proportionate to a particular purpose and consequence, rather than the vague conclusion that performance is good overall.
Conclusion: benchmarks compare; workplace evaluations authorise
A benchmark result is useful evidence when selecting candidate technologies. It does not directly confer permission to enter real work.
The organisational distinction should be:
A public benchmark asks what a model can do on standard tasks. A contextual evaluation asks what this complete AI system may be allowed to do in a particular workflow, for particular users and consequences.
Only the second can support decisions about access, release and oversight. The stronger the model becomes, the less adequate it is to substitute a general score for validation of the actual workflow. Stronger generation also makes an error easier to present as work ready for adoption.
Primary sources and further reading
- NIST: Assessing Risks and Impacts of AI
- NIST: ARIA Pilot Evaluation Report
- NIST: AI Test, Evaluation, Validation and Verification
- METR: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- METR: We are Changing our Developer Productivity Experiment Design
Continue reading: Explore the After AI Enters the Workflow series.
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.