After AI Enters the Workflow · Season One: “From Answering Questions to Participating in Work” · Article 1
Imagine an employee preparing a supplier assessment. He could ask an AI system, “Which indicators are normally used to assess a company’s financial stability?” The system might give him a well-organised list covering cash flow, leverage, profitability and payment history. The answer could be useful, but the AI has not yet entered the actual work. It has answered a general question detached from the suppliers being assessed.
Now suppose he provides the financial statements of three suppliers, the organisation’s procurement policy, its risk criteria and the reporting template. He asks the AI to compare the figures, identify missing documents, prepare a draft and attach a source to every conclusion. The output is no longer merely general advice. It is a candidate work product intended to enter a real file.
The situation changes again if the AI can sign in to the procurement system, request missing information from a supplier, alter a risk status or submit the assessment for approval. An error is no longer confined to a chat window. It can change an official record, trigger a communication or affect a transaction.
All three situations are commonly described as “using AI”, but they involve very different degrees of delegation. The first treats AI as a reference tool, the second inserts it into production, and the third permits it to act. Unless those positions are distinguished, a persuasive answer can easily be mistaken for reliable work, and the ability to generate can be confused with the authority to decide.
Answering a question is not completing a task
A question usually ends when an adequate answer has been supplied. A task has a longer structure: why it is being performed, which materials apply, which rules must be observed, what artefact must be produced, how completion will be demonstrated, who will approve the result and how an error can be reversed.
“What documents are generally required to register a company in Australia?” is a question. “Prepare the registration material for this applicant and confirm the proposed name, share structure, directors’ eligibility and declarations” is a task. The second requires more than relevant knowledge. It requires correct identification of the parties, current rules, treatment of missing information, consistency across documents and clarity about who is authorised to make a legal declaration.
Likewise, “How should a rejection email to a job applicant be written?” is a question. “Send the formal outcome of this recruitment process to this applicant” is a task. The wording is only one link in the chain. The system must know that a decision has actually been made, which reasons may be disclosed, whether the recipient is correct and what record must be retained.
A task therefore introduces at least five matters that ordinary question answering does not:
- State: which stage the work has reached and what has already been completed;
- Context: which facts, rules and prior records govern this particular case;
- Tools: which external objects the system may read or change;
- Completion criteria: which evidence proves that the task was completed correctly;
- Consequences: who is affected when the result is adopted and who must repair a mistake.
AI enters work when its output acquires a recognised place in this chain—not when its answer becomes longer or sounds more expert.
From conversational tool to workflow participant
Discussions of AI often use “workflow” and “agent” as if they meant the same thing. They do not.
In a predetermined workflow, people or conventional software have already selected the sequence: collect documents, classify them, prepare a draft and obtain approval. A model may perform one or more stages, but the route is largely imposed from outside.
An AI agent receives more discretion over process. Anthropic’s work on trustworthy agents describes agents in terms of systems that can decide how to pursue a task and use tools, rather than merely execute a fixed script. The important feature is not conversational fluency; it is the ability to choose the next action. Anthropic, “Trustworthy agents in practice”
An agent might search for material, open a file, notice that evidence is missing, revise its plan, call another tool and then prepare a result. It maintains a goal, a working state and a chain of actions over time.
Greater resemblance to work, however, does not mean greater trustworthiness. A small error can propagate through the chain. The wrong source contaminates the analysis built on it. The wrong customer identity turns an otherwise correct email into a privacy incident. A misunderstood file path can cause a precise edit to damage the wrong object.
The central change introduced by agents is therefore not that “AI has become a colleague”. It is that software is receiving a measure of process discretion previously exercised by an operator. Permissions, records, checkpoints and stopping mechanisms must grow with that discretion.
What productivity studies do—and do not—show
There is credible evidence that generative AI can improve performance on some tasks. In an experiment by Shakked Noy and Whitney Zhang, university-educated professionals completed mid-level occupational writing tasks. Participants with access to generative AI finished faster on average and received higher quality ratings. Their time shifted away from rough drafting and towards planning and editing. Science, “Experimental evidence on the productivity effects of generative artificial intelligence”
A study of customer-support workers found that an AI assistant increased the number of issues resolved per hour, with larger gains among less experienced workers. One explanation was that the tool made effective practices embedded in prior conversations available at the moment of work. NBER, “Generative AI at Work”
These studies matter because they show that AI is not confined to impressive demonstrations. It can redistribute labour in real or realistic work settings. They do not establish that “AI can complete knowledge work” in general. The writing experiment did not require mastery of an organisation’s complete history or make rigorous fact-checking the central task. The support environment contained recurring products and questions. Each study measured a particular tool, task, population and period.
Another field experiment introduced the idea of a “jagged technological frontier”: on tasks within the model’s capabilities, AI could improve speed and quality, while on apparently similar tasks beyond that frontier it could make users more likely to reach incorrect answers. Harvard Business School, “Navigating the Jagged Technological Frontier”
Once AI enters a workflow, productivity cannot be measured only by words generated or minutes saved. We must ask what kind of task it entered, whether completion is independently testable, whether failure will be visible and at what point a person can still intervene.
Work moves rather than simply disappears
When AI takes over drafting, classification or search, human labour often moves elsewhere.
A report writer may spend less time arranging sentences but more time checking whether the record is complete, whether the criteria apply and whether the conclusion follows from the evidence. A programmer may type less code while retaining responsibility for requirements, tests, architectural fit and deployment. A support worker may receive a ready-made answer while still judging exceptions, sensitive customers and the limits of what the organisation may promise.
This is often described as a move from doing to checking, but effective checking is not easy. A reviewer must understand the task, have access to the underlying evidence, possess enough time to examine it and be authorised to reject the output. If a person merely approves a queue of AI-produced work, “human oversight” becomes a way of placing a human name on an automated result.
AI consequently creates or enlarges four less visible forms of work:
- Task design: turning an ambiguous request into bounded work with testable completion criteria;
- Context management: deciding what the model may see and which material has authority;
- Result evaluation: using evidence, tests and professional judgement to distinguish fluency from correctness;
- Responsibility control: deciding which actions may proceed automatically and which must stop for approval.
These functions may be distributed across business staff, technical teams, security, compliance and management. The interface may be a single chat box, but the actual system consists of many design and authorisation decisions.
Three levels of entry into work
A practical distinction can be made among three levels.
At the first level, reference, AI supplies explanations, examples and suggestions. The user may disregard them completely, and no official record changes automatically.
At the second level, production, AI output becomes a candidate report, program, analysis or communication. It participates in work, but the candidate must pass an identifiable process of verification and approval.
At the third level, execution, AI can call external tools, alter files, send messages, submit transactions or trigger another process. This level requires least privilege, activity logs, reversible operations, risk classification and explicit human checkpoints.
The model’s brand or nominal capability does not determine the level. The same model can be a low-risk reference tool in one setting and an executor with access to email or databases in another. Risk depends less on how intelligent the system sounds than on the position it occupies.
The US National Institute of Standards and Technology similarly treats generative-AI risk across the lifecycle, asking organisations to establish context, measure risk and govern use rather than judge a system from one answer. NIST, “Generative Artificial Intelligence Profile”
Conclusion: AI enters work when its output acquires an institutional position
When AI answers a question, it supplies language that may be useful. When it undertakes a task, it receives some combination of goals, materials, procedures, tools and permissions, and produces an artefact intended for a real workflow. The dividing line is not answer length or the use of labels such as “reasoning” and “agent”.
The real test is this:
Has the AI’s output become a recognised intermediate product, or can its action change the real state of the work system?
If the answer is yes, the organisation cannot continue to manage it as an ordinary chat tool. The task needs boundaries, context needs provenance, completion needs evidence, action needs permission, errors must be stoppable and reversible, and an identified person or institution must remain responsible for adoption.
AI does not thereby become a worker in the legal or moral sense. People and organisations have delegated work functions to models and software. The more steps those systems can complete independently, the less adequate it is to ask only whether they can do the work. We must also ask how we will know that it was done correctly—and who will act when it was not.
That is the threshold between answering questions and participating in work, and the starting point for the rest of this series.
Primary sources and further reading
- NIST: Artificial Intelligence Risk Management Framework—Generative Artificial Intelligence Profile
- Anthropic: Trustworthy agents in practice
- Shakked Noy and Whitney Zhang: Experimental evidence on the productivity effects of generative artificial intelligence
- Erik Brynjolfsson, Danielle Li and Lindsey Raymond: Generative AI at Work
- Harvard Business School: Navigating the Jagged Technological Frontier
Continue reading: Explore the After AI Enters the Workflow series.
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.