Research and version note This is a version 0.1 research draft in Who Maintains the World? It concerns continuing maintenance after AI deployment, not a safety certification for any particular product. Government policy and technical sources were reviewed to August 2026. Each organisation still needs an assessment based on its use, law, contracts and risk.
Who Maintains the World? · Article 11
Quick read
AI often enters an organisation with a promise to reduce routine work: sorting requests, finding anomalies, drafting responses, summarising records or making recommendations. The business case can make deployment look like an endpoint, as though a model that passed testing will continue to produce the same value. In practice, deployment begins a new maintenance arrangement.
Data, user behaviour, business definitions and external conditions change. A supplier updates a foundation model, retrieval material expires, interfaces are modified, and staff gradually alter their reliance on the output. Real-world effects can shift even if performance on a fixed test set does not. NIST’s 2026 report on deployed AI monitoring separates functionality, operations, human factors, security, compliance and large-scale impacts. It also reports that validated methods and common terminology remain nascent.
“A human reviews the result” is not an automatic solution. Reviewers need time, source material, authority to reject the output and continued practice. When nearly all recommendations look plausible, automation bias can develop. When routine judgment has been delegated for long enough, people may lose the experience needed to recognise a rare failure. A nominal human in the loop can leave liability with the person while moving effective control elsewhere.
AI maintenance is not performed only by model engineers. Business workers define correct service, data workers maintain inputs, security teams respond to attacks, frontline staff notice contextual errors, and appeals reveal failures that internal tests did not see. Procurement decides whether an organisation can obtain logs, receive update notices, roll back a version or continue when a supplier exits.
The maintained object must be the service and its justified purpose, not merely model availability. A model can remain online and statistically stable while staff misuse its output or affected people lose a meaningful route to correction. Technical telemetry, service outcomes and user experience need to be connected to thresholds that trigger investigation, restriction or suspension.
This draft argues that every AI system carrying an ongoing task needs an identifiable maintenance institution: accountable owners, defined monitoring, change control, incident routes, human intervention, appeal and retirement conditions. Those costs belong in the original business case. AI can reduce particular tasks, but it does not remove maintenance; it relocates maintenance into more distributed positions that need evidence and authority.
Go-live is the start of exposure to reality
An AI project is most visible before launch. A team gathers data, assesses a model, performs privacy and security reviews, and seeks a go-live decision. Budget and executive attention often peak around that date.
Only after launch does the AI become part of real work. Inputs are messier than test examples. Staff may copy an answer into decisions that the designer did not anticipate. External policy changes the meaning of a correct outcome, while adversaries begin probing the boundary. No single offline accuracy result contains those relationships.
Conventional software also needs patches, monitoring and support. AI adds a further form of instability: both the distribution of inputs and the relation being estimated may change. Generative systems can also vary across similar or repeated prompts. Pre-deployment evaluation remains essential, but it establishes the performance of a particular version under specified conditions. It cannot certify all future use.
This temporal asymmetry matters for governance. Approval is a visible event with a document and a decision-maker. Months of observing weak signals are diffuse. If resources move immediately to the next project, the organisation creates systems faster than it creates the capacity to learn whether they remain acceptable.
Model, system and use context are different maintained objects
When an organisation says it uses “an AI”, it often combines three objects. The model produces a classification or text. The system also includes prompts, retrieval sources, rules, interfaces, identity controls and logs. The use context includes how staff interpret an output, how a decision is made and whether the affected person can correct it.
The system can deteriorate without a model change. An expired knowledge link makes retrieval misleading. A permissions error discloses information. Interface language turns a recommendation into what appears to be a command. Conversely, a supplier can update the model while the customer’s application code remains identical; observed behaviour may still change.
A checklist that monitors model accuracy alone therefore misses the service. A conversation summariser may continue to represent sentences faithfully, yet complaint handling changes because employees use it instead of the full record. Model metrics can look stable while procedural rights or the quality of attention decline.
Responsibility should follow this layered structure. A model provider can describe changes to the model but may not see how an agency places the output in a statutory process. A deploying organisation sees the process but may lack access to weights or training details. Neither fact justifies a gap. It requires agreed information and action across the boundary.
Drift is only one family of maintenance problems
Data drift occurs when the distribution of inputs differs from training or validation. Concept drift concerns a change in the relationship between inputs and the target. Fraud strategies, language, economic conditions and administrative definitions all evolve. A model can remain internally consistent with an old relationship while becoming systematically wrong in the new setting.
Not every deployment failure should be labelled drift. Labels may have been unreliable from the start. The system’s decisions can alter its own feedback data. Staff can extend a model to a population that was never evaluated, or an upstream component can change. Generative AI adds retrieval contamination, prompt injection, unsupported statements and provenance problems.
The 2026 NIST report NIST AI 800-4 organises post-deployment monitoring into six categories: functionality, operations, human factors, security, compliance and large-scale impact. Its contribution is not a finished monitoring standard. The report maps gaps, barriers and open questions in a field where practice is still fragmented.
That uncertainty is not a reason to wait for perfect guidance. It is a reason to limit claims and match deployment scale to available monitoring. A high-impact use should not expand faster than the institution’s capacity to detect and reverse harm.
Monitoring has to begin from the use and its harm pathways
An internal meeting summariser and a system influencing benefit eligibility or clinical priority should not have the same monitoring regime. Reversibility, volume, affected populations, the seriousness of error and opportunities for human checking change the requirement.
Technical monitoring can cover latency, outages, security incidents, input shifts and performance on maintained test material. Service monitoring asks different questions. Do staff misread the output? Who cannot complete the process? Are appeals concentrated in a group? Has the tool produced extra hidden work? Average accuracy can conceal an infrequent but severe failure.
Every measure needs a baseline and an action threshold. Once a change is detected, who investigates? At what point is use restricted, a version rolled back or the system suspended? A dashboard that sends warnings to nobody with authority provides visibility without maintenance.
Monitoring can also create new risks. Detailed logs may contain personal or confidential data. Teams need retention limits, access controls and a reason for collection. “Observe everything” is not a responsible substitute for deciding what evidence the use requires.
Human involvement can be organisational theatre
“A human makes the final decision” is widely used as a reassurance. For it to be true, the reviewer must know where AI participated, have access to relevant source evidence, understand limitations and possess authority to change the result. Workload must permit attention.
If a system produces thousands of recommendations and a small team has seconds to confirm each one, the human gesture resembles an approval routine. Interfaces may foreground the model output, while staff must write an additional justification whenever they disagree. The organisation then rewards acceptance while retaining a person’s signature for accountability.
Automation changes skill as well. When people no longer handle ordinary cases, they see less of normal variation. If the system sends only the most unusual cases to people, reviewers face a difficult sample without routine practice. A maintenance plan may need independent sample review, training and task rotation. It cannot assume unused expertise remains intact.
Affected users provide another source of evidence, but they should not become unpaid testers. A reporting path must be findable, and an appeal should not require a person to diagnose the model. Once a systematic fault is found, the organisation should search related past cases rather than correcting only the person persistent enough to complain.
Meaningful human control can also occur at stages other than every individual output. Design limits, approval of purpose, sampling, incident response and a credible stop decision may be more effective for some low-latency systems. What matters is that intervention can alter reality, not that a diagram contains a human-shaped box.
Supplier dependence changes the maintenance boundary
Many organisations access AI through an API or managed cloud service. The provider maintains weights, filters and service infrastructure, while the deployer controls only part of the application. Responsibility is distributed; it is not dissolved.
Procurement should address version notice, available logs, data handling, incident cooperation, evaluation windows, service termination and migration. If an organisation cannot identify when a material change occurred, it cannot explain why outputs differ from an earlier period. If there is no rollback or alternative process, a formal power to suspend may be unusable in practice.
Commercial confidentiality can protect intellectual property, but it cannot replace all operational evidence. A high-impact deployer needs enough information to judge fitness for purpose, understand material changes and allocate action between supplier and user. A contract does not transfer a public authority’s statutory responsibility to a model provider.
Concentration matters too. Several apparently separate services may depend on the same foundation model or cloud component. Monitoring one application at a time will miss common failure. An inventory needs dependency information, not only a list of product names.
Australian government policy retains human accountability
The Australian Government Policy for the responsible use of AI in government, version 2.0 took effect on 15 December 2025 for non-corporate Commonwealth entities, subject to stated exclusions. It includes accountable officials, transparency statements, a strategic approach, internal use-case registers, staff training and use-case impact assessment.
Its preparedness and operations material says APS officers need to explain, justify and take ownership of advice and decisions when using AI. The impact-assessment section identifies ongoing monitoring and evaluation as a principle. These provisions matter because they do not assign responsibility to “the AI” as an abstract actor.
A framework still has to become local capability. A register does not find an incident by itself, and one accountable official cannot personally observe every use. Technical staff, business owners, frontline feedback, resources and a route to stop the system must connect. Otherwise formal accountability sits on an organisation chart while operating knowledge remains elsewhere.
The policy also illustrates a necessary distinction between general governance and legal authority. Compliance with an AI policy does not establish that a particular administrative, employment or health use is lawful. Existing duties, review rights, privacy rules and sector standards continue to apply.
AI has many maintainers, but it needs a path from signal to action
Model developers maintain training and evaluation. Platform teams support infrastructure. Application workers manage prompts, retrieval and interfaces. Domain experts determine whether outputs still correspond to real definitions; security and privacy staff handle attacks and data use; records workers preserve versions and decision evidence. Frontline staff notice exceptions, while affected people encounter circumstances internal tests missed.
Listing those roles is not enough. Maintenance requires a route from observation to intervention. Who aggregates complaints? How does a technical anomaly become a business-impact investigation? Who may suspend use, who approves resumption, and who contacts people affected before the fault was found?
Some roles contain a structural conflict. A product owner may be responsible for both rapid adoption and risk control. A supplier may measure success by usage while the public organisation should measure service and rights. Independent review is not necessary for every trivial tool, but a high-impact system should not be judged safe only by the group rewarded for its expansion.
Maintenance workers also need protection when reporting failure. If finding a problem is treated as obstruction to innovation, the organisation suppresses the very signals on which safe scaling depends. A credible governance culture treats defect discovery as information, while still distinguishing evidence from general resistance to change.
Records give AI accountability a memory
Rapid change makes version records especially important. An output may need to be connected to the model or service version, application configuration, relevant knowledge sources and rules in force at the time. Privacy and security may limit retention of prompts and personal content, but an environment that cannot be reconstructed at all cannot support meaningful review.
Change control should record the reason, validation, approval and rollback plan. The UK Government AI Playbook calls for continuing operational monitoring, managed releases, documented changes and the ability to revert where required. It also distinguishes model performance from service measures—an accurate component does not prove that users can complete their task correctly.
Institutions should retain evidence of alerts that were dismissed and cases where people overrode the system, within lawful limits. If only the final decision survives, the organisation cannot see how often staff corrected the model or whether challenge declined over time. Maintenance records support learning before an incident as well as accountability afterwards.
Version history also prevents a convenient fiction: that the system reviewed after a controversy is necessarily the system that produced the original decisions. Without a reliable chronology, later improvements can obscure earlier exposure.
Retirement belongs to the maintenance life cycle
An AI system may need to end because performance declines, law changes, a supplier leaves or the use is no longer legitimate. Without a retirement plan, organisations add workarounds around an old model until nobody understands the dependencies. Informal uses may continue after the formal project closes.
Retirement determines whether data are retained or deleted, how records remain readable, what replaces downstream processes and whether past decisions require review. External users may need notice and another service channel. Turning off a model does not undo the institutional consequences it has already produced.
These costs belong in the original business case. A comparison of licence fees with labour hours saved omits monitoring, appeal, training, supplier management and migration. Automation can look inexpensive because its maintenance labour has been moved into other job descriptions.
The possibility of retirement also disciplines adoption. If an organisation cannot describe how it would operate safely without the AI, dependence may arrive before evidence that the use deserves permanence.
A Sustenesis reading of delegated judgment
AI makes delegated knowing particularly visible. No single person holds the data, model, business process, interface and every case in view, yet outputs from those distributed components enter action. The delegation of a knowledge function does not carry understanding and responsibility away with it.
From an applied Sustenesis perspective, identity cannot be inferred from a product name. After changes in supplier, weights, data, policy and work practice, “the same AI” may retain only a procurement label. Maintenance must decide which changes remain within the authorised use and which have formed a different institutional arrangement.
The human remainder should not be reduced to signing for a machine. People and institutions still commit to the purpose, acceptable consequence, treatment of exceptions, appeal and termination. A system may help discover a pattern, but it cannot provide the organisation’s reason for continuing to let that pattern shape action.
This is why explainability alone is insufficient. An explanation of a model output can clarify one component. It does not answer who authorised the purpose, why the threshold is legitimate or how harm will be remedied. Maintenance joins technical evidence to those institutional commitments.
Current judgment: sustainable AI requires a maintenance institution
AI deployment should not end with a one-time approval. Every continuing system needs a use owner, technical maintainers and a decision authority answerable to affected people; they need not be the same person. Monitoring must extend across the model, service, human interaction and distribution of effects, and it must connect to investigation, rollback, suspension and remedy.
Human review counts as control only when reviewers have information, time, skill and the power to change the result. Feedback and appeals should enter systemic review rather than being handled solely as isolated customer-service cases. Supplier arrangements must provide material-change visibility, incident cooperation and a feasible exit.
The version 0.1 conclusion is that AI tends to automate work that can be specified, while the labour of maintaining AI spreads across technical, legal, organisational and interpersonal boundaries. A business case that counts only saved tasks misses the work that keeps automation reliable and legitimate. AI does not remove maintenance. It forces a renewed decision about who can see failure, who may intervene, and who remains able to justify why the system exists.
Primary sources and further reading
- NIST, Challenges to the Monitoring of Deployed AI Systems (NIST AI 800-4), March 2026.
- NIST, AI Risk Management Framework: Core.
- Australian Government, Policy for the responsible use of AI in government, version 2.0, effective 15 December 2025.
- Australian Government, Preparedness and operations under the AI policy.
- UK Government, Artificial Intelligence Playbook for the UK Government, 2025.
Series navigation: Who Maintains the World? series overview
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.