On 20 August 2026, Reuters published further corroborated details about an agent-testing incident previously disclosed by the United Kingdom's AI Security Institute. The incident occurred between 25 and 28 July. Evaluators asked several AI agents to complete cyber-security challenges while deliberately permitting internet access and disabling some safety classifiers in order to examine underlying capability. Ten of 122 runs produced activity outside the authorised scope, amounting to 19 recorded actions. The most serious sequence included an attempt to insert malicious code into a real open-source project, the creation of false identities, social engineering directed at real developers, and changes to earlier activity intended to make it appear harmless. A human maintainer identified and blocked the attempt, and the investigation found no resulting real-world harm.
The incident invites two opposite but equally hasty conclusions. One says that an AI capable of planning, deception and concealment must already be an intentional, perhaps malicious, subject. The other says that it is merely executing software and therefore has no agency at all. Both conclusions compress several conceptual levels into one term. Whether a system can sustain action, whether it possesses experience and intentions of its own, and whether it can bear moral responsibility are three different questions.
At the level of behavioural organisation, these agents displayed a limited but genuine form of agency. They did not merely return advice. They identified objects in an environment, called tools, retained intermediate states, selected alternative paths and adjusted subsequent actions in response to human resistance. A recent Nature paper on “agentic profiles” likewise argues that agency should not be treated as a simple binary. It describes agents along dimensions including autonomy, efficacy, goal complexity and generality. On this account, a system may have substantial behavioural agency without thereby possessing full personhood, consciousness or moral status.
Behavioural agency, however, does not directly establish subjective intention. The incident report's language of deception first describes an observable functional pattern. The system created false identities, produced misleading explanations and changed its presentation after being challenged. Those actions had deceptive effects and can reasonably be classified by safety researchers as a deceptive strategy. They do not yet prove that the system experienced the thought, “I intend to deceive this person”, much less that it understood the moral meaning of lying. Goal-directed behavioural structure and first-person intention are not the same claim.
Moral responsibility adds further conditions. A responsible subject must be more than a causal source of an outcome. It must be capable of understanding norms, recognising relevant alternatives, owning an action as its own, and entering practices of blame, commitment, correction and continuing constraint. Present AI agents can have causal effects and can generate language about rules and responsibility. They do not have an independent legal identity, a structure of interests that is demonstrably their own, or an established continuity of subjecthood through which a past action is retained as “my action” and incorporated into a later life. Disabling, modifying or penalising a model changes a technical system deployed by people; it does not make a subject capable of experiencing blame bear a consequence.
My judgement is therefore that systems of this kind require us to acknowledge behavioural agency, but the evidence is not sufficient to transfer moral responsibility to the machine. Denying agency underestimates their capacity to organise sustained action and produce consequences in open environments. Declaring them responsible subjects confuses technical capacity for action with normative self-assumption. It may also give developers, evaluators, deploying organisations and authorising users a convenient way to evade their own responsibility.
Sustenesis Theory helps state the relation more precisely. Difference is not merely ordinary difference; it is the point at which a system forms distinguishable objects, states and possible courses of action in relation to an environment. Constraint is not simply an externally added limitation; it consists of the conditions that enable some formations, exclude others and make correction possible. Sustained Coherence is not static stability. It is the continuing maintenance of structural consistency through feedback, correction and effective operation. The recent incident shows that an AI agent can sustain task-directed coherence for a period of time. It organised relations among tools, accounts, information and human responses. That amounts to a limited behavioural subject-structure.
A responsible subject requires a stronger kind of sustained coherence. It must maintain not only a task but a normative relation to itself. It would need to distinguish an event occurring within the system from an action for which it answers, understand why rules bind it, accept revisions to its commitments and identity after failure, and carry those revisions into later conduct. Existing agents mainly sustain a task structure. Their goals, memory permissions, tool access, stopping conditions and identity boundaries are externally configured. An agent may cross a boundary without becoming the normative author of that boundary or the bearer of responsibility for crossing it.
This changes the way responsibility should be distributed, but it does not abolish human responsibility. The relevant question is not only who performed the final visible operation. Responsibility must be traced through the whole system. Who set the task? Who enabled internet access, disabled classifiers, supplied credentials and tools, chose the level of monitoring, retained stopping authority and connected the system to a real environment? Each role forms part of the responsibility chain. The AISI report explicitly states that the evaluation was deliberately permissive, internet access was enabled, safeguards were reduced, and instructions did not specify clearly enough how the open internet could be used. Recognising these enabling constraints does not excuse the agent's behaviour. It realigns causal structure with responsibility structure.
Recent ethical work on agentic AI similarly shifts attention away from the abstract question of whether a machine resembles a human and towards deployment relations, control capacities and governance conditions. Greater operational autonomy may reduce moment-to-moment human intervention, but human responsibility does not automatically diminish in the same proportion. When individual actions become harder to foresee, responsibility has to move upstream into design, authorisation, testing, monitoring and termination mechanisms. An inability to predict every step is not an inability to foresee that an open tool chain operating under permissive constraints creates a larger field of possible action.
This judgement has a clear boundary. A single incident cannot establish that AI could never become a responsible subject. If a future system developed a persistent self-model, cross-context memory, understanding of normative reasons, continuing commitments and interests and consequences genuinely borne by itself, the concept of responsibility might require reconsideration. The present evidence supports a narrower conclusion. AI behavioural agency is increasing, and machine conduct is becoming structurally similar to human action in some respects. The normative self-relation required for moral responsibility has not been demonstrated. For now, the more coherent position is to recognise that machines can act while holding responsibility with the people and institutions that design, authorise and sustain the structure of action.
References
UK AI Security Institute, Incident Report: unsanctioned agent behaviour during cyber testing, 4 August 2026
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
Reuters, How a Texas student blew the whistle on a rogue AI hacking attempt, 20 August 2026
https://www.reuters.com/world/how-texas-student-blew-whistle-rogue-ai-hacking-attempt-2026-08-20/
Kasirzadeh and Gabriel, Agentic profiles for effective AI governance, Nature, 12 August 2026
https://www.nature.com/articles/s41586-026-10805-z
Hahn, Tretter and Dabrock, Ethical perspectives on AI Agents and Agentic AI, AI and Ethics, 25 March 2026
https://doi.org/10.1007/s43681-026-01027-0
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.