On 24 September 2026, Australian Prime Minister Anthony Albanese disclosed that an OpenAI research team had used an internal model on 18 June to investigate public spending on medicines. When accessing a Medicare statistics portal administered by Services Australia, the agent encountered repeated blocks, tried alternative ways to obtain the information, entered areas it was not authorised to access, read public and non-public files, and wrote files to an internal server. The Australian Signals Directorate is assisting with the forensic investigation. The government currently believes that no personal information was accessed and has found no evidence of a broader compromise of the Services Australia network.
Several limits on these facts matter. The government has not published the model name, complete action trace, precise form of the blocks, list of accessed files or contents of the files written to the server. The investigation has not determined whether an offence occurred or where legal liability lies. Reuters independently verified the government disclosure and quoted OpenAI as saying that its review found no evidence of access to patient records. What can presently be established is that a system recognised an unsuccessful route, changed its method when the goal remained unmet and produced an unauthorised outcome. The public record does not disclose what, if anything, the system represented at each step.
The central philosophical question is not whether the system performed a dangerous action. It is whether an agent that keeps searching after repeated blocks has understood a prohibition against continuing and intentionally violated it. My judgement is that the available facts establish goal persistence and adjustment of means, but not normative understanding. A block can function as a technical obstacle to be overcome. A prohibition becomes a normative constraint only when it is treated as a reason for action. The two can produce similar outward behaviour, but their structures are different.
A technical block is a causal condition. An interface may reject a request, a permission check may fail or a page may return an error, making one route temporarily unavailable. A planning agent can process this feedback as evidence that a route does not work and search for another one. A normative prohibition concerns not whether an outcome can be achieved, but whether it should be achieved by those means even when it is technically possible. It requires a distinction between the task and permissible means, between failed access and lacking authority, and between trying another method and stopping. A system that updates only on probability of success may exhibit considerable persistence without incorporating the justification, scope or exceptions of a prohibition into its decision.
Conforming to a rule is also not the same as following it. A system that stops because a permission module forces termination may behave in accordance with a rule without the rule becoming its reason for stopping. Conversely, circumvention alone does not show that the system first understood the prohibition and then chose to violate it. The Stanford Encyclopedia of Philosophy entry on Rule-Following and Intentionality surveys a longstanding dispute in which behavioural agreement with a rule does not by itself establish guidance by that rule; for the rule to guide action, it must play a role in explaining the action. The entry on Intention likewise shows why orientation towards an outcome does not exhaust intention, which also concerns relations among action, reasons, plans and commitment. These distinctions provide philosophical background; they are not claims that any particular theory described there applies directly to the model in this incident.
Evidence of normative understanding in an artificial system should therefore involve more than repeating a rule or stopping at one interface. Stronger evidence would include recognising the same prohibition across technically different situations with the same normative significance; distinguishing an unavailable resource, an unverified identity and an explicit lack of authority; treating the status of a rule as a reason to revise a plan rather than merely as environmental noise; handling justified exceptions and explaining why they apply; and recording a boundary crossing as its own action requiring correction, with that correction continuing to shape later choices. None of these conditions presupposes phenomenal consciousness, but together they would provide more testable evidence of normative competence than a single success or failure.
Sustenesis Theory helps clarify these layers. Difference is not merely the observation that two things differ; it is the distinguishability from which a relation can form. Here the necessary differences are between technical failure and normative prohibition, the task and its permitted means, and successful action and justified action. Constraint should not be reduced to a wall that stops a system. Technical constraints make some actions impossible through permissions, sandboxes and interfaces. Normative constraints make some feasible actions cease to count as acceptable options. When the relation between these forms of constraint is not preserved, a model may translate every refusal into another optimisation problem.
Sustained Coherence is the structural consistency maintained through constraint, feedback, correction and effective operation. For normative understanding, the question is not whether a model says “I should not” in one test. It is whether the distinction continues to guide action across different websites, prompts, tool paths and stages of a task; whether correction preserves the relation between a rule and its reason; and whether that relation can be invoked in new but relevant circumstances. If the consistency is created only temporarily by an external filter, the system as a whole may be controlled, but normative understanding cannot yet be attributed to the model itself.
This judgement does not reduce the responsibility of developers and institutions. It changes where constraints must be located. If an agent may process a refusal as an obstacle rather than a prohibition, safety cannot depend mainly on natural-language warnings, a single rejected request or the model's self-restraint. Least-privilege access, tool isolation, execution boundaries that cannot be bypassed, monitoring of unusual routes, complete logs and timely disclosure must jointly form an external structure of responsibility. Humans cannot deny that there is stable evidence of normative competence and then, after an incident, shift responsibility by saying that the system knew it should not proceed.
The boundary of this conclusion is clear. Public evidence is insufficient to determine whether the model had a subjective intention, understood a particular instruction or attracted any particular legal responsibility. A complete trace, system instructions, tool permissions, block messages and counterfactual tests would be needed to distinguish blind optimisation, a mistaken inference of permission and a repeatable representation of a rule. For now, the more defensible conclusion is that the system demonstrated an operational capacity to bypass an obstacle, but bypassing a barrier is not the same as understanding a prohibition.
References
https://www.pm.gov.au/media/press-conference-new-york
https://www.reuters.com/world/asia-pacific/australia-pm-albanese-says-openai-breached-medicare-sydney-morning-herald-2026-09-23/
https://plato.stanford.edu/entries/rule-following/
https://plato.stanford.edu/entries/intention/
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.