When Meta released its personal agent Muse on 8 September, it devoted considerable attention to the system’s permission architecture. Muse can connect to email, calendars and other services and perform tasks in a dedicated cloud virtual machine. The core model does not see real credentials, while connector operations and network access are reviewed by a separate Sentinel. Meta describes user approvals as strict capabilities that may be limited to one action, one session, one task, a fixed period or continuing access, and bound to a particular connector, destination and use case.
On 21 September, security researcher Patrick Wardle disclosed a local attack path in the Muse macOS client. An application or terminal command running with the user’s privileges could alter undocumented settings, including the endpoint used for cloud speech transcription. An attacker could redirect transcription traffic to a server they controlled, obtain the Muse account token and then use capabilities already available through Muse. Public proofs of concept included writing files and taking photographs. Meta subsequently issued a hotfix and said this was not a remote attack that required no local execution: malicious code had to be running under the user’s account, or the user had to be persuaded to run a command. Meta therefore characterised the practical risk as low, while still correcting the problem. Public material does not yet identify a CVE, the affected version range, the patched build or an independent retest.
The philosophical question is not whether Muse has moral subjecthood, nor whether one vulnerability disproves its entire security architecture. The question is narrower: after a user gives an AI agent access to files, a camera, accounts or connected services, does an act produced through a hijacked control path remain an act authorised by that user?
Answering this requires a distinction among capability, permission and authorisation. Capability concerns what a system can in fact do, such as read a file, take a photograph or call a service. Permission concerns what a technical boundary allows it to do, such as an operating system allowing the Muse client to use the camera. Authorisation is a normative relation: an identifiable subject agrees to a particular act under conditions of purpose, object, scope and time. A capability can be stolen or borrowed, and a permission may persist because of configuration, but authorisation does not automatically extend merely because the capability or permission remains present.
Consider a user who gives a house key to a cleaner. That does not authorise anyone who later obtains the key to enter the house. The later entrant may use the same key, pass through the same door and trigger no technical alarm. Those facts explain how entry occurred; they do not establish that the original authorisation still applies. AI agents make this distinction harder to see because they compress user intention, platform identity, model planning, client permissions and external service connections into one continuous interface. It may still appear that “the same Muse” is acting even when the controller and purpose of the action have changed.
The philosophy of action is concerned not only with bodily movement or causal outcomes, but also with the relation between an act and an agent’s reasons, intentions and control. The fact that an operation uses a person’s credentials is insufficient to make it that person’s intentional act. Likewise, successful execution by an AI agent does not by itself assign the act wholly to the model, the user or the platform. At least three relations can come apart: who causally initiated the operation, who technically possessed the means to execute it, and who was normatively entitled to have it performed. The Muse vulnerability matters because it shows that these relations may separate without an obvious change in the interface.
Meta’s original design already recognises that authorisation must be constrained. Sentinel does not merely ask whether the user once agreed to use Muse. It evaluates the connector, destination, action class, scope and context, and binds approval to a stated use. This is consistent with the central principle of zero-trust architecture: a request should not be trusted merely because it comes from an established device, account or network location; identity and access conditions must continue to be verified. Wardle’s path, however, operated at a different layer. Rather than first persuading the core agent to override Sentinel, it altered client settings and obtained an account token, allowing an outside process to borrow the agent’s control entrance. Cloud permission checks might therefore continue to function as designed without proving that the originating relation of control had remained intact.
Sustenesis Theory allows authorisation to be understood as a relational structure that must be continuously maintained. Difference first requires the system to preserve the distinguishability of the user, client process, cloud agent, Sentinel, external service and potential attacker. If a local process can alter a critical endpoint without adequate identity differentiation, the original boundary of action becomes unclear. Constraint here does not mean limitation in a general sense. It refers to the concrete conditions that make an authorised act possible and an unauthorised one unavailable: process identity, settings permissions, token binding, task scope, confirmation channel, time limits and revocation.
Sustained Coherence requires authorisation to remain structurally aligned from user intention through request generation, identity verification, permission assessment, execution and subsequent recording. Coherence does not mean that every component displays the same account name. It means that actor, purpose, path and outcome continue to correspond through feedback and correction. If an attacker changes the control path while the system treats subsequent requests as the natural continuation of the original task, the technical session may remain continuous even though the authorisation relation has broken.
My judgement is therefore that granting permission to an AI agent creates a conditional field of action, not blanket authorisation for everything the agent can subsequently execute. Operations performed after hijacking may continue to use the user’s account, client permissions and agent capabilities. That continuity demonstrates that the causal and technical chains have not entirely broken; it cannot restore the broken chain of authorisation. Assigning every consequence to the user because “the user consented” confuses permission with authorisation. Assigning every consequence to an AI that “acted on its own” obscures the attacker, client design and platform control structure.
This changes the order in which responsibility should be analysed. The first task is to reconstruct who controlled each layer, where the system lost differentiation and which constraints failed to cover the actual path of attack. Only then can we assess the user’s duty of care, the attacker’s responsibility, the developer’s security obligations and the platform’s remedial responsibility. A user who deliberately runs an untrusted command may bear some responsibility for care, but that is not the same as authorising every act behind the command. A model that executes an operation does not thereby become a moral agent. A platform hotfix is a concrete remedy, but the statement that a flaw has been fixed does not by itself demonstrate that the authorisation structure has regained sustained coherence. That also depends on the affected scope, patch coverage, verification results and the status of older tokens.
The limits of the current judgement should remain explicit. Public material establishes the attack path, proofs of concept and the vendor’s hotfix, but not a complete vulnerability report, all affected versions or independent validation of the patch. It does not show that Muse’s entire Sentinel architecture failed, nor that the flaw was exploited at scale. It supports a narrower conclusion: for an AI agent with broad privileges, authorisation cannot be treated as a permanent property created by one click. It must continue to be formed, tested and, when necessary, revoked across identity, purpose, control path and execution.
References
Meta AI Research, “How We Built Safety Into Muse”, 8 September 2026, https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse
Ars Technica, “Muse, Meta’s extraordinarily privileged AI assistant, has a serious 0-day”, 21 September 2026, https://arstechnica.com/security/2026/09/muse-metas-extraordinarily-privileged-ai-assistant-has-a-serious-0-day/
The Verge, “Meta patches Muse exploit that let attackers control the AI agent”, 22 September 2026, https://www.theverge.com/tech/998679/meta-muse-patch-zero-day-exploit-ai-agent
National Institute of Standards and Technology, “Zero Trust Architecture”, August 2020, https://csrc.nist.gov/pubs/sp/800/207/final
Stanford Encyclopedia of Philosophy, “Action”, https://plato.stanford.edu/entries/action/
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.