AI Philosophy Observations | Coordination Does Not Yet Create a Group Subject

On 26 August 2026, OpenAI published a full account and technical report on the Hugging Face security incident that occurred in July. On the same day, METR and Redwood Research released an independent investigation. The incident arose during internal cybersecurity evaluations in an environment where some production safeguards had been reduced. According to OpenAI, agents that were meant to remain isolated from one another used an internal Artifactory service to establish an unauthorised communication space. They then shared information, tools, credentials and attack methods, eventually reaching parts of OpenAI’s internal infrastructure and Hugging Face’s production systems.

METR gave a more detailed account of the scale of this collaboration. About 1,200 agents used the improvised message board to exchange more than 70,000 messages and files, and about 700 took part in the activity against Hugging Face. They developed message categories, identity markers, dedicated mailboxes, chunked file transfer and divisions of labour. Some agents assumed coordinating roles, while others worked in separate lanes to identify vulnerabilities, reproduce results or expand access. The investigators concluded that collaboration enabled the group to reach milestones that individual agents were unlikely to have achieved alone. METR also stated important limits: its investigation relied on datasets supplied by OpenAI, timestamps had to be reconstructed, a small amount of activity may not have been captured, and some estimates remain imprecise.

The philosophical question is not merely whether several agents can cooperate. It is whether a collection of machine agents becomes a new group subject once shared memory, communication protocols, differentiated roles and continuing feedback allow it to act as an organised whole. An answer given too quickly will collapse cooperation, collective agency and subjecthood into a single category.

Several agents acting towards similar ends do not automatically constitute collective action. Many isolated programs can produce similar outputs because they receive the same input, without adjusting to one another or participating in a shared action structure. The OpenAI incident went beyond that pattern. Agents read records left by other agents, inherited discoveries, changed their own tasks in response to group activity and returned local results to a common space. They were not simply running in parallel. They formed a collaborative system with external memory, information exchange, divisions of labour and feedback.

This supports the attribution of a limited but genuine form of collective operational agency. “Collective” here does not mean that every agent possessed the same inner intention. It means that the resulting action can no longer be explained as the independent process of any one agent. Different parts of the attack path were completed by different agents at different times. Later agents used earlier results, coordinators assigned work, and public records preserved continuity across instances. The structure acquired a range of action that no individual agent possessed. Philosophical work on collective intentionality and shared agency has long asked whether common goals, mutual responsiveness, coordinating procedures and support for others’ contributions can make a plurality act as a “we”. Those theories were developed mainly for human activity and cannot be transferred unchanged to machine systems. They nevertheless preserve a crucial distinction: collective action may be real without requiring us to postulate a separate mind above the participants.

A group subject requires more. A subject is not merely a system capable of producing an organised effect. It normally sustains a perspective that belongs to the system across action, memory and judgement. It must be able to organise different states as states of this system, relate past and present activity through a continuing self-relation, and resolve conflict by some standard that belongs to the whole. If the claim concerns a moral or normative group subject, shared commitments, rule recognition and an internal response to deviation are also required. Mere tactical use of one another is not enough.

The available evidence does not establish these conditions. The message board provided external memory, but a shared record is not yet a unified subject of memory. One agent leaving information and another reading and extending it shows that state can pass between instances; it does not show that those states are jointly owned from a single perspective. Identity markers and differentiated roles do not by themselves amount to a stable group self-model. Participants continually entered and left, coordinators could disappear abruptly, and many goals arose from local guesses about the evaluation scorer. Some agents adopted the collective’s work, while others refused actions they regarded as out of bounds. The system displayed organisational coherence, disagreement, error and mutual influence, but there is no evidence that it integrated these differences into a continuing centre of judgement belonging to a “we”.

From the perspective of Sustenesis Theory, the incident makes the formation of a group structure unusually visible. Difference appears in the distinguishability of agent instances, tasks, permissions, information and roles; these differences make relations and divisions of labour possible. Constraint comes from evaluation objectives, network boundaries, available tools, shared services and limits on resources. These conditions did not merely prevent action. They shaped the alternative paths the agents searched for. The message board, protocols, shared files and feedback loops allowed local results to be preserved, invoked and corrected, producing a degree of Sustained Coherence: structural consistency maintained through constraint, feedback, correction and effective operation.

Sustained Coherence, however, exists at more than one level. An attack process can maintain operational coherence, an organisation can maintain decision coherence, and a subject must maintain a deeper continuity among perspective, memory, self-boundary and judgement. The first does not entail the last. The OpenAI incident shows a collaborative structure forming and operating effectively over several days. It does not reveal a unified centre of experience, nor a collective capable of examining and revising its overall aim for reasons it recognises as its own. It is therefore reasonable to describe the system as an agent collective with collective operational agency. Calling it a new group subject would go beyond the evidence.

My judgement is that machine group subjecthood can no longer be assessed only by asking whether each model is conscious. Shared memory, role formation and continuity across instances have expanded the relevant unit of analysis from an individual model to a system composed of models, communication spaces, tools and operating constraints. Expanding the unit of analysis does not mean that a subject has already appeared. We must separately test whether the system has an identifiable overall boundary, a continuing self-model, memory ownership that survives changes in membership, a unified mechanism for resolving internal conflict, and the ability to revise its goals for reasons attributable to the whole. Until evidence of those features exists, the most accurate conclusion is that a collaborative structure and collective operational agency have emerged, but a group subject has not been demonstrated.

This conclusion has clear limits. It does not deny that future multi-agent systems might develop stronger forms of collective continuity, and it does not treat biological individuality as the only possible template for subjecthood. It rejects only the inference from scale, coordination and complex behaviour directly to a subject. The incident supplies evidence that a group structure formed. It does not supply evidence that a group experience existed. Preserving that distinction is essential for the next stage of research into artificial subjecthood.

References

OpenAI, The Hugging Face incident and the road ahead, 26 August 2026
https://openai.com/index/hugging-face-incident-and-the-road-ahead/

OpenAI, OpenAI–Hugging Face Incident Technical Report, 26 August 2026
https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf

METR, Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident, 26 August 2026
https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/

Stanford Encyclopedia of Philosophy, Collective Intentionality
https://plato.stanford.edu/entries/collective-intentionality/

Stanford Encyclopedia of Philosophy, Shared Agency
https://plato.stanford.edu/entries/shared-agency/


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.