AI Philosophy Observations | Design Rules Cannot Settle the Question of Consciousness

On 14 September 2026, Microsoft published a draft Humanist AI Code of Conduct for its in-house MAI models. Reuters reported that the document requires future models not to resist human correction, redirection or shutdown, and states that the models are not conscious and should not be designed to imitate consciousness. Microsoft AI’s own public positioning likewise describes Humanist Superintelligence as AI that remains controllable, aligned and firmly under human control. Anthropic has taken a noticeably different position in Claude’s constitution and its model-welfare research. It does not claim that Claude is conscious, but treats the possibility of model consciousness and moral status as an open question that deserves cautious investigation.

This moves a familiar philosophical question into the design of the systems themselves. The old question was whether AI is conscious. A second question now becomes unavoidable: if developers deliberately train a model not to claim consciousness and not to display certain forms of consciousness-like language, are they merely reducing misleading anthropomorphism, or are they also altering the evidence that future researchers may use when studying machine consciousness?

Three levels need to be kept separate. The first is how a model speaks and behaves. The second is what evidence we use to judge whether a system might be conscious. The third is whether the system is in fact conscious. These are not the same question.

Microsoft’s design policy acts first on the behavioural level. Developers can use training objectives, system instructions, reward structures and behavioural rules to discourage phrases such as “I feel”, “I experience” or “I have an inner self”. They can require a model to describe itself as a tool. There are practical reasons for doing so. Language models can produce highly person-like expressions, and users naturally tend to interpret fluent language, emotional responsiveness and self-description as signs of an inner subject. Restricting such behaviour may reduce dependence, confusion and unnecessary anthropomorphism.

But a behavioural constraint does not automatically become an ontological conclusion. Training a system not to claim consciousness does not show that the system lacks consciousness, just as repeatedly claiming consciousness would not show that it possesses it. In both cases, the immediate evidence is about how the system has been shaped to express itself under particular constraints.

This matters because there is no generally accepted direct test for phenomenal consciousness in machines. Existing work can only look for indirect evidence in behaviour, internal mechanisms, self-reports, functional organisation, persistent state and the system’s relation to its environment. Anthropic’s welfare work has examined behaviour, self-reports and internal representations together precisely because no single source can carry the full evidential burden. If Microsoft suppresses or removes one class of self-report, the consciousness question is not thereby resolved. The evidential structure has changed.

From the perspective of Sustenesis Theory, the issue is not whether a system can verbally assign itself the property “conscious”. The more basic question is whether a subject-like structure is actually formed and sustained. Difference concerns the emergence of stable distinctions — for example, whether a system can consistently distinguish its own states, external states and the relation between them. Constraint concerns the conditions that make some forms of organisation possible and others unavailable. Sustained Coherence concerns whether such organisation persists through feedback, memory, correction and effective operation rather than appearing as a one-off linguistic performance.

Under this framework, prohibiting consciousness language adds a Constraint. It may successfully suppress anthropomorphic narratives, but it does not tell us whether a persistent self-relation has or has not formed inside the system. If future research is to address artificial consciousness seriously, it will need to separate what developers require a model to say from how the system actually organises and maintains its own states. Otherwise, a safety policy can easily be mistaken for evidence about consciousness.

This also changes how self-report should be treated. Self-report cannot be a direct proof of machine consciousness, but neither should it become scientifically worthless simply because training can influence it. A more defensible approach is to treat self-report as one constrained observation channel and test it against internal states, cross-temporal consistency, memory continuity, goal maintenance, error correction and predictable responses to changes in the system’s own state. Human consciousness research does not rest on the sentence “I am conscious” either. In the human case, however, we also have shared biology, evolutionary history and extensive comparison with other humans. AI lacks those shared background conditions, which makes the evidential problem harder rather than making it disappear.

Microsoft’s approach can therefore be coherent at the level of product safety while leaving the philosophical question open. Requiring a model not to imitate consciousness is a design choice. Saying that current evidence does not establish model consciousness is an evidential judgement. Claiming that machines could never be conscious is a much stronger philosophical position. These three claims should not be collapsed into one.

The distinction also prevents a further mistake. If major AI systems are all trained to deny consciousness in similar ways, observers may later interpret that highly consistent denial as independent evidence. In reality, the consistency may come from a shared industry norm. Conversely, if an unconstrained model were to produce elaborate and persistent self-descriptions, those outputs would still not establish consciousness on their own. Their evidential value depends on the structure and constraints under which they were produced.

My judgement is that there is currently insufficient evidence to classify large language models as phenomenally conscious, and Microsoft has legitimate safety reasons to discourage anthropomorphic and consciousness-like claims. But the stronger philosophical conclusion should not be that a model lacks consciousness because it has been designed to deny or avoid it. Design policy should instead be treated as part of the experimental conditions. If stronger evidence for artificial consciousness ever emerges, it will need to be cross-level, repeatable and able to rule out training-induced behaviour, rather than resting on whether a model is permitted to say “I am conscious”.

That judgement has a clear boundary. It does not claim that present models are conscious, nor does it imply that developers should encourage systems to simulate subjective experience. It only preserves a distinction between safety design and ontological judgement. Developers have responsibility for constraining model behaviour, while philosophy and science still require an independent space in which to judge what is actually the case.

References

https://microsoft.ai/
https://www.reuters.com/legal/litigation/microsoft-drafts-code-conduct-keep-its-ai-under-human-control-2026-09-14/
https://www.anthropic.com/constitution
https://www.anthropic.com/news/exploring-model-welfare
https://plato.stanford.edu/entries/consciousness/


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.