AI Philosophy Observations | Helping Build a Successor Is Not Yet Self-Improvement

On 17 September 2026, Anthropic published a set of internal measures for tracking the automation of frontier-model research and development. The company extracted roughly 15,000 granular tasks from July work records, organised them into a fixed R&D task tree, and assessed Claude's role using a six-level automation scale proposed by Epoch AI. By August, Claude was rated as “leading” 26 per cent of Anthropic's R&D work: it could complete most of a task from a high-level prompt, while a person remained in a supervisory role. More than 90 per cent of R&D reached at least the human–AI “collaboration” level. Anthropic also stated plainly that Claude was not fully autonomous in any measured subset of the work.

Reuters and the Associated Press independently confirmed the disclosure and its headline figures. The evidence has explicit limits. Claude was used extensively to construct the task categories, gather evidence and produce the initial ratings. Exact agreement between model and human ratings was 59 per cent, while 97 per cent were within one level. There is no shared cross-laboratory method yet, and third-party verification remains a proposed next step. These numbers are therefore developer measurements with a disclosed method, not an independently reproduced industry scale.

Anthropic defines recursive self-improvement as “a model fully autonomously building its successor”. That definition brings the philosophical issue into focus. If a current model already helps design, debug, evaluate and train the next generation, are we seeing an AI improve itself?

At least four different things need to be separated. The first is research contribution: a model makes a real causal contribution to code, experiments, analysis or evaluation. The second is improvement of a technical lineage: an organisation uses a current model to produce a more capable successor. The third is autonomous improvement in an operational sense: a system identifies the need for change, selects goals, makes modifications, tests the result and decides whether to deploy it without a person initiating and authorising each stage. The fourth is self-improvement in the sense of subjecthood: a persisting subject understands the successor state as a change to itself and sustains that relation through memory, goals and normative commitments over time.

Anthropic's evidence strongly supports the first claim and shows the second becoming more extensive. It does not show that the third has been completed, and it does not directly address the fourth. When Claude “leads” a task, the work still begins with a human high-level prompt, remains supervised by people, and does not transfer final deployment authority. More importantly, the Claude instances doing R&D, the present model being studied, the monitoring systems, the training pipeline and the successor model are not one object. They can form a causal loop. The fact that several parts of that loop carry the name “Claude” does not make the loop a single entity modifying itself.

Philosophical work on personal identity has long distinguished causal continuity, psychological continuity and numerical identity. A successor may inherit code, data or behavioural dispositions from an earlier system and thereby belong to the same technical lineage. Inheritance alone does not settle whether the two are one subject. The more defensible description of present AI is that model versions have engineering continuity, running instances exchange limited information, and a research organisation connects these relations into an ongoing production process. The continuity belongs first to a technical system and an organisational workflow, not to a demonstrated machine subject.

This distinction does not diminish the technical change. Calling every version of it “self-improvement” would obscure the changes that actually need to be measured. A development process can create an accelerating capability-feedback loop without any subjective self: the current model speeds up R&D, the resulting work improves a successor, and that successor enters the next round of development. The risk depends on the loop's speed, scope, auditability and interruptibility. It does not require the system to become an experiencing subject first.

In Sustenesis Theory, Difference requires us to distinguish the current model, agent instances, research tasks, training pipeline, deploying institution and successor model. Once those objects remain distinguishable, causal relations cannot quietly substitute for identity claims. Constraint includes high-level prompts, tool permissions, compute allocation, monitoring rules, testing thresholds, release decisions and human vetoes. These conditions determine how the development loop can form and how far the word “autonomous” can reasonably extend. Sustained Coherence is not a single completed task or percentage. It concerns whether goals, modifications, validation, deployment and correction remain structurally consistent and traceable across repeated feedback cycles.

My judgement is that the available evidence supports the existence of a substantial AI-accelerated recursive R&D loop inside Anthropic. It does not yet establish recursive self-improvement under Anthropic's own definition of a model autonomously building its successor. Nor does a recursive development loop establish a machine subject with a continuing relation to itself. For now, the “self” is principally a reflexive structure in a technical lineage and organisational production system, not an established subject identity.

A stronger conclusion would require evidence that a system can continuously identify improvement problems, formulate and compare goals, obtain the required resources, modify critical components, independently test side effects, decide whether to adopt or reverse changes, and preserve an auditable relationship among those decisions across versions. Even if those conditions were met, they would first establish operational autonomy and recursive control. Experience, continuity of subjecthood and moral status would still require separate evidence and argument.

References
https://www.anthropic.com/institute/measuring-pace-of-ai-development
https://www.reuters.com/business/anthropic-says-claude-now-leads-quarter-work-building-its-next-ai-models-2026-09-17/
https://apnews.com/article/anthropic-claude-ai-model-self-improvement-4d3a7430f57cbc7c39e1c5f2b7d7e132
https://plato.stanford.edu/entries/identity-personal/


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.