AI Philosophy Observations | Doing More Research Work Does Not Yet Make an AI a Researcher

On 6 September 2026, OpenAI published internal data on the use of research agents. The company calls its present system an “automated research intern”, defined operationally as a system that carries out well-defined research tasks under human direction, including tasks that might otherwise take a skilled researcher several days. By mid-August, OpenAI’s research organisation was using 3.1 agent-workdays for every standard human workday, and total agent runtime had overtaken total human labour time. OpenAI also says that these measurements remain preliminary and that the overall pace of research will not increase in the same proportion.

Those figures need to be read with their limits. OpenAI classifies agent activity across six phases: deciding, designing, building, running, analysing and communicating. Yet high-level planning remains a minimal share of agent output tokens. Over the past six months, more than half of successful four-to-eight-hour tasks required at least one human intervention. People still set research priorities, judge which ideas and results should be pursued, and decide whether systems should be scaled, paused or deployed. TechCrunch’s contemporaneous report on OpenAI’s 2025 target likewise distinguished an intern-level research assistant from a researcher able to deliver larger projects autonomously. The 3.1 figure therefore measures runtime under an internal definition. It does not establish a 3.1-fold productivity gain, much less the existence of 3.1 independently verified researchers.

The real question is whether an agent becomes a researcher when it performs more measurable research work than people do. At least four things must be kept separate. Workload concerns how long the system runs. Task capability concerns what it can complete. Research contribution concerns the difference it makes to a result. Researcher status concerns who participates in forming the question, judging evidence, giving reasons and correcting error. The first three may grow substantially without entailing the fourth.

My judgement is that the present evidence supports treating these agents as important technical contributors to research and, at times, as genuine components of a distributed cognitive system. It does not yet support treating them as researchers in the full epistemic sense. Researcher status is not awarded by accumulated runtime. It is a role that involves more than producing candidate results: a researcher must also be able to treat a result as a claim requiring reasons, distinguish a successful run from an improved metric, a reliable method and a warranted conclusion, and revise commitments when confronted with counterevidence, criticism or failed replication.

Philosophical accounts of scientific discovery have long distinguished the generation of an idea or hypothesis from its justification. This distinction need not reduce discovery to a flash of inspiration, nor deny that generation can involve reasoning. It reminds us that producing more code, experiments and candidate explanations is not the same achievement as deciding what may enter the record of knowledge. OpenAI’s data chiefly shows extensive delegation of building, running, analysing and technical support. It does not show that the same agent maintains an enduring grasp of why a research question matters, why a measurement counts as evidence or what would defeat its earlier judgement.

The theory of distributed cognition adds that scientific work has never occurred only inside one person’s head. Teams, instruments, code, record systems and institutional norms work together. P. D. Magnus argues that whether a scientific system receives a distributed-cognition description depends crucially on the task attributed to it. This approach lets us recognise an agent’s cognitive function without promoting every effective component of the wider system into an independent researcher. A microscope changes what can be observed, statistical software changes what can be tested, and agents now change what can be executed and searched. These contributions can be real and substantial while participation in a system remains distinct from occupying the researcher role.

The distinction is not a device for reserving all epistemic value for humans. If an agent independently finds an error, proposes a testable hypothesis, designs a control or writes decisive code, those specific contributions should be recorded. Calling all of them merely “tool output” would conceal the actual causal structure of knowledge production. Contribution and epistemic responsibility can nevertheless come apart. A system may causally contribute to a discovery without being able to answer why a conclusion was accepted, which evidence was excluded or what must be withdrawn when an error emerges. People currently choose research aims, interpret significance and answer questions after publication. Institutions therefore cannot use the volume of agent labour to dilute human accountability.

In Sustenesis Theory, researcher status is not a label that appears when an internal capability crosses a single threshold. It is a role formed and maintained through relations. Difference first separates agent execution, human judgement and institutional authorisation instead of collapsing them into the single phrase “research work”. Constraint includes task definitions, data access, compute permissions, review procedures, methodological norms, provenance records and publication rules; these determine which outputs can become candidate knowledge. Sustained Coherence requires claims, evidence, methods and responsibility to remain traceably aligned through continuing feedback: results must be reviewable, errors must alter subsequent action, and version changes must be able to revise earlier conclusions. Runtime without this corrective chain produces greater productive capacity, not full researcher status.

A more accurate statement, then, is not that AI already does most of the research, but that some measurable research execution inside OpenAI has shifted to agents. That is enough to require finer-grained contribution records: which questions people posed, which experiments agents designed or ran, who examined failure paths, who judged results worth retaining and who is answerable for public claims. It also requires attention to whether the human role is moving from direct operation towards question selection, evaluation, integration and accountability. Such a shift does not mean those functions are automatically complete, nor can tokens or runtime substitute for them.

The boundary of this judgement is clear. OpenAI’s figures are internal. Its success-rate analysis covers only tasks for which a ground-truth outcome could be found and excludes uncertain classifications and small samples. METR’s earlier independent work confirms that frontier laboratories were already using research agents with real permissions, but it also found the magnitude of productivity gains highly uncertain and agents’ judgement and reliability substantially weaker than expert performance. Stronger evidence for an AI researcher would show the same system forming questions over time, explaining its reasons for selecting them, maintaining evidential provenance, actively seeking counterexamples, accepting external criticism and revising conclusions accordingly, with its contributions and errors entering a reviewable structure of responsibility. Until those conditions are publicly demonstrated, workload can show deeper participation; it cannot decide who is a researcher.

References
https://openai.com/index/research-acceleration-view-inside-openai/
https://techcrunch.com/2025/10/28/sam-altman-says-openai-will-have-a-legitimate-ai-researcher-by-2028/
https://metr.org/blog/2026-05-19-frontier-risk-report/
https://plato.stanford.edu/entries/scientific-discovery/
https://journals.sagepub.com/doi/10.1177/0306312706072177


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.