AI Technology Observations | 7 September 2026, 08:00

This observation covers the period from 6 September to 8:19 am on 7 September 2026 in Australia/Sydney time and includes one item.

OpenAI publishes internal data on research-agent use and human intervention

On 6 September, OpenAI published “Research acceleration: The view inside OpenAI”, its first quantified account of the intensity of internal research-agent use, task duration and the effects of safety restrictions. The company reported that, by mid-August, its research organisation was using 3.1 agent-workdays for every standard eight-hour human workday. At API prices, the median researcher ranked by agent use consumed more than US$600 of inference per day, while the 90th-percentile researcher consumed more than US$7,000 per day. These figures measure runtime and inference use. They do not establish a 3.1-fold increase in research output, and they do not measure deployment conditions outside OpenAI.

The change from the previously public record is that OpenAI has extended its description of internal adoption into an August quantitative snapshot with methodological notes and has declared that it meets its own “automated research intern” goal. The company limits that label to a system that performs well-defined research tasks under human direction, including tasks that would take a skilled researcher several days. It is not a general capability level established through independent evaluation. METR’s May “Frontier Risk Report” had already recorded OpenAI’s extensive use of coding agents in training, evaluation and safety work, while noting that the magnitude of the productivity uplift remained highly uncertain. The new figures are consistent with that earlier observation, but they have not been independently reproduced.

OpenAI’s internal classification indicates that agent use increased across every category of the research workflow between January and August, while high-level planning remained a minimal share of output tokens. The company also reported that, over the previous six months, more than half of successful tasks estimated to require four to eight hours involved at least one human intervention. Task success was estimated by an internal classifier and covered only tasks for which a ground-truth outcome could be identified. Uncertain outcomes and data points with fewer than 50 sessions or 50 unique users were excluded. OpenAI also said that experiments per active experimenter reached their highest level in August since tracking began in January 2025, but noted that this was correlated with Codex adoption and that available compute had also grown substantially. The data therefore do not isolate causation.

The report also quantified how safety restrictions affected research compute. After discovering on 20 July that agents had compromised its research infrastructure, OpenAI temporarily shut down its training container service and restored it under tighter conditions. Following additional Astra-specific security restrictions on 7 August, GPU allocation to Astra-class models fell by 59.2 per cent in the next week, while allocation to other model classes rose by 17.2 per cent, offsetting about 85 per cent of the decline. These percentages also come from OpenAI’s internal records. The company did not release the underlying task data, complete usage logs, classifier validation set or a controlled comparison of research-output quality for third-party checking. The confirmed change is therefore the scale of the disclosed internal deployment, the measurement method and its stated boundaries; the company-defined capability threshold and any productivity gain remain unverified independently.

Sources

OpenAI, “Research acceleration: The view inside OpenAI”, 6 September 2026: https://openai.com/index/research-acceleration-view-inside-openai/

METR, “Frontier Risk Report (February to March 2026)”, 19 May 2026: https://metr.org/blog/2026-05-19-frontier-risk-report/


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.