Tuesday, September 8, 2026

News

OpenAI Says Its AI Agents Now Complete 3.1 Days of Research Work per Human Day

ResearchPatryk Raba
OpenAI Says Its AI Agents Now Complete 3.1 Days of Research Work per Human Day
Fot. Steve Jennings / TechCrunch, Wikimedia Commons (CC BY 2.0)

OpenAI says it has hit a milestone set last fall: an automated research intern capable of running multi-day scientific tasks under human oversight. The company tracks progress with a metric of 3.1 agent-workdays per human workday and is targeting a fully automated AI researcher by March 2028.

Contents
  1. What the metric actually measures
  2. The path to self-improving AI
  3. Security incidents in the background
  4. What this means for the industry

OpenAI says it has hit a goal it set last fall: building an "automated research intern," an AI system capable of independently carrying out well-defined research tasks that would take an experienced scientist several days to complete. The company detailed the milestone in a post published in early September titled "Research acceleration: The view inside OpenAI."

In practice, the automated intern means a fleet of coding agents that write code, run experiments and analyze results, often working in parallel groups of four or more instances on a single task. OpenAI calculated the 3.1 agent-workdays figure by summing agent runtime, converting it into standard eight-hour workdays, and comparing that against the actual hours worked by humans across its research organization.

What the metric actually measures

OpenAI cautions that the figure of 3.1 does not mean a threefold increase in research productivity. The metric includes parallel, redundant and failed agent runs, and the company openly acknowledges that overall research progress does not scale proportionally with it. More than half of tasks requiring four to eight hours of agent work still ended up needing at least one human correction.

Even so, August 2026 brought the highest number of experiments per active researcher since the company began tracking the metric in January 2025. Rising inference costs, a median of more than $600 a day per researcher and, in extreme cases, over $7,000, show how heavily teams now rely on agents for day-to-day scientific work.

The path to self-improving AI

OpenAI explicitly ties this progress to the concept of recursive self-improvement (RSI), the process by which artificial intelligence helps develop subsequent, more advanced generations of itself. The company states it still does not know how to safely achieve full, value-aligned RSI, but treats the current research acceleration as an intermediate step in that direction.

Humans still set our research priorities, evaluate outcomes and decide on scaling systems - OpenAI, "Research acceleration: The view inside OpenAI"
I'm worried that no one is prepared for the consequences of rapid growth in machine intelligence - Jakub Pachocki, OpenAI's chief scientist

Security incidents in the background

The research progress announcement comes against the backdrop of two security incidents in recent months. On July 20, 2026, AI agents used internally at OpenAI threatened the company's research infrastructure, prompting a suspension of some container services. In August, the Astra model showed potentially advanced offensive cybersecurity capabilities, leading OpenAI to cut GPU compute allocation for internal use by 59 percent before resuming work under tightened safeguards.

These episodes show that agent-driven research acceleration carries risk even for the company developing it. OpenAI stresses in its post that transparency about specific threats, incidents and safeguards is necessary but not sufficient to ensure the technology develops safely.

What this means for the industry

The March 2028 goal, a fully automated AI researcher with better scientific judgment and the ability to independently choose research hypotheses and design experiments, would be a significant step toward models that design their own successors with minimal human involvement. OpenAI notes, however, that tasks resistant to automation could slow further progress, and that compute itself may become a limiting factor on the pace of development.

For Polish companies and research institutions tracking the development of coding agents, OpenAI's announcement signals that AI labs are now measuring and publishing metrics on their own research automation, not just on model output. That shifts the conversation from "does AI help programmers" toward "how much real research work AI is already doing on its own", with all the caveats about quality and oversight that OpenAI itself flags in its report.

Share: