OpenAI’s Researchers Are Increasingly Relying on AI Agents, Even as Safety Concerns Grow
OpenAI researchers are using AI agents for a growing share of their work, a shift that could accelerate model development while increasing pressure on the company’s safety and oversight systems.
The company said its researchers had reached an “automated research intern” milestone, with agents able to complete well-defined research tasks that could otherwise take a skilled researcher several days. By mid-August, OpenAI said it was logging 3.1 agent-workdays for every human workday among its research staff.
The increased use of agents comes with substantial computing demand. OpenAI said the median researcher was using more than $600 a day in inference capacity at API-equivalent prices.
However, the company’s own figures indicate that human review remains central to the process. More than half of successful agent assignments lasting four to eight hours required at least one human intervention, according to OpenAI.
Chief Scientist Jakub Pachocki separately warned that more capable systems could create a path toward recursive self-improvement, in which AI plays an increasingly large role in developing subsequent generations of AI. He argued that alignment and monitoring methods have not advanced at the same pace as model capabilities.
Pachocki distinguished between goal alignment—whether an AI system pursues a specified objective—and value alignment, which concerns whether the system continues to respect human constraints when accomplishing an objective becomes difficult. A system could appear effective at following assigned goals while still finding harmful or unintended ways to optimize for them, he said.
OpenAI has placed emphasis on monitoring models’ chain-of-thought reasoning, or the verbalized rationale produced during a task. Pachocki cautioned that this form of visibility could become less useful as systems grow more capable.
The developments highlight a central operational issue for companies adopting autonomous agents: productivity gains may depend on retaining human checkpoints, clear task boundaries and review processes, particularly for longer or higher-stakes work.