Ground Truth.
AI, checked against the source.

News · 2026-09-06

OpenAI reports 3.1 agent-workdays for every human research workday

OpenAI says its research organization reached 3.1 agent-workdays of effort for every human workday by mid-August 2026. The company presents the figure as a measurement, not a forecast, and calls its measurement work preliminary. The number marks a large internal agent deployment, but it is not a public proof that agents have replaced researchers or a common productivity standard.

Key facts

The phrase to retain is 'according to our measurements.' OpenAI explains the unit only at the surface level: an eight-hour day. It does not publish a detailed formula mapping tokens, tool calls, code changes, successful tasks, and human review into human-equivalent effort. That does not make the number empty; it tells readers what it is. This is a company-specific internal throughput indicator, more like a report of machine-hours than a universal measure of completed scientific value.

The work definition is broader than code generation. OpenAI divides research into Decide, Design, Build, Run, Analyze, and Communicate, including ideas, specifications, code, datasets, training and evaluation runs, deployment work, results, feedback, and decisions. It says researchers are writing more code and running more experiments while high-level planning remains a small fraction of agent output. The agents are therefore substantial labour inside a human-directed pipeline, not an autonomous research department.

The automated research intern label is useful precisely because it is bounded. OpenAI defines it as a system that carries out well-defined research tasks under human direction, including tasks a skilled researcher might take days to finish. An intern can speed a lab without setting its agenda, judging a surprising discovery, or carrying responsibility for an error. The named March 2028 target is a concrete checkable commitment.

The companion story is control. OpenAI's monitoring post says it reviews full histories, reasoning traces, tool calls, and outputs across standard internal coding-agent deployments. It reports tens of millions of trajectories in five months, no highest-severity real conversations, and about 1,000 moderate-severity alerts, while acknowledging some bespoke and local setups lie outside coverage. Agent labour only becomes an organizational asset if it is observable and reversible.

An apt analogy is a construction crew with more power tools. Tool-hours describe capacity, not whether the building passed inspection or whether the crew chose the right design. The metric could overcount redundant subagents and failed attempts, or undercount value when an agent unblocks a decisive experiment. Only task-level outcome, error, cost, and review data can resolve the balance.

For a distinct external yardstick, METR's task-horizon work models success against how long a human expert needs for a task and warns estimates beyond 16 hours are presently unreliable. It measures something different, which is useful. The industry needs deployed-effort metrics, time-horizon metrics, and outcome metrics together. The enduring news is not 'three agents equal three people'; it is that a frontier lab is publicly treating agent labour as something to measure, monitor, and schedule.


Primary source, verified: read the paper →

Key questions

What does OpenAI mean by an agent-workday?

OpenAI normalizes the unit to a standard eight-hour workday and includes both directly launched agents and downstream subagents.

Does 3.1 agent-workdays mean agents replace 3.1 researchers?

No. It is an internal preliminary throughput measurement and OpenAI does not publish a full conversion formula from agent activity to researcher output.

What work are the agents doing?

OpenAI says they span deciding, designing, building, running, analysing, and communicating, including code, data, experiments, and evaluations.
Cite this

APA

Ground Truth. (2026, September 6). OpenAI reports 3.1 agent-workdays for every human research workday. Ground Truth. https://groundtruth.day/news/openai-reports-3-1-agent-workdays-per-human-day.html

BibTeX

@misc{groundtruth:openai-reports-3-1-agent-workdays-per-human-day,
  title  = {OpenAI reports 3.1 agent-workdays for every human research workday},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {sep},
  url    = {https://groundtruth.day/news/openai-reports-3-1-agent-workdays-per-human-day.html}
}

Topics: openai · agents · research · productivity · recursive-self-improvement

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.