News · 2026-08-19
One person with AI matched a two-person team at Procter and Gamble
A randomized field experiment inside Procter and Gamble found that one professional working with an AI assistant produced solutions rated as high as those from two-person teams working without one. The study, published in Organization Science on June 12, 2026, ran a preregistered two-by-two design across 826 participants at the company, with blind expert evaluators scoring the output. The task was real product development work, not a benchmark.
Key facts
- The paper is The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork, published online in Organization Science on June 12, 2026 and summarized by Harvard Business School.
- 826 P and G participants, with analyses focusing on 791 professionals, randomly assigned to individual or two-person team, with or without AI.
- Randomization was stratified by business unit and geography; team conditions paired one commercial and one R and D professional.
- Blind experts scored quality on a 1 to 10 scale, plus novelty, feasibility, impact and business potential.
Most claims about AI and productivity come from one of two weak places: a vendor benchmark, or a survey asking people how productive they feel. This is neither. It is a preregistered randomized controlled trial run inside a real company on real work, which makes it one of the more credible measurements available. The AI tool was GPT-4-based, accessed through Microsoft Azure, and the study was run by researchers from HBS, Wharton, ESSEC and Warwick together with P and G collaborators.
The design is what earns the result. Participants worked on genuine problems from their own business units during a one-day virtual workshop, following a workflow that mirrored the company's actual early-stage development routine: generate ideas, select one, develop it. Team assignments deliberately crossed functions, pairing a commercial professional with an R and D professional, because the standard argument for teams in product development is precisely that they integrate expertise from different silos. One team member was randomly designated to share their screen and submit.
The result is that the AI-assisted individual closed that gap. A single person with an assistant reached the same quality level as the cross-functional pair. The authors also report that the gain concentrated in idea generation rather than in selection, which is the more interesting detail: the assistant was good at producing options and less decisive in choosing among them, where human judgment still carried the work.
The natural analogy is a good reference librarian rather than a second engineer. The librarian does not do your job, and cannot decide which of two directions your project should take. What they do is put five options in front of you within minutes instead of you finding two by yourself over an afternoon.
The framing matters enormously here, and the researchers are more careful than the headlines about it. Their conclusion is about collaboration and expertise integration: AI can replicate some benefits of human collaboration on this innovation task, reduce functional silos, and change social engagement patterns. The defensible sentence is "one person with AI matched a two-person team on a product-innovation workshop task." The sentence "AI took two jobs" is not supported by anything in the paper. A one-day workshop is not a quarter of shipped work, and matching on a blind quality score is not the same as replacing a colleague's ongoing contribution to a team.
That distinction is doing real work right now, because the labor sentiment is running well ahead of the evidence. Gallup reported on July 28, 2026 that adults 18 to 29 are increasingly skeptical of AI and more likely to expect it to reduce US jobs over the next decade. The Federal Reserve's May 2026 household well-being report found that workers under 30 were more likely to worry AI will replace their job than to say it will improve their career. Studies like this one are cited in both directions in that argument, usually with the caveats removed.
There is also a gap between measured performance and perceived performance that keeps showing up. We covered a case where the benchmarks said Opus 5 improved and the people using it disagreed, and OpenAI's own usage analysis, which we covered as one in six work prompts crosses job lines, suggests the actual pattern of AI at work is messier than either the optimistic or pessimistic story.
The honest caveats are the ones any field experiment carries. The AI tool was a 2024-generation model, so the result is a floor rather than a current measurement. One company, one day, one workflow, in one industry. Blind expert scoring is a good proxy for solution quality but not a measurement of whether the solution shipped or made money. And a workshop is designed to produce ideas in a day, which is exactly the task shape where an idea-generating assistant should look best.
What the study genuinely establishes is narrower and still important: on early-stage innovation work, an assistant substituted for a specific benefit of teamwork, the cross-functional perspective, well enough that trained evaluators could not tell the difference in quality.
Key questions
What exactly did the experiment measure?
Was this a realistic task or a benchmark?
Does this mean AI replaces a colleague?
Cite this
APA
Ground Truth. (2026, August 19). One person with AI matched a two-person team at Procter and Gamble. Ground Truth. https://groundtruth.day/news/one-person-with-ai-matched-a-two-person-team-at-procter-and-gamble.html
BibTeX
@misc{groundtruth:one-person-with-ai-matched-a-two-person-team-at-procter-and-gamble,
title = {One person with AI matched a two-person team at Procter and Gamble},
author = {{Ground Truth}},
year = {2026},
month = {aug},
url = {https://groundtruth.day/news/one-person-with-ai-matched-a-two-person-team-at-procter-and-gamble.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.