News · 2026-08-23
Gallup is testing AI agents that answer surveys for real people
Gallup has begun independently validating whether AI agents built from long interviews can predict how real people answer survey questions, in a partnership with the Stanford-founded startup Simile. About 1,000 members of Gallup's probability-based U.S. panel sat for in-depth interviews starting in fall 2025, and each agent is grounded in one person's transcript plus their survey answers. Gallup says the simulated responses will never appear in its published population estimates.
Key facts
- Roughly 1,000 Gallup Panel members completed in-depth interviews used to build the agents; validation began this year.
- Simile has raised over $200 million at a $2 billion post-money valuation, co-led by Greenoaks and Index Ventures.
- Gallup's stated limit: simulated responses "will not be used to produce Gallup's published population estimates."
- Primary source: Gallup's methodology blog post.
The idea has a paper behind it. In Generative Agent Simulations of 1,000 People, a Stanford-led team interviewed 1,052 Americans for two hours each, built a language-model agent for every participant grounded in that transcript, and then asked the agents questions the humans had also answered but the agents had never seen. The headline result is usually quoted as 85 percent accuracy, and the number needs unpacking, because it is not measured against truth. It is measured against the humans themselves: the same people were re-surveyed two weeks later, and they did not perfectly agree with their own earlier answers. Against that ceiling, interview-grounded agents hit 83 percent, survey-grounded agents 82 percent, and agents with both 86 percent -- while agents built only from demographics managed 74 percent.
That framing is the whole story. The benchmark is human inconsistency, and the claim is that an agent grounded in what you actually said about yourself gets closer to predicting you than a stereotype built from your age, income, and zip code. The paper's less-quoted finding is why it works: part of the gain is simple lookup from the transcript, but part is genuine inference from unrelated answers. Strip out the questions an agent could answer by searching the transcript and interview-grounded agents still beat demographic ones. Strip out the inference-friendly ones too and the gap narrows but does not close. Our lesson on simulating people with language models covers the method in more depth.
The commercial version has moved faster than the science. Simile, founded by researchers behind that paper, says on its Series B announcement that it has grown revenue fivefold in five months, run tens of millions of simulations for Fortune 100 enterprises, and raised over $200 million at a $2 billion post-money valuation. Its customer list names CVS Health, Wealthfront, Deloitte, and Gallup itself. The company states its mission with no hedging at all: "to simulate all eight billion people on earth, accurately and honestly."
Gallup's posture is the interesting counterweight, and it comes from inside the partnership rather than outside it. The organization is explicit that it is testing "where these AI-generated agents perform well in predicting people's responses, where they fall short and how they compare to established methods." It reports early findings that on general population estimates for topics close to what the interviews covered, the simulated distributions were "close enough to approximating human responses for us to warrant continued exploration" -- which is about as restrained as an encouraging result can be phrased. And it names the risk directly: "We acknowledge that this technology has the potential to erode public trust."
Why it matters: survey research is expensive, slow, and getting harder as response rates fall, and simulated respondents are the most plausible shortcut anyone has proposed. If they work even for narrow uses -- pre-testing questionnaire wording, exploring hard-to-reach populations, sizing a hypothesis before fielding it -- that changes the economics of a large industry. If they get adopted for the uses they do not work for, it changes what "a poll said" means.
The strongest counter-argument comes from Joni Salminen, who has catalogued five failure modes of synthetic users: they compress human variety, extrapolate badly outside what they were grounded in, mirror the stances implied by the prompt, fail precisely when real behavior would have surprised you, and create a validation paradox -- you can only confirm the simulation is right by running the human study it was supposed to replace. The paper's own data supports the caution: the authors note the accuracy gains flatten once enough grounding evidence is in hand, so richer interviews do not buy unlimited fidelity. Gallup's line is the one to hold onto: simulated and human responses "are not interchangeable."
Key questions
Will Gallup publish poll results generated by AI?
How are the agents built?
How accurate is the underlying method?
Cite this
APA
Ground Truth. (2026, August 23). Gallup is testing AI agents that answer surveys for real people. Ground Truth. https://groundtruth.day/news/gallup-is-testing-ai-agents-that-answer-surveys-for-real-people.html
BibTeX
@misc{groundtruth:gallup-is-testing-ai-agents-that-answer-surveys-for-real-people,
title = {Gallup is testing AI agents that answer surveys for real people},
author = {{Ground Truth}},
year = {2026},
month = {aug},
url = {https://groundtruth.day/news/gallup-is-testing-ai-agents-that-answer-surveys-for-real-people.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.