Pew Study Finds AI Synthetic Surveys Cannot Reliably Replace Human Respondents
Summary
Pew Research Center tested whether AI-generated “digital twins” could reproduce the results of high-quality public opinion surveys. Researchers gave an AI model demographic information and earlier survey answers from members of the American Trends Panel, then asked it to answer the same questions as those panelists across three survey waves conducted in the first half of 2026. Nearly 300 questions showed an average absolute difference of 12 percentage points between synthetic and human respondents, and the gap exceeded 15 points on about 28% of questions. Average error was especially high among Republicans and Republican-leaning respondents, Black adults and other subgroups. The model also missed current political and public-interest questions, including views on immigration officers wearing face coverings, approval of President Donald Trump’s job performance, and awareness of data centers. Synthetic respondents often avoided answer choices representing strong support or opposition, reproduced subgroup stereotypes as near-universal views, and selected “not sure” far less often than humans. They also answered factual knowledge questions much more accurately than human respondents, suggesting that the model overstated what the public knows. Results varied by model: in a comparison with OpenAI’s GPT-5.1, Claude Opus 4.6 portrayed a more middle-of-the-road public while GPT-5.1 produced more extreme estimates, with both differing from reality. Pew concludes that AI respondents are not currently an adequate replacement for rigorous, probability-based surveys of real people, although AI can still assist with tasks such as coding open-ended responses and analyzing survey data.