Synthetic Polls Fall Short

Updated: 2026.09.30 39MIN ago 1 sources
Large language models can imitate some public‑opinion answers but produce systematic, model‑dependent errors on many questions; accuracy varies by question type, prompting, and model version. Pew’s replicated tests (GPT‑5.1 and Claude Opus 4.6) show some questions are well matched while others diverge by multiple percentage points, so synthetic samples cannot yet replace human interviewers for reliable polling. — If media, campaigns, or researchers accept synthetic survey outputs uncritically, they risk amplifying biased or unstable measurements that could misinform policy debates, election coverage, and scholarly work.

Sources

Appendix B: Additional tables
Janakee Chavda 2026.09.30 100% relevant
Pew Research Center synthetic‑respondent experiments comparing GPT‑5.1 and Claude Opus 4.6 against ATP Wave 185 real survey data (Jan–Apr 2026), reporting per‑question average absolute errors and model differences.
← Back to all ideas