AI/ ai · llms · policy monitoring · research

LLMs Match Humans on Structured Policy Survey Answers

A new study finds AI models replicate human survey responses on policy documents with 84-95% agreement, though free-text answers still diverge.

AI can now sit in for the human analysts who fill out international policy surveys - and it gets the multiple-choice-style questions right most of the time.

A pipeline described in a preprint posted to arXiv (arXiv:2609.29370, posted September 25, 2026, and not yet peer reviewed) feeds public policy documents to a large language model and asks it to answer the same structured survey questions that human researchers currently fill out by hand - things like what policy instruments a program uses, who it targets, and what theme it falls under. A second LLM checks the first model's answers for relevance and supporting evidence before they get compared against real human-generated responses across a multi-country dataset. On structured, multiple-choice-style questions, the AI's answers matched humans' 84 to 95 percent of the time. On open-ended, free-text questions, the model tended to write longer, more procedural descriptions than human respondents did.

That gap matters because science and innovation policy surveys are currently a manual, country-by-country slog - expensive to run and hard to keep consistent, which is why cross-country policy comparisons often lag years behind the policies themselves. An 84-95% match rate on the parts of the survey that feed into cross-country databases suggests the costliest part of this work could be automated, with humans reserved for spot-checks and messier qualitative questions.

Worth flagging: this is one unreviewed preprint testing one pipeline, not a validated replacement for policy researchers. And a model over-explaining its free-text answers is a familiar LLM habit, not proof it understands the policy better than the humans it's being graded against.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →