AI/ ai · llm research · eu ai act · research methods

LLM Text Scoring Is Consistent but Misses the Point, Paper Finds

An arXiv study of EU AI Act consultation responses shows LLM text annotations can be highly consistent yet still measure the wrong thing entirely.

Consistent answers from an AI model are not the same as correct ones.

A new paper on arXiv, "Reproducibility is not construct validity: LLM measurement of institutionally situated communication" (arXiv:2609.19866, posted September 18, 2026), tests what happens when you use an LLM to score real-world documents instead of survey responses. The paper's authors linked structured survey answers from stakeholders in the European Commission's AI Act consultation to free-text submissions from those same stakeholders, then had an LLM rate the free-text for concern about AI risk. The LLM's scores were remarkably consistent, with intraclass correlations above 0.99, but they lined up poorly with what the same stakeholders reported on the survey. Business associations voiced far more AI-risk concern in their written submissions than in their survey answers, a gap of about one standard deviation, while public authorities and several nonbusiness groups showed smaller or reversed gaps.

That is a construct-validity problem, not an accuracy problem: the model was reliably measuring something, just not the thing the survey measured. It is a warning for anyone using LLMs to score policy submissions, employee feedback, or open-ended survey text at scale, since a model can output the same result every time and still be systematically wrong about what it is capturing. The paper also found the divergence clusters geographically across European countries (Moran's I = 0.347), suggesting national communication norms, not just underlying attitudes, shape how stakeholders write versus how they answer checkboxes.

Consistency has become the easy selling point for AI research tools. This paper is a reminder that consistency was never the hard part.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →