AI tools that summarize employee feedback for executives are quietly deleting the opinions nobody else repeated.
Researchers built a Voice Retention Ratio to measure which employee comments survive when an LLM condenses raw feedback into a leader-facing summary. They tested it on 2,586 bilingual English and German free-text responses from a global professional-services company, run through 45 actual leader-summaries. The pattern: criticism makes it through fine, but a concern raised by just one person gets cut 86% of the time. Short comments fared far worse than long ones, retained at a rate of 0.14 versus 0.74, and the researchers separately flag a preliminary, not-yet-statistically-robust finding that German-only comments showed a similar drop-off.
The real finding here is not sentiment bias, it is popularity bias. Once the team controlled for how often a comment was repeated, sentiment stopped predicting what survived at all; length and frequency did the work. That matters because employees already under-report praise relative to criticism by a factor of 82 to 1 in this dataset, so a summarization layer that further prunes anything said only once compounds an existing distortion instead of correcting it.
Call it an engagement algorithm wearing an engagement-survey costume: the loudest complaint wins, the quiet-but-real one disappears, and a targeted prompt fix only recovers themes someone already thought to name.