AI/ finance-nlp · sentiment-analysis · llm-evaluation · research-methods

Sentiment Tools Can Match Humans Without Predicting Stock Moves

A study of 70,500 X posts tied to securities lawsuits finds that sentiment tools matching human judgment don't reliably predict stock returns.

A new study finds that a sentiment tool matching human judgment tells you almost nothing about whether it can predict stock moves.

Researchers ran five sentiment-scoring tools - VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator - through the same pipeline on 70,500 X posts tied to securities class-action lawsuits filed between 2002 and 2025. They compared each tool's output against abnormal stock returns and against a human-labeled sample of the same messages. The result changed depending on how messages were sampled and how scores were represented: under standard sampling, human agreement tracked same-day return moves better than next-day moves, but on a more controlled fixed-size sample, agreement looked similar at both horizons while coarse buy-or-sell rankings stayed weak throughout. The dataset itself was messy too - 17.6% of the messages were spam, and how many messages showed up had no relationship to how much money was lost or how big the eventual settlement was.

The finance-NLP field has leaned on a shortcut: if a sentiment model matches human labels, treat it as validated for predicting markets too. This paper shows that shortcut can fail quietly, because the two kinds of validity respond differently to sampling choices that most published benchmarks don't report.

Sentiment scores make for tidy charts, but this is a reminder that sounding right to a human annotator and actually moving markets are separate claims - and swapping one for the other is exactly the kind of shortcut that gets exposed by outcomes, not benchmarks.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →