Researchers have found a way to watch a language model change its mind, one word at a time.
The team took a transformer-based language model, RoBERTa, and recomputed a token's contextual word embedding (the number string a model uses to represent a word's meaning in context) every time a new word was added to a sentence. Strung together, those snapshots form a trajectory showing how the model's read on a word shifts as the sentence unfolds. They tested this on garden-path sentences, the classic trick sentences like "the horse raced past the barn fell," which lead a reader toward one meaning before forcing a last-second reinterpretation. The trajectories showed a sharp disruption right at the point where the sentence flips on the reader, and reliably separated garden-path sentences from otherwise similar sentences with no such twist.
The more interesting finding is where that disruption shows up. Researchers usually check a single summary token (called CLS) to see what a model "thinks" about a whole sentence. Here, the ambiguity signal also showed up in ordinary, everyday tokens scattered through the sentence, meaning a model's sense of confusion is not stored in one tidy spot but spread across the sentence as it reads.
It is a clean result, but a narrow one. RoBERTa is a few generations behind the chatbots people use today, and garden-path sentences are a well-worn test case in linguistics. Whether the same word-by-word trajectories show up in larger, more conversational models dealing with messier, real-world ambiguity is still an open question.