AI/ ai · llm-reasoning · interpretability · research

Researchers Build a Ruler for AI Models That Ramble

SLIDER uses information theory to flag when an AI model's reasoning step just repeats earlier ones, and the signal can train leaner models.

A new interpretability tool called SLIDER can finally tell you when an AI model's reasoning is just spinning its wheels.

Researchers built SLIDER using partial information decomposition, a method from information theory, to split the information in each reasoning step into three parts: what's genuinely new, what's redundant with earlier steps, and what only makes sense combined with earlier steps. From that split they derive Step-RRI, a score for whether one step is mostly just restating what came before, and Trajectory-RRI, an aggregate score for an entire reasoning chain. On the redundancy portion of the PRMBench dataset, Step-RRI beat embedding-similarity and information-gain baselines by more than 10 points at catching repetitive steps. Across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and GPT-4.1, Trajectory-RRI tracked closely with how long each model's reasoning actually ran.

Reasoning models are notorious for burning tokens re-deriving the same logic before landing on an answer, and that waste adds up in cost and latency at scale. A measurable, theory-grounded signal for redundancy - rather than an eyeballed "this looks repetitive" - gives teams a concrete way to filter training data toward leaner reasoning. The paper reports that fine-tuning on Trajectory-RRI-selected data cut down rambling while largely holding onto task performance.

It's a diagnostic, not a fix: the paper doesn't claim SLIDER makes a model smarter, only that it tells you exactly where its reasoning is stalling.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →