AI/ ai · llm-training · machine-learning · research

A Fix for AI Models That Reason Without Words

A new training method called REST catches errors invisible to standard AI training, boosting accuracy on math and coding tests by up to 7.5 points.

AI systems that reason in raw numbers instead of words have a hidden training gap, and a new paper claims to close it.

Some large language models now reason by recycling their own internal, numeric hidden states, either looping on them internally or passing them between multiple AI agents, instead of writing out each step as text. The catch: training only checks the final decoded answer, never the intermediate "thought" itself. A new arXiv paper documents four specific ways that blind spot backfires, including thoughts that blur together across different questions and thoughts that hang onto irrelevant details. The fix, called REST (REpresentation-Supervised Thoughts), adds four extra training penalties, for causality, minimality, separability, and stability, on top of standard training, without changing the model's architecture or adding parameters.

Across seven benchmarks covering math, science, medicine, and code, REST lifted accuracy by up to 7.5 percentage points and got models to converge on a final answer 30 percent more often, using identical data and compute. Just as important, REST made those internal thoughts easier to decode into something resembling the model's actual intent, which matters because reasoning that happens outside visible text is otherwise a black box.

Latent reasoning has been pitched as a faster, cheaper alternative to writing out every step in English. This paper is a reminder that efficiency and reliability are not the same thing, and that a model that thinks without showing its work needs its own scaffolding to stay honest. It is a single, not-yet-peer-reviewed preprint, so the real test is whether other labs can reproduce the gains outside the authors' own benchmarks.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →