AI/ ai · llm reasoning · test-time compute · beam search

A New Way to Make AI Models Second-Guess Themselves

A new decoding method uses hypothetical good and bad reasoning prompts to steer language models toward better answers without any retraining.

A new decoding trick makes language models act as their own reward model, with no training required.

The method, called test-time self-distillation, adapts an existing training-time technique known as Self-Distillation Fine-Tuning, which had models learn from a demonstration via an implicit reward signal but only worked with gradient updates and expert examples. The new approach swaps that demonstration for two fixed text templates instead, one written to prime the model for excellent reasoning and one for poor reasoning, then measures the log-odds ratio of a candidate answer's likelihood under each. That ratio becomes the reward. Because applying it token-by-token would ignore how good the finished answer actually is, the researchers approximate the ideal output using beam search. Tested on MATH500 math problems, HumanEval coding tasks, and GPQA science questions across multiple model sizes, it beat standard sampling, low-temperature sampling, plain beam search, and power sampling, on average.

Test-time compute, getting more out of a model without retraining it, is the industry's preferred cost-saving trick right now, and most labs are chasing some version of it. What stands out here is the source of the steering signal: no separate reward model, no human feedback, just two prompt templates the model critiques its own output against. If the results hold up outside these three benchmarks, it is a cheap route to self-consistency-style gains without generating dozens of samples and voting.

"On average" wins across three benchmarks are a real result, not a rounding error, but beam search means more compute spent per answer, and the paper does not say how that overhead compares to simply asking the model twice and picking the better response.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →