A new controller can make large language models reason faster and get things right more often, without touching their weights.
A paper posted to arXiv, 'MetaCtrl: Your Large Language Models Can Reason Better and More Concisely with a Metacognitive Controller' (arxiv.org/abs/2609.37304), describes a lightweight add-on trained with reinforcement learning to watch a reasoning model's step-by-step output and decide, in real time, whether to keep going, simplify, skip ahead, or stop. The reasoner itself stays frozen - MetaCtrl doesn't require retraining it or setting a fixed token budget in advance. Tested across seven math, science, and coding benchmarks, the paper reports MetaCtrl raised DeepSeek-R1-Distill-Qwen-7B's average accuracy by 4.7 points while cutting its output length by 53.3%. Without further training, the same controller transferred to a different model, Qwen3-14B, improving accuracy by 2.9 points and shortening responses by 50.3%, according to the paper.
Reasoning models burn tokens, and money, generating long chains of thought that don't reliably track with correctness - sometimes the extra steps help, sometimes they're just padding. A controller that learns separately when to stop thinking, instead of baking that judgment into the base model, is a cheaper lever than retraining every reasoner from scratch. The more interesting claim is the cross-model transfer: it suggests knowing when to stop can be a portable skill rather than something each model has to relearn.
Efficiency claims like these often look tidier in a paper's own benchmark suite than they do once other labs try to reproduce them, so treat the percentages as one team's numbers until that happens.