AI/ reinforcement learning · ai training · language models · ai research

New Method Speeds Up AI Reasoning Training by 63 Percent

Researchers propose T5, a twin-critic technique that speeds up reinforcement mid-training for AI reasoning models without sacrificing accuracy.

A new technique called T5 cuts the cost of teaching language models to reason, without the usual accuracy tradeoff.

Researchers posted a paper to arXiv describing T5, short for Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training. It targets reinforcement mid-training, the stage where models learn internal reasoning steps from plain, unlabeled text. The dominant approach today generates multiple candidate responses per prompt just to figure out which tokens deserve credit, which is slow and expensive. T5 instead uses two learned critic models that each estimate how good a token-level decision was, then blends the two estimates using weights learned through a conditional-moment saddle-point objective, a calibration step the authors say corrects a bias that creeps into simpler single-critic estimates and compounds over training.

Mid-training is the unglamorous plumbing behind the step-by-step reasoning style now common in newer models, and its cost scales with how many rollouts a method needs per prompt. A single-rollout method that matches or beats multi-rollout ones is a direct hit to training budgets, not just a leaderboard score. The paper reports a 7.8% average benchmark improvement and training steps up to 63.4% faster than the best critic-free baseline, numbers large enough to matter for labs without frontier-scale compute.

Those figures come from the authors' own experiments against one baseline category, so treat them as a ceiling, not a guarantee, until outside labs reproduce them.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →