AI/ ai · test-time-scaling · llm-reasoning · open-source-ai

New Framework Lets AI Models Decide Their Own Reasoning Steps

Researchers built Hermes, a training method that teaches smaller open-source models to manage their own inference compute instead of following rigid rules.

A new paper argues that AI models should decide for themselves how to use extra "thinking time," rather than following a fixed script written by their software harness.

Researchers introduce a family of harnesses called Hermes that let a model control how it allocates and reuses context windows during inference, plus a training method called Hermes-Learn that teaches models this skill directly instead of hard-coding the rules into the surrounding system. In tests, models that already reasoned well could exploit that flexibility to get more out of extra inference-time compute. Smaller open-source models could not do this on their own. After training with Hermes-Learn, those smaller models learned to adapt their strategy based on both the specific problem and how far their reasoning had already progressed.

Test-time scaling - giving a model more compute per query instead of training a bigger one - is one of the cheaper levers labs have for squeezing out better performance without a full retrain. This work suggests the limiting factor isn't just how much extra compute you hand a model, but whether it knows what to do with it, a skill that currently splits along lines of model size and training approach. The reported gains held across different benchmarks and models, and reportedly carried over to other test-time scaling methods beyond Hermes itself.

Promising, but it's one paper without named benchmarks or released code in hand, so "closes the gap" is a claim worth watching, not a verdict.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →