AI/ test-time-training · self-improvement · llm-research · agentic-ai

A model that decides when to retrain itself

Agentic-TTT adds a policy layer that decides when test-time training helps, nearly doubling utility over the base model in benchmarks.

A model that decides for itself when to retrain on the fly - and when not to bother.

Agentic-TTT is a policy trained to govern test-time training (TTT), the technique of updating a model's parameters using signals from the inputs it is currently working on. TTT already produces gains in narrow settings like IMO-style competitions, but it is finicky - the wrong TTT algorithm applied to the wrong problem can waste compute or even hurt performance. Agentic-TTT turns TTT procedures into callable tools and treats the skills a model accumulates as an evolving environment, training its policy on the actual utility gained from each decision. In testing, it nearly doubled the backbone model's utility and still generalized to problem domains it never saw during training.

The real contribution here is not the retraining trick itself, it is teaching a model when to use it, which is a basic requirement for any kind of durable self-improvement. Most "self-improving" demos are one-off showcases tuned for a single competition or benchmark. A policy that weighs compute cost against expected benefit is a step toward systems that manage their own learning loop during deployment, not just inside a lab setup built for a paper.

Whether that judgment holds up against messy real-world deployment traffic, instead of the curated domains used here, is the harder question the benchmark numbers do not answer.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →