A new reinforcement learning method teaches AI agents to flag their own blind spots instead of just reacting after the fact.
Researchers introduce Prospective Hindsight (PH), a training technique that adds a self-calibration signal on top of standard reinforcement learning. Before an agent acts, it predicts how things will go; after the environment responds, that prediction gets checked against what actually happened. PH uses the size of that gap, which the paper calls surprise, to decide which training examples deserve more weight, with a stop-gradient trick keeping the self-prediction separate from the main policy update. Tested on single-turn tasks and a multi-turn personal-agent task, under GRPO, on-policy distillation, and a combination of both, PH improved task performance and calibration across different model scales.
The more interesting finding is that the type of miscalibration it fixes depends on context. In single-turn tasks, agents tend to be overconfident right before they fail. In multi-turn tasks, they tend to be underconfident right before they succeed. PH corrects both with the same mechanism, which suggests the fix targets self-knowledge generally rather than patching one specific failure mode.
That is a real departure from most RL work, which spends its effort tuning how well an agent predicts the world, not how well it predicts itself. Whether this holds up outside the paper's own benchmarks is the obvious next question, and the preprint does not pretend to answer it.