Ask a large language model to predict a number - a chemical's melting point, a snippet of code's runtime - and it'll give you one. Ask it how confident it is, and the answer is usually garbage.
Researchers have introduced Distribution-Aware Reward, or DAR, a reinforcement learning method that trains language models by scoring groups of predictions together instead of one at a time. Each prediction gets credit based on how much it improves the overall spread of guesses for a given input, calculated by checking what happens when that prediction is left out. The goal is a set of predictions that clusters around the right answer without being either overconfident or needlessly scattered. The team tested DAR on a synthetic benchmark plus two real scientific tasks involving code and molecular data, comparing it against standard supervised fine-tuning and simpler pointwise reinforcement learning.
This matters because LLMs are quietly becoming general-purpose regressors, plugged into scientific pipelines to estimate properties from messy, mixed-format data where a built-in uncertainty estimate is the whole point. A model that nails the average prediction but is wildly over- or under-confident about its own error bars is not actually useful for that job. DAR reduced prediction error and sharpened uncertainty estimates across all three tasks compared to the alternatives.
Three benchmarks is not production deployment, and calibration measured inside a paper's own evaluation is not the same as trustworthy in the wild. Still, teaching a model to reason about its own spread of guesses, rather than just its single best guess, is the kind of unglamorous fix that tends to matter more than it sounds.