A new paper shows a simple fix for an overlooked flaw in how AI systems rank candidate actions: the order those candidates are listed in can quietly change the score, even when nothing about the actual choices has changed.
Researchers built an attention mechanism called candidate-independent block-causal attention. It keeps the model's normal left-to-right reading pattern for the shared context and within each individual candidate, but blocks candidates from seeing each other and resets their position numbering so none gets an unfair head start. The team tested the approach against standard attention on three open-source model backbones: Gemma 3 1B, Qwen3 1.7B, and Qwen3 4B. Across all three, the new method made scores far less sensitive to how candidates were ordered, without giving up decision quality. Follow-up tests isolated candidate blocking, not position resetting, as the bigger contributor to that stability. The team also ran a larger Qwen3 4B experiment with more training data and released both code and a trained model publicly.
This targets a problem that rarely shows up in standard benchmarks but matters once models start acting as judges or selectors inside bigger AI systems, picking between tool calls, generated replies, or retrieved documents. If a ranking flips just because candidates got shuffled, every downstream decision inherits that noise, and comparisons between systems become harder to trust.
The fix is narrow by design, and that is its strength: it changes how candidates attend to each other, not the model's core capabilities. Whether it holds up at the scale of the much larger models actually doing this job in production is, for now, untested.