An AI model can decide on its answer long before it finishes writing out the reasoning behind that answer.
A new arXiv paper studies "masked diffusion" multimodal language models, which build answers and explanations by filling in blanks across a shared canvas rather than writing word by word, and tests one such model, LaViDa, across three visual question-answering benchmarks in a single-block, EOS-suppressed setup. In those LaViDa runs, 89.4% to 98.1% of the rationale canvas was still unwritten at the exact moment the model's answer stabilized. On the VBench benchmark specifically, shrinking LaViDa's generation block size from 128 tokens to 8 dropped that unwritten fraction from 89.4% to just 1.7%. In a separate comparison using a different model, Nemotron, adding direct instructions under EOS-enabled prompting raised its accuracy by 15.0 and 19.5 percentage points on two other benchmarks, while the same technique cut LaViDa's VBench accuracy by 11.0 points - a difference the paper traces mostly to answer coverage rather than reasoning quality.
That LaViDa-specific stabilization number matters because it undercuts the pitch that diffusion models "think out loud." If the answer locks in before most of the explanation exists, the explanation is not driving the decision - it is decoration generated alongside it. The researchers do not claim this pattern generalizes beyond LaViDa's tested configuration, and that restraint is doing real work here.
Call it the opposite of showing your work: writing it after you have already handed in the test.