A new study finds diffusion language models are good at writing code edits, but bad at figuring out where those edits should go.
Researchers tested masked diffusion language models on code editing using four interfaces - whole-file rewriting, search-and-replace, locate-then-infill, and token-level editing. Using the CanItEdit benchmark, they found what they call a composition gap: when a model is told exactly which lines need to change, it generates coordinated edits successfully. But when it has to predict those edit locations itself, much of that ability disappears. Seeing the full original file helps the model fill in multiple edit regions at once, but it does not solve the harder problem of choosing where those regions should be. A similar test on sentence-level Wikipedia edits showed the same pattern outside of code.
The finding undercuts a common assumption: that a model good at filling in blanks will be good at editing. The real bottleneck is location selection. Miss a required spot and the change never happens; widen the net to be safe and the model starts needlessly regenerating code that should have been left alone. That is a useful distinction for anyone building AI coding tools, since it suggests infilling benchmarks alone overstate real-world editing reliability.
It is a reminder that a model which looks sharp on narrow fill-in-the-blank tasks can still make a mess once it has to decide, unprompted, what to touch.