A new decoding trick claims to make AI coding assistants actually follow the rules you give them, without retraining anything.
Researchers found that large language models get worse at following instructions as constraints pile up. The model does register what you asked for in its internal logits, but too weakly to reliably steer its output. IntentCoding fixes this by masking out the instructions during generation, comparing that against normal generation, and amplifying the difference with a multi-strength ensemble at each decoding step; no fine-tuning is required, and it drops into any existing model's decoding pipeline. The team also built a new benchmark, CodeConstraints, to measure compliance as the number of rules increases, testing the method on IFEvalCode, HumanEval, and LiveCodeBench.
Most fixes for instruction-following focus on better prompts or more training data; this one works purely at generation time, which is cheaper and needs no labeled examples or retraining. The reported gains are large: up to 71.0% relative improvement on CodeConstraints, 67.3% on IFEvalCode, and 29.3% in pass@1 on HumanEval and LiveCodeBench versus plain greedy decoding. If that holds up outside the authors' own benchmarks, it's a real step toward handing an AI coding tool an actual spec instead of a vague prompt.
Relative-improvement percentages without baseline accuracy numbers are an old trick in the research-paper playbook, and this is a fresh, not-yet-peer-reviewed arXiv preprint, so treat it as an interesting idea rather than a shipped feature.