A new paper argues that the wiring of a neural network's shortcut connections matters as much as how deep the network is.
Researchers behind the method, called ANCRe (adaptive neural connection reassignment), revisited residual connections - the skip-the-layer shortcuts that let deep networks train at all. Their analysis found that how those connections are laid out can create an exponential gap in how fast a model converges during training. ANCRe fixes this by learning the connection layout from the data itself, adding less than 1% in extra compute and memory overhead. The team tested it on large language model pretraining, diffusion models, and ResNets, and reported faster convergence and better performance across all three.
This matters because the field has mostly scaled depth by brute force - stack more layers, hope the gradient reaches them. Recent work has shown many of those deep layers barely get used, meaning some of that compute is effectively wasted. ANCRe reframes the problem as connectivity design rather than raw scale, a cheaper lever to pull than training ever-bigger models.
Cheap and effective is also exactly the pitch every optimization paper makes before anyone outside the authors' lab tries to reproduce it on a frontier-scale model.