Make a neural network wide enough, and it barely needs backpropagation at all.
Researchers tested convolutional networks trained two different ways: standard end-to-end backpropagation, where error signals travel backward through every layer to coordinate updates, and greedy layer-wise training, where each layer learns on its own with no coordination from the rest of the network. Testing both approaches in a self-supervised learning setup, the team varied network depth and width. In relatively shallow but very wide networks, layer-wise training didn't just catch up to backpropagation, it outperformed it. A closer look at how each type of network organizes information internally showed that the wide, layer-wise-trained networks built a more favorable representational structure than the backpropagation-trained ones did.
That's notable because backpropagation's need for lockstep updates across layers is a real constraint on how AI training gets scaled and parallelized across hardware. If width can substitute for that coordinated signal, it points to training large self-supervised models in independent pieces, potentially spread across separate devices or schedules, without giving up performance. Most recent efficiency gains in AI training have come from better chips and better data pipelines; this is a reminder that architecture choices can do some of that work too.
The catch: this is a narrow experimental setup, wide and shallow convolutional networks, not a blueprint for training the next giant language model, and nobody's ripping backpropagation out of PyTorch on the strength of one paper.