AI/ neural-networks · deep-learning-theory · optimization · arxiv

Study Finds a Simple Fix for a Stubborn Neural Network Flaw

Researchers show that adding a skip connection eliminates a class of training traps that otherwise survive no matter how large a neural network gets.

A new theoretical paper identifies a fix for a type of training flaw in simple neural networks that doesn't go away just by making the network bigger.

Researchers studied shallow, bias-free ReLU networks in a teacher-student setup, where a smaller teacher network's behavior is approximated by a student network. They found that for teacher networks with certain planar features, adding a learned linear skip connection removes all spurious local minima, as long as the student is at least as wide as the teacher. Without that skip connection, they constructed a specific three-neuron teacher network whose bad local minima persist in the student no matter how wide the student network gets, even at arbitrary overparameterization. The team also showed that student networks with positive output weights always end up learning within the subspace defined by the teacher's features, and that in two dimensions, even very wide student networks have an effective complexity capped by the teacher's width.

This matters because "just add more parameters" is a common fix for training problems in practice, and here is a documented case where more parameters alone doesn't help. The paper gives a concrete, provable reason why a small architectural tweak, a skip connection, succeeds where scale fails, which is relevant to anyone trying to understand why some networks train reliably and others don't.

Skip connections already power architectures like ResNets, but this work explains one specific reason they help, in a toy setting simple enough to prove things about. Don't expect this to change how anyone builds production models tomorrow, but it's a useful data point for the theory catching up to practice.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →