A new study proves that a standard way of growing decision trees mid-training literally cannot learn anything new.
Researchers tested four ways that classifiers add structure as they train - adding a tree level, fitting a new hidden unit, splitting a leaf, or requiring statistical significance before a split - under one fixed protocol. They found that the most common way to deepen a soft decision tree, turning a leaf into a gate whose two children copy the parent's class distribution, leaves the gradient of every new gate at exactly zero. That means the added level cannot learn through gradient descent at all. Tested across three datasets over three seeds of five-fold cross-validation, trees built this way lost 19.6 accuracy points on Iris, 19.1 on Wine, and 55.6 on Digits compared with a tree of the same depth trained from scratch.
This isn't a minor inefficiency - it's a structural bug that would silently cripple any system using this growth trick, since nothing looks obviously wrong during training. The fix is almost embarrassingly simple: add a small random nudge to break the symmetry, after which gradients flow normally. The other three growth methods tested had real trade-offs rather than being broken: fitting a new unit to residual error shrinks networks without improving accuracy, splitting the highest-error leaf buys sparsity but loses 4.3 points on a harder problem, and gating splits behind statistical-significance tests buys nothing at all.
The paper's own suggested fix is a one-line sanity check - assert that new parameters have a nonzero gradient - cheap enough that its absence from standard practice is the more interesting finding here.