AI/ ai · machine-learning · transformers · research

A New Method Claims to Train Transformers Without Backprop

A new project claims it can pretrain transformers without backpropagation, though details and discussion remain thin so far.

A small research project called Dust says it can pretrain transformer models without using backpropagation, the algorithm that trains nearly every neural network in production today.

The project, published under the name Dust at qlabs.sh, lays out a method for pretraining transformers that skips backpropagation entirely. Backpropagation is the process that computes how much each weight in a network should change after every training step, and it has been the backbone of deep learning since the 1980s. The post surfaced on Hacker News on October 5, where it picked up a modest 33 points and exactly one comment - a sign the claim hasn't yet been put through its paces by the wider machine learning community.

Backpropagation works, but it's expensive: it requires storing activations across a model's full depth and computing gradients layer by layer, which is part of why training large transformers needs so much memory and compute. A credible alternative would matter because it could change the hardware and memory math of training, not just add a research footnote.

Backprop-free training has been promised before, including Geoffrey Hinton's forward-forward algorithm a few years back, and none of it has replaced standard training at scale yet. A method with one Hacker News comment hasn't cleared that bar either.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →