AI/ ai · model-training · optimizers · hyperdimensional-computing

A New Training Trick Cuts AI Memory Use by Up to 23x

Phase-HDC trains compact AI classifiers via gradient thresholds instead of optimizer memory, cutting storage up to 23x for about five points of accuracy.

A new training method strips out the memory-hungry optimizer that most AI models carry during training.

Researchers built Phase-HDC for a type of compact classifier called a hyperdimensional model, where each learned parameter is just a low-bit angle - think of it as a dial with a handful of settings instead of a precise decimal number. Normally, training such a model with the Adam optimizer requires storing a running history of past gradients for every parameter, and that history can take up several times more memory than the model itself. Phase-HDC skips the history entirely: it nudges each angle one step opposite its current gradient, and only does so when that gradient is big enough to clear a threshold. The authors show this threshold rule is the exact answer to a simplified version of the training problem where every parameter change carries a fixed cost.

Across eleven datasets spanning images, tabular data, and text, Phase-HDC used 16 to 23 times less memory than standard 32-bit Adam and 4 to 6 times less than a lower-precision 8-bit version of Adam. That comes at a cost of roughly five accuracy points on average versus full-precision Adam - but Phase-HDC actually beat 8-bit Adam on six of the eleven datasets, including byte-level text prediction, where 8-bit Adam broke down completely.

That last result is the real finding here: for these discrete-grid models, the optimizer's memory was mostly deciding whether a parameter should move at all - a decision a simple gradient threshold can make for free. Compressing that memory into fewer bits, as 8-bit Adam does, apparently destroys the fine distinctions needed for that decision on sparser inputs like text.

Worth noting: the memory savings reported are logical state, meaning parameter counts, not measured reductions on actual hardware. And this only applies to hyperdimensional models with angle-only parameters, not the dense neural networks most AI teams actually train.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →