A new tensor type promises to stop PyTorch from wasting compute on padding for oddly shaped batches.
Researchers built DanLing NestedTensor, an abstraction that tracks ragged shapes directly on the tensor instead of padding every input to match the largest one in a batch. Padding forces a GPU to run calculations on empty filler whenever sequences, images, or other inputs vary in size, and that waste multiplies fast in pairwise operations. The paper reports geometric-mean speedups of 2.74x in eager mode and 3.39x compiled on an A100 across four BERT model sizes, 1.97x eager across four FCN backbones, and a Pairformer-style workload running 2.40 to 4.32x faster while peak memory use dropped from 38.08 GiB to 5.41 GiB on its most uneven batch. The team says code will be released publicly once the paper is published, so none of this is installable or testable yet.
Variable-length data - protein structures, sentences of different lengths, images of different sizes - shows up constantly in deep learning, and the standard workaround has long been to pad and accept the wasted compute. If the reported speed and memory gains hold up once outside researchers can run their own tests, this could meaningfully cut training costs for groups working with irregular batches, especially in biology and NLP workloads where input sizes swing widely.
The benchmarks so far are self-reported and the code is not yet public, so this is a preview of a claim, not an independently verified result.