A new algorithm speeds up the matrix multiplication at the heart of AI models without the manual tuning vendor libraries usually require.
Researchers built a system called SFC-CA GEMM. GEMM stands for General Matrix Multiplication, the dense math operation that underlies deep learning and high-performance computing. SFC stands for space-filling curves and CA stands for communication-avoiding - the paper uses space-filling curves to split matrix work into locality-preserving chunks automatically, then adds communication-avoiding techniques so the same scheme stays efficient whether a matrix is square or rectangular. Tested on four x86 and Arm chips, it beat vendor-tuned libraries by up to 5.5x on individual matrix shapes and 1.8x on average (weighted harmonic mean) across a mix of shapes. Plugged into real workloads, it sped up the prefill stage of LLM inference by up to 1.85x and distributed-memory matrix multiplication by up to 2.3x compared with the leading framework in each case.
Chipmakers hand-tune their math libraries for each new hardware generation and matrix shape, a process that does not scale and leaves gaps whenever new silicon ships faster than the tuning work does. A scheme that works across platforms and shapes without that manual effort - matching or beating vendor-tuned code - would matter most to smaller chip vendors and HPC teams that cannot staff large tuning operations the way bigger players can.
It is one arXiv paper tested on CPU architectures, not a shipped library - the real test is whether the approach holds up on GPUs, where most AI training spend actually happens.