Hardware/ fp8 · fp64-emulation · gpu-performance · nvidia-b300

Study Finds Fp8 Trick for Fast Fp64 Math Has Limits

A new paper adds a missing cost term to a GPU math speedup model, showing memory bound workloads gain far less than claimed.

A new analysis pokes a hole in a GPU trick for faking 64 bit precision with cheap 8 bit hardware.

The trick, called Ozaki Scheme II, uses fp8 tensor cores to emulate fp64 math, and a prior paper modeled its speed as the higher of two costs, tensor core throughput and memory bandwidth, plus a small cost to reconstruct each output. This note points out a missing cost: every fp64 number streamed in has to be scaled, rounded, and split into several moduli before any multiply can happen. Adding that term and calibrating it against NVIDIA's cuBLAS emulation code yields a hard floor, roughly 0.56 FLOP per byte on a B300 GPU, below which the fp8 trick cannot beat plain fp64 no matter how fast the tensor cores are.

That floor lands squarely on the workloads the original paper highlighted. Memory bound operations like GEMV and SpMV, the kinds of math used throughout scientific computing and sparse linear algebra, only reach 0.3 to 0.9 times native fp64 speed, and a common stencil computation gets 1.8x instead of the claimed 3.1x. Dense matrix multiplication, which is not memory bound, is untouched.

It is a reminder that tensor core math tricks are great at the things tensor cores are great at, and nothing more. Precompute the residues ahead of time, the authors note, and the cost just moves to the memory bandwidth term instead of vanishing.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →