A new analysis pokes a hole in a GPU trick for faking 64 bit precision with cheap 8 bit hardware.
The trick, called Ozaki Scheme II, uses fp8 tensor cores to emulate fp64 math, and a prior paper modeled its speed as the higher of two costs, tensor core throughput and memory bandwidth, plus a small cost to reconstruct each output. This note points out a missing cost: every fp64 number streamed in has to be scaled, rounded, and split into several moduli before any multiply can happen. Adding that term and calibrating it against NVIDIA's cuBLAS emulation code yields a hard floor, roughly 0.56 FLOP per byte on a B300 GPU, below which the fp8 trick cannot beat plain fp64 no matter how fast the tensor cores are.
That floor lands squarely on the workloads the original paper highlighted. Memory bound operations like GEMV and SpMV, the kinds of math used throughout scientific computing and sparse linear algebra, only reach 0.3 to 0.9 times native fp64 speed, and a common stencil computation gets 1.8x instead of the claimed 3.1x. Dense matrix multiplication, which is not memory bound, is untouched.
It is a reminder that tensor core math tricks are great at the things tensor cores are great at, and nothing more. Precompute the residues ahead of time, the authors note, and the cost just moves to the memory bandwidth term instead of vanishing.