Training large language models burns almost as much time moving data between chips as it does computing anything useful. A new technique called EDGC claims to claw back a chunk of that wasted time.
EDGC, short for entropy-driven dynamic gradient compression, adjusts how aggressively it compresses gradient data based on how predictable, or entropic, that data is at any given moment in training. It estimates entropy by sampling gradients, applies a theoretical model that ties entropy to a safe compression rank under a bounded-error constraint, and adjusts that rank in windows across pipeline stages rather than fixing it once. Researchers tested it on GPT2 models with 2.5 billion and 12.1 billion parameters, running on clusters of 32 V100s and 64 H100s. The result: up to 46.45% less communication latency and 16.13% faster end-to-end training, with no drop in model quality.
This matters because communication overhead, not raw compute, is often what slows down multi-GPU training runs. Most compression schemes pick one fixed rate and stick with it, which either wastes bandwidth when gradients are highly compressible or corrupts the model when they are not. Letting the compression rate track the gradients' actual entropy is a more surgical fix than the usual blunt-force approach.
Still, these are GPT2-scale tests on modest-by-frontier-lab-standards hardware, not the trillion-parameter runs where communication costs really hurt. Whether the gains survive that jump is the question nobody has answered yet.