A single blast of random noise, it turns out, is enough to tell you how badly compression will hurt a given layer of a language model.
Researchers tested the idea, called RAM, by firing one random Gaussian probe through each of 1,683 tensors pulled from a 35B mixture-of-experts model and a 9B dense model. The probe gives an unbiased read on how much error rounding will introduce, because that error turns out to be spectrally flat rather than concentrated in a few directions. A single probe pegs a tensor's sensitivity to within 4 to 7 percent of the true value; running twenty probes tightens that to 1.3 to 1.4 percent. RAM then feeds those scores into a knapsack solver that assigns each tensor one of six bit-widths to hit an exact byte budget, with guardrails to stop any tensor from being crushed down to a ruinous 2 bits.
The payoff is that this works with zero calibration data. Where methods like GPTQ and HAWQ-V2 need real activations run through the model to judge sensitivity, RAM's probe can score a tensor in isolation and still rank-correlate 0.81 to 0.83 with GPTQ's own real-activation objective once propagated through the network. Across seven model families from 8B to 122B parameters, it undercuts size-matched uniform 4-bit builds by 3.5 to 13.6 percent lower perplexity, and probing a 400B model takes nine minutes on one workstation.
That last number is the real headline: calibration-based quantization has always carried a data and compute tax, and this result suggests most of that tax was unnecessary. Whether random noise holds up as well on models the authors did not test is the open question.