A new technique lets 3D Gaussian avatars animate in real time without the expensive neural inference that normally slows them down.
Researchers introduced GALA (Gaussian Animation via Linear Approximation), a distillation method that swaps per-frame neural decoding for a shallow coefficient predictor and a linear blend of identity-independent blendshapes. The team builds the blendshape basis using block-local PCA under a rendering-aware metric and a fixed memory budget, then trains a small MLP to predict blend coefficients for each frame. The method attaches to existing avatar models without retraining them. The researchers tested it on three separate avatar architectures covering facial expressions and full-body animation with clothing dynamics, and across all three, GALA cut CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality, even for identities the models hadn't seen during training.
The real finding here isn't just speed - it's that avatar models trained separately, on different architectures, appear to converge on a similar linear internal structure. Practically, that difference is between an avatar that needs a GPU and one that runs at 60fps on a phone, which matters for anyone trying to ship these in AR glasses, video calls, or games instead of conference demos.
Real-time 3D avatars have been a research mainstay for years, usually chasing better rendering. This paper's bet is that the bottleneck was never the rendering - it was the animation math sitting in front of it.