A new transformer model can retarget motion between wildly different skeletons without ever seeing matched examples of the two during training.
The approach, described in a new arXiv paper, uses a transformer autoencoder to learn a single latent space that ignores a skeleton's specific topology and position. Its key trick is a learnable flattening of the skeleton's graph structure, paired with graph-based positional encodings folded in multiplicatively rather than simply added to token content, as standard transformers do. The whole system trains unsupervised, with no paired retargeting data required. In zero-shot tests on skeleton types the model had never seen, it cut global joint position error by 43-47% compared to current benchmarks, and a 37-person user study that included professional animators ranked its output highest for motion alignment and physical plausibility.
That zero-shot result matters because transformer-based retargeting has had a rough track record: prior attempts have consistently lost to older, specialized geometric methods built for this exact problem. If the reported gains hold up, this is the first transformer architecture to actually beat those geometric baselines rather than just approach them, which matters for any game studio or animation pipeline juggling characters with mismatched rigs.
Still, this is one paper's benchmark numbers and a 37-person study, not a shipped tool. Whether it survives contact with the messy, inconsistent rigs animators actually use in production is a separate question.