A new analysis explains why robots trained by imitation often settle on one 'safe' behavior instead of learning the multiple valid ways humans actually complete the same task.
Researchers examined two families of generative behavioral-cloning policies: latent-variable models like variational autoencoders (VAEs), and action-space generative models such as diffusion or flow-matching policies. Each has a different failure mode. Latent-variable policies need action-specific information baked into their latent codes to keep separate demonstrated behaviors distinct, but the common trick of forcing that latent space to match a simple prior distribution can wash the information out; loosen the constraint too much, though, and the policy risks depending on latent regions it was never trained to use at deployment. Action-space generative models hit a geometric wall instead: a smooth mapping from random noise to actions can't stretch cleanly across widely separated valid behaviors, so it either needs abrupt jumps in its input space or physically unrealistic in-between actions. Tests on a synthetic navigation task and a physical robot manipulation task with two valid solutions confirmed both failure patterns.
The more uncomfortable finding is about how we measure progress. Popular simulated robotics benchmarks turn out to have little genuine behavioral choice baked in - most tasks have close to one correct answer, so a plain deterministic model matches fancier generative ones on them.
Worth remembering next time a paper claims its diffusion policy or VAE-based controller "handles multimodality": check whether the benchmark ever gave the robot two good options to begin with.