AI/ robotics · transformers · ai-research · rigid-body-mechanics

A Tiny Robot Policy Bakes In Physics Instead of Learning It

A 16,000-parameter transformer that encodes rigid-body physics directly beats networks 27 times its size on a standard robot manipulation benchmark.

A new transformer layer skips the trial-and-error part of teaching robots geometry and just gives them the physics upfront.

Researchers describe Screw Attention, a transformer layer that replaces the usual graph-edge relationship between tokens with an actual spatial transform. Each pair of tokens carries the relative pose between two bodies, and for robot joints, the joint's screw axis. Messages get transported into the receiving token's frame before any attention happens, while the attention scores themselves only see quantities that don't change when you shift reference frames. On the LIBERO-Spatial benchmark, a policy with just 16,000 parameters trained from object poses alone hit 97.3% success, beating graph networks, standard transformers, and flat networks of the same size, plus a flat network with 27 times more parameters. The researchers say ablations confirm the gain comes specifically from transporting the right relations, not from some other side effect.

That equivariance matters because most learned manipulation policies quietly memorize how a particular dataset labeled its coordinate frames. Change the convention and today's policies tend to fall apart, while this one, by construction, doesn't care. It also held up under pose noise and calibration errors about as well as a hand-built analytic controller, and improved a contact-rich insertion task when bolted on as a gated residual to one.

It's a reminder that bigger isn't always the answer. A 27x parameter penalty for skipping basic rigid-body mechanics is a steep price for a model to pay just to rediscover physics it could have been given for free.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →