A new caching system called rMuscle speeds up the AI models that control factory robots, without changing how well they perform.
Researchers built rMuscle as an inference framework for Vision-Language-Action (VLA) models, the dominant approach for programming robots that do repetitive factory tasks. They noticed that when a robot repeats a job, not just its movements but the model's internal computations look similar from one run to the next. rMuscle exploits that redundancy with two caches: a Context Cache that reuses visual-token outputs, and an Action Cache that reuses neuron activation patterns to cut down on weight accesses. Tested on an RTX 4090 and a Jetson Thor chip across the LIBERO and RoboTwin benchmarks plus physical manipulation tasks, the system delivered a 1.29x to 1.42x speedup while keeping the robots' task success rates unchanged.
Inference latency is the hidden tax on embodied AI. A robot that thinks slowly moves jerkily, which matters on an assembly line where timing is everything. rMuscle's bet is that factory work, precisely because it's repetitive, is exactly the kind of workload that benefits from caching tricks borrowed from how human bodies offload routine motions to muscle memory rather than conscious thought.
It's a modest speedup, not a breakthrough, but in robotics modest and reliable usually beats flashy and fragile.