AI/ ai · reinforcement-learning · infrastructure · bytedance

ByteDance Built a Faster Way to Move LLM Weights Across GPUs

TensorHub, deployed inside ByteDance, cuts GPU stall time up to 6.7x and speeds elastic weight updates up to 4.8x during reinforcement learning training.

ByteDance has built a storage trick that lets AI training clusters share GPU memory instead of copying it.

Researchers introduced TensorHub, a system built on an idea called Reference-Oriented Storage (ROS). Instead of physically storing copies of a large language model's weights, ROS tracks which GPU workers already hold a given version in memory and serves read requests straight from them. The team layered topology-aware data transfer, consistency checks for model-parallel setups, and fault tolerance on top of that idea to make it production-ready. In testing across three different reinforcement learning rollout patterns, TensorHub saturated the available RDMA network bandwidth with little extra engineering work.

Training large models with reinforcement learning means constantly pushing fresh weights to thousands of GPUs, and clusters often scale up or down mid-job - a logistics problem that can leave expensive hardware idle waiting for data. TensorHub's numbers suggest that bottleneck is real: it cut GPU stall time by up to 6.7x for standalone rollouts, sped up weight updates by up to 4.8x in elastic setups, and slashed cross-datacenter stall time by up to 19x.

It is already running in ByteDance's production RL training, which matters more than the benchmark chart - infrastructure papers from major labs tend to describe what already shipped, not what might someday.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →