A new research framework called BabelCoder splits code translation across three specialized AI agents instead of relying on one model to do everything.
BabelCoder assigns one agent to generate translated code, a second to test it against the original program's behavior, and a third to repair whatever breaks. Researchers evaluated the system on four benchmark datasets against four existing state-of-the-art translation methods. BabelCoder won by 0.5 to 13.5 percent in 94 percent of test cases, reaching an average accuracy of 94.16 percent.
Code migration - moving an app from Python to Go, or modernizing a legacy system - is tedious, expensive, and easy to get subtly wrong. Splitting the job into generate, test, and repair roles mirrors how a careful engineer actually works, and the accuracy gain suggests structured decomposition, not just bigger models, is where translation quality is headed.
A 94 percent benchmark score is tidy, but these test suites are far cleaner than the sprawling, dependency-heavy codebases most companies actually want to migrate - so this reads as a promising lab result, not a tool you hand your legacy stack to yet.