An AI system now studies an entire TV series before it translates a single subtitle.
Researchers built SMART, a multi-agent system that builds a persistent memory of a show's terminology and tone, then translates a sample of sentences through a dynamic router and specialized agents that check terminology, verify subtitle timing, and pull context from earlier episodes. A judge agent scores the results and rewrites the other agents' instructions based on what went wrong, without retraining the underlying language models, before the system translates the rest of the series. The team tested it on Subtitle Arena, a new benchmark spanning 14 genres, shows from 1959 to 2023, and 15 target languages. SMART scored best across all 15 of those language pairs, cutting average translation errors by 6.9 percent compared with the next-best agent system, and also topped a separate public benchmark with a human-rated score of 4.50 out of 5.
Subtitle translation has a specific failure mode: sentence-by-sentence tools lose track of a character's nickname, a running joke, or a show's register over dozens of episodes. Letting a system build and refine its own playbook for a series before committing to a full translation addresses a problem that generic translation tools were never built to solve.
The catch: the benchmark is the researchers' own creation, and a system that improves itself by critiquing its own output can also learn to agree with itself.