A new training trick teaches AI to describe time series data by testing whether another model can pick the right dataset from the caption alone.
Researchers built LineupRL, a reinforcement-learning pipeline for time series captioning that scores a generated caption not by asking a model to judge its quality, but by seeing if a frozen large language model can match that caption back to the correct time series out of several lookalikes. The verifier reads only raw numeric values, never a chart, and picks which series the caption describes. That task, matching, is easier for an off-the-shelf LLM to perform reliably than writing comprehension questions or rating caption quality directly. Across two captioning benchmarks and two downstream tasks, forecasting and reconstruction, where the predictor sees only the caption, LineupRL beat both supervised fine-tuning and prior RL approaches on every metric.
The standout result is efficiency: a 3B-parameter vision-language model trained with LineupRL outperformed a 72B model whose captions had been used to supervise the smaller baseline, using 1/24 the parameters. That is a meaningful data point in the broader argument that reward design, not just scale, determines how well captioning models generalize beyond their training data. It also suggests verifiable, identification-style rewards could transfer to other data types where judging quality is harder than confirming a match.
Researchers say the approach resists reward hacking and produces captions that track trends and name values at key points. Whether that holds once it leaves the benchmark, and someone tries to game it in production, remains to be seen.