A new research benchmark forces AI agents to survive weeks of real-world chaos, not just answer one tidy prompt.
Researchers introduced ReLiveGym, an evaluation environment that replays weeks of real news, financial markets, and social media activity in chronological order while an AI agent monitors the stream and acts only occasionally, the way it might on an ongoing market-analysis task. The tasks vary in how time-sensitive they are, how much reasoning they demand, and how often they recur. The team tested eight base language models to see how both the underlying model and the surrounding software harness shape performance over these long stretches. They also tested whether letting an agent learn from hindsight feedback on its own past actions improved results.
Most agent benchmarks drop a model into a frozen snapshot of the world and grade one response, which says little about agents meant to run unattended for days or weeks. ReLiveGym's headline finding is that knowing when to act is a bigger design problem than which model you pick - the best timing approach changed depending on the task and sometimes the model too. That is a pointed finding for anyone selling autonomous monitoring agents: the harness around a model may matter more than the model's name.
The code is on GitHub for anyone who wants to replay the chaos themselves, though this is a diagnostic tool, not proof that any agent is ready to run unsupervised for a month.