A new benchmark called Forecast-Dojo puts language models to the test as forecasters, and the market still wins.
Researchers built Forecast-Dojo as a replayable environment combining 1,568 resolved Polymarket prediction-market events with 18.8 million dated news articles. The system lets an AI agent research a question, make a forecast, then revisit that forecast at later historical dates as new news arrives, without waiting for real-world events to actually resolve. The team evaluated 12 models this way. Giving those models research tools lowered their Brier score, a standard measure of forecast accuracy, across all 12, and predictions kept improving as more dated evidence accumulated.
Despite those gains, every model still trailed historical market forecasts on both accuracy and Brier score. That is a notable data point in the debate over whether AI agents can match crowds and markets at judgment calls, not just retrieval or coding tasks. The researchers also tested a persistent belief notebook carried between dates; it lowered research costs but did not reliably improve forecast quality.
Prediction markets have beaten pundits at forecasting for decades; this benchmark suggests they have not lost that edge to chatbots yet.