AI agents can pool scattered information through trading, but only when the puzzle stays simple.
Researchers ran controlled prediction markets where large language model agents received private signals and traded against each other, testing four information structures that got progressively harder to reason through. The market aggregated information well in the easier setups, but accuracy fell apart once success required tracking more than two levels of what-do-others-know reasoning - a ceiling that lines up with prior results in human subjects. None of the usual fixes helped: letting agents chat before trading, giving them more time to trade, or prompting them to think strategically all left performance unchanged. A follow-up round of markets, run three months later with newer, more capable models, did better on the three easier structures, but on the hardest one they mostly swapped confidently wrong answers for cautious near-50/50 hedges.
That ceiling matters beyond academic curiosity. Prediction markets and multi-agent forecasting tools are increasingly pitched as a way to let AI systems pool distributed knowledge - internal corporate forecasting, research prioritization, even geopolitical risk modeling. This paper suggests that pitch holds up for simple aggregation tasks but breaks down exactly where the interesting problems live, in situations that require reasoning about what other agents know about what you know.
Smarter models scored better and traded more profitably, yet giving them feedback on past performance didn't budge their aggregation ability - a reminder that scaling up a model doesn't automatically scale up its theory of mind.