Researchers built an AI portfolio manager that can tell the difference between a good decision and a good day in the market.
The new method, called MemTrial, targets a specific flaw in LLM agents that manage stock portfolios by learning from their own trading history. Normally these agents judge each past experience in memory by the outcome of decisions that used it - but that outcome is mostly just the market's overall move on a given day, shared by every decision regardless of which experience shaped it. So agents end up crediting memories for what the market did, not what they did, and often underperform a plain equal-weight buy-and-hold portfolio. MemTrial fixes this by drafting the same decision multiple times, with and without a given experience, under identical market conditions, then isolates each experience's real contribution using a game-theoretic credit score (a Banzhaf value) pooled across dates with a hierarchical Bayesian model, defaulting back to the equal-weight portfolio whenever the evidence is too thin to trust.
This targets a quieter failure mode than the usual LLM-agent hype: an agent that looks like it is learning while actually memorizing market noise and mistaking it for insight. On benchmarks including PortBench and InvestorBench, MemTrial capped its worst-case losses at 2.2% below the equal-weight baseline, versus 15-38% for the best prior experience-learning agent, and lifted average utility by 21.2% across five test settings.
In a field that has piled LLMs onto trading desks mostly on faith, a method whose main trick is defaulting to doing nothing when it isn't sure is the least flashy, most useful idea here.