A new paper puts PPO, one of the most widely used reinforcement learning algorithms, into a trading scenario where the correct strategy is already known from math.
Researchers built a continuous-time broker-trader simulation: a PPO agent plays the broker, choosing how fast to trade while an informed trader and random background order flow move the market. They derived the reward signal directly from the broker's known payoff equation and checked it carefully against the math. When there is no random background trading, PPO's neural-network policy learns to match the known-optimal strategy almost exactly. Add realistic randomness back in, and PPO, whether built on a simple feedforward network or an LSTM, gets noticeably worse, even though its network is capable of representing the right answer, confirmed separately via supervised learning. The researchers traced the failure to PPO's critic, the component that judges how good an action is, which could not reliably tell a slightly better trade from a slightly worse one. A separate, non-learning controller that uses only the broker's observable trading history stayed much closer to the correct answer under the same imperfect-information conditions.
This matters because it is a rare, controlled test of an AI method already used in real trading systems against a problem where correct is not a guess but a derived mathematical fact. PPO had an exactly correct reward and still could not reliably learn the right behavior once the environment got noisy, which says more about PPO's internal mechanics than about the environment. The researchers also tried a middle path: freeze the known-correct policy and let PPO only adjust it after trading costs changed. That closed just 2.22 percent of the gap to the new optimum, a small but repeatable gain.
If a textbook-perfect reward cannot get a standard RL algorithm to the textbook-perfect answer, that is a useful reality check for anyone pitching a trading desk on an AI that learned to trade.