A new trading algorithm claims to fix reinforcement learning's oldest finance problem: how to chase gains without getting wrecked in a selloff.
Researchers built PPO-HRAP, a system that pairs Proximal Policy Optimization with a regime-aware policy. The agent watches market data and its own portfolio state, then blends its learned trading action with a volatility-driven target exposure. Its reward function mixes portfolio log return with penalties for VIX-linked drawdown spikes, exposure drift, and excess trading. On a held-out SPY test window from 2020 to 2022, it returned 27.62% total and 8.48% annualized, with a Sharpe ratio of 0.64 and a max drawdown of 18.47%, well below the 34.10% drawdown a buy-and-hold strategy suffered over the same stretch. In single-run tests on QQQ and DIA, it also topped both total return and Sharpe ratio.
The result matters because the two standard approaches to RL trading both have obvious failure modes: reward purely on profit and the agent learns to just hold stocks through every crash, or penalize risk heavily and it turns overly timid during the volatile stretches where active management should earn its keep. PPO-HRAP's blended-action approach is a reasonably elegant middle path, and nearly halving drawdown while still beating a passive benchmark is a real result, not a rounding error.
Still, this is an arXiv preprint, not a live trading record. The authors admit turnover stays high, and the cross-asset wins on QQQ and DIA each come from a single run, not the five-seed stability test SPY got. Every backtested trading bot looks disciplined until it meets live slippage, fees, and a regime nobody tested for.