AI/ reinforcement-learning · procgen · benchmarking · ai-research

Study Finds RL Benchmark Scores Depend on a Hidden Baseline

Researchers show ProcGen generalization scores swing wildly based on test-time action choice and whether anyone bothered to measure a random policy baseline.

A new paper argues that reinforcement learning's favorite generalization benchmark has been grading itself against nothing.

Researchers tested PPO agents on eight ProcGen environments under a compute-limited budget of 8 million steps (25 million for three games), then compared standard train-versus-test generalization gaps against a new reference point: the score a uniform-random policy gets on the same levels. The random floor changes what the usual numbers mean. In the miner environment, a policy that samples its actions at test time scores 5.1 times the random floor on unseen levels, while the same policy's greedy (argmax) version scores below that floor in every run tested. Greedy evaluation also puts two of the eight environments significantly below random performance, and policy entropy, commonly used as a stand-in for how converged training is, turns out to be 32 to 66 percent inflated by actions that have identical effects on the environment.

That matters because the test-time action rule, sampling actions versus taking the most likely one, is usually treated as an implementation detail, not a variable worth reporting. An audit of twelve public ProcGen codebases found nine make that choice implicitly, with no explicit decision at the point where evaluation happens. If the rule and the baseline aren't fixed and reported, two labs can publish different generalization gaps on the same benchmark that aren't actually comparable.

Call it benchmark hygiene, not a new algorithm: the fix here is a checklist, not a breakthrough, which is probably why it took an audit of twelve codebases to notice nobody was doing it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →