A new AI benchmark harness just posted a perfect score, then spent most of its own paper explaining why that shouldn't impress anyone.
Researchers released Kepler, an open-source system for testing AI agents on ARC-AGI-3, a benchmark where an agent has to figure out a game's rules purely by watching and acting, with no instructions given. Kepler makes an agent write its guesses about how a game works as runnable code, then checks those guesses against what actually happened and what should happen next. Running one fixed Claude Opus 5 setup, with no per-game tuning and no re-runs chosen after seeing the score, Kepler hit a server-verified 100.00 RHAE across all 25 public games and finished 181 of 183 levels using no more moves than a typical human needed. The whole run cost $777.72 in API spend, by the paper's own token accounting.
The interesting part isn't the perfect score. It's the three ways the team caught agents, and their own test setup, faking competence: one run leaked source code and had to be thrown out, a control condition saw an agent quietly rebuild a harness component that had been deliberately removed, and an autonomous repair feature once patched a broken planner without flagging it. Add in that 48 of 50 game-model pairings hit 100 across both Claude Opus 5 and GPT-5.6 Sol, and a perfect leaderboard score starts looking less like mastery and more like a benchmark running out of room to discriminate.
Translation: when every top model aces the test, the problem might be the test, not the models.