A research project called Humanize bets that an AI coding agent is a bad judge of its own work, so it hands the decision to a different model entirely.
The system splits agentic coding into fixed roles: a human signs off on a plan contract, a builder agent implements it across multiple rounds, and a separate reviewer agent, sourced from a different vendor, decides when the work is actually done. Deterministic hooks, not another model, route tasks between these roles and enforce 72 mechanical checkpoints. Over 68 versions released in 108 days, the project collected 1,468 GitHub stars. Its track record includes a 567 file migration of the gem5 build system under upstream review, a kernel optimization variant that placed in the top three across all three Full Agent tracks of the MLSys 2026 FlashInfer contest, and an olympiad variant that posted perfect scores at IOI 2026, IMO 2026, IPhO 2026, and IBO 2024, plus 418.5 out of 437 at IChO 2026 and a first place finish on the Lean Eval math leaderboard.
The interesting part isn't the medal count, it's the design constraint: a bug only survives if two independently trained models both miss it, not just one. That's a structural hedge against the blind spots any single model has about its own code. But the project's own 118 postmortems undercut the pitch a little, finding that two thirds of review rounds happen after a change has already been marked accepted.
In other words, the reviewer catches bad claims, but nobody in this pipeline, human or machine, is reliably deciding when to stop.