Researchers have a new way to test whether AI agents can actually edit a working piece of software, not just build one from scratch: make them mod Minecraft and Terraria.
The approach comes from a paper introducing IGMWorld and IGMBench, a benchmark of 110 game-modding tasks and over 1,100 checks spanning Minecraft and Terraria. Tasks range from simple property tweaks to edits that touch entities, dynamics, or whole systems at once, a scale the researchers call intervention depth. The best-performing agent setup solved 78.2% of tasks outright and passed 94.8% of individual criteria, but performance dropped as edits required touching more interdependent parts of a game. Most failed edits still built and ran fine; they just didn't behave the way they were supposed to.
That distinction matters more than it sounds. The industry has spent two years hyping AI's ability to generate entire apps and games from a prompt, but generation is the easy mode: there's no existing system to accidentally break. Editing is what most real software work actually is, and this benchmark suggests current agents are shakier at surgical changes than at wholesale creation, especially once a change ripples through connected systems rather than staying contained.
The benchmark also found something marketing decks tend to skip: even when an edit worked as coded, visual consistency held up less than half the time. A mod that runs correctly but looks broken is still broken.