A new research agent called Mimir used roughly half as much water as standard irrigation schedules - without ever touching a crop.
Researchers built Mimir, an LLM-based agent that manages irrigation decisions over full growing seasons across multiple sites, crops, and years. Rather than letting the language model act directly, Mimir routes every proposed watering decision through a deterministic physics simulator that checks, revises, and bounds the action before anything executes. A slower-running process separately tracks recurring mistakes and turns them into persistent rules that shape future proposals, while the underlying physics model and safety limits stay fixed. Tested against a common retrospective evaluator, Mimir produced the lowest aggregate control cost among the systems compared and used about 51% less irrigation than simply replaying historical schedules.
The interesting part isn't the water savings - it's what the ablation tests reveal about deploying LLMs in long-horizon physical control, where errors compound over months instead of resetting after one bad step. Strip out the forward simulation, the verified revision step, or the persistent memory, and performance gets measurably worse. Making the underlying LLM bigger, meanwhile, didn't reliably help at all.
That's a useful check on the assumption that scaling the model is always the lever worth pulling - sometimes the cheaper fix is simply not trusting the model to act unchecked.