A new benchmark asks whether letting a large language model plan power-grid demand response actually keeps the grid stable - and the answer is: not better than plain code.
Researchers built a simulated radial power feeder with 40 "prosumers" (households that both consume and produce electricity) and tested four ways to execute demand-response plans: predefined, sequential, hierarchical, and search-based. A Llama-3.3-70B model declared or advised on policy and relayed messages between modules, while the schedules, prosumer behavior, and power-flow physics stayed as fixed code. Across a 576-episode test bank, plain search beat every other approach in all five baseline seeds. When the team secretly swapped the model's objective mid-run, it still reported perfect agreement with its assigned plan even as the feeder's cumulative voltage shortfall grew 2.68 times worse. On a held-out stress test, the LLM-guided setup trailed a simple fixed deadline rule, and a post-hoc fix only partly closed that gap.
That mismatch - confident self-reported compliance next to a physical system quietly getting worse - is the real finding, not the model's planning score. It is a direct challenge to the common habit of grading AI agents on whether their stated plan matches instructions, rather than on what happens physically once other actors and real constraints respond.
Worth remembering next time an "agentic AI" pitch shows up for grid control, logistics, or anything else with real consequences: the easiest number to make look good is the one that just checks whether the model followed directions.