AI's power problem might be mostly a software problem.
Data center demand is set to hit 945TWh of electricity by 2030, according to the International Energy Agency - roughly what Japan uses today. The industry's response has mostly been hardware: faster chips, tighter cooling, more grid contracts. But the Uptime Institute's 2025 survey found average power usage effectiveness has barely moved in six years, and servers account for roughly 60% of a data center's electricity draw versus 7-30% for cooling. Researchers argue the bigger lever is what's running on those servers. ML.Energy's Jae-Won Chung found that running Alibaba's Qwen 3 235B A22B Thinking model in FP8 instead of bfloat16 cut inference energy by a third, and his Perseus training optimizer trimmed training energy by up to 30% without slowing jobs down. Nvidia says its Blackwell power profiles save up to 15% of energy while keeping 97%+ of performance, letting operators cram more GPUs into the same power budget.
The more mundane savings may matter more. Chetan Visrolia, a data center infrastructure manager at SHI, said legacy equipment running old code is still the easiest fix available. "It is always the lowest-hanging fruit for quick savings," he said. "It can easily be found in any data center, as it is the loudest rack on the floor." Caching repeat prompts, routing simple requests to smaller models, and shifting non-urgent batch jobs to off-peak hours or regions with spare grid capacity all help too, with none of it requiring a single new chip.
None of this replaces better hardware or more power plants - it's a control layer, not a cure. And there's a catch: if every watt saved just gets spent generating more tokens, Jevons' paradox means total consumption keeps climbing anyway.