Researchers found that how you phrase a coding prompt can cut the energy your AI-generated code burns by more than half.
The study tested 21 prompting strategies for energy-efficient code generation, narrowed them to 8 effective ones, and ran them across 10 widely used open-weight and proprietary LLMs. Each model generated Python and C++ code with and without the optimized prompts, and the researchers measured the energy difference against a baseline prompt. Averaged across all 10 models, the best strategies cut energy use by up to 25% for Python and 17% for C++. The gains varied sharply by model: Python savings topped out at 50% for Granite-4.0-H-Small, 39% for Claude 4.5 Haiku, and 28% for MiniMax M3, while C++ savings ranged from just 7% for Qwen3-Coder-480B-A35B-Instruct up to 56% for Granite-4.0-H-Small.
Most code-generation benchmarks still grade LLMs purely on whether the output runs and passes tests. This study is a reminder that two functionally identical programs can have very different energy footprints, and that the prompt itself, not just the model, is a lever worth pulling as AI-written code becomes a bigger share of what runs in production.
The catch: a strategy that saves 56% on one model barely moves the needle on another. That is not a one-line fix you can paste into every prompt library. It is a research finding that still needs someone to turn it into a product.