The AI Power Crunch
Data center energy demand is shooting constantly upwards. By 2030, some 945TWh of electricity is expected to be used to meet AI’s demand, which the International Energy Agency notes is roughly equivalent to the electricity usage of Japan today. The energy crunch is inescapable, and the industry’s answer has largely been to focus on the physical side: build more efficient processors and secure more power supply from an already tight grid. However, researchers suggest that optimizing the software stack could yield immediate power savings.
Average power usage effectiveness has barely changed for six successive years, according to industry surveys. While power usage metrics show whether cooling and power systems are wasteful, they do not account for whether the software running on servers is doing useful work. Given that servers account for around 60% of electricity demand in a modern data center, focusing on software efficiency offers significant potential for improvement.
Jae-Won Chung, a computer science researcher at the University of Michigan and a member of the ML.Energy initiative, confirms that software and algorithms can meaningfully help reduce energy consumption.
Computing systems can be viewed as a stack with hardware at the base, supported by systems software, algorithms, and applications. While hardware is slow and costly to replace, the upper three layers are easier to modify, and efficiency gains made there can compound rapidly.
Algorithmic Efficiency
Tests conducted by ML.Energy on large language models indicate that running inference in FP8, a lower-precision format, consumes significantly less energy than bfloat16 versions on problem-solving tasks. Training optimizers can also identify parts of a large-model training job with lighter workloads and slow them down to finish alongside busier tasks, cutting training energy consumption without sacrificing throughput or altering the hardware.
Major hardware vendors are implementing similar strategies. For instance, advanced power profiles fine-tune everything from processor compute and memory frequencies to power limits and cache settings to fit specific workloads, allowing power-constrained facilities to run more hardware efficiently.
Workload Management
Industry experts note that significant energy waste often stems from legacy equipment supporting outdated code, presenting a clear target for quick efficiency gains.
Cloud instances can be right-sized, and frequently used AI prompts can be stored in cache rather than recomputed continuously. Additional efficiencies can be achieved by routing routine requests to smaller models while reserving larger models for complex tasks, alongside better batching and compilers to raise accelerator utilization.
Workload Shifting and Grid Interaction
Software can also alter when and where electricity is drawn. Researchers studying workload shifting across global data-center fleets suggest that the timing and location of power usage, as well as how facilities interact with the local power grid, are critical factors.
If batch jobs are not time-sensitive, they can be delayed until local electrical demand falls or routed to a region with spare capacity and lower-carbon energy sources. Day-ahead planning and real-time scheduling help data centers respond dynamically to grid signals while preserving performance guarantees.
Such proactive scheduling was less critical when hyperscalers maintained excess capacity and the grid was not a primary bottleneck, but surging AI demand has fundamentally altered operational parameters.
However, operational limits remain. Data sovereignty regulations can prevent cross-border workload migrations, and moving large datasets across networks introduces logistical hurdles for inference providers.
Conclusion
Given the scale of AI-driven data center demand, software is not a complete substitute for improved silicon and expanded electrical infrastructure. Nevertheless, it serves as a vital control layer that extracts maximum efficiency from expensive assets and helps redirect demand to regions with adequate capacity.
Experts also emphasize the relevance of Jevons' paradox, noting that efficiency gains do not automatically reduce total electricity use if every watt saved is immediately deployed to generate more output. Power remains the core bottleneck for modern AI infrastructure.




