How moving non-urgent AI tasks in time can help the power grid
This paper looks at whether AI data centers can help the electric grid by shifting the timing of non-urgent computing jobs. Modern AI centers use large numbers of GPUs and their power use can be large and jumpy. The authors propose a way to measure how much of that demand can be shifted without breaking the computing work.
The researchers focus on so-called batch workloads. These are offline jobs, like large model training or backups, that can wait a bit before they run. They first convert very fine-grained usage records (seconds-level CPU, GPU and memory traces) into coarser time blocks used by power systems (for example 5, 15, or 60 minute intervals). This “averaging-based” processing makes the computing data compatible with grid operation timescales.
They then build a scheduling model that can move the start times of batch jobs inside each job’s allowable window. The model keeps execution continuous (non-preemptive, meaning a job runs to completion once it starts), enforces delay limits, and respects server capacity for CPU, GPU and memory. A server power model that depends on utilization translates each scheduling plan into an electricity demand profile.
To judge flexibility, the paper uses two metrics: short-term peak demand shaving (how much a peak can be reduced) and the maximum duration of a sustained power reduction. Tests using real GPU cluster traces show that shifting batch workloads can create measurable, grid-compatible demand flexibility while causing limited disruption to the computing jobs.
There are important caveats. The method only measures flexibility from batch (delay-tolerant) jobs, not from delay-sensitive online services. The processing assumes each instance uses a fixed amount of resources while it runs and that schedulers follow a predefined algorithm. Results also depend on the chosen time resolution and on the accuracy of the server power model. The paper provides a framework and initial numerical evidence, but real deployments would need integration with data-center schedulers and grid operators and more study of practical constraints.