Trace of 155,410 GPUs shows AI data centers can shed only a few megawatts and less reliably over longer events
This paper studies how much power large AI clusters can realistically stop using when the grid asks for help. The authors reconstruct 4,439 hourly power observations from a public 185‑day trace of 155,410 GPUs across 17 clusters. They turn scheduler records into a planner‑facing estimate of demand and of the portion of that demand that is “eligible” to be curtailed — meaning work that could be paused while the GPU remains allocated and draws its idle floor.
To make power from the trace, the team built workload‑conditioned curves for training, online inference, offline inference and other work. They applied a modest data‑center conversion (including a 1.2 power‑usage effectiveness multiplier) and ran 300 model draws to expose parameter uncertainty. They checked anchors against 117 published measurements, 92 of which matched categories in the trace. The result is a duration–reliability–portfolio surface: estimates of how many megawatts can be removed for a given length of event and a given confidence level.
Key numbers are concrete and modest. The fleet’s median facility demand is about 55.8 MW. Immediate eligible curtailment — the power you could remove right away while keeping allocated GPUs powered at idle — averages 3.55 MW. That is 12.1% of the workload power but only 6.35% of median facility power. If every eligible watt could actually be realized, the amount available with 95% reliability falls from 2.51 MW for one hour to 2.32 MW for four hours and to 1.95 MW for 24 hours. The authors note that if only a fraction q of eligible power is realizable, every number on the surface simply scales by q.
A major finding is that a single fixed percentage of load is a poor shorthand. Using a single mean‑calibrated scalar would overstate the one‑, four‑, and 24‑hour 95% products by 17%, 25%, and 47%. Calibrating a scalar to be conservative at four hours instead makes the one‑hour product 6% too low and the 24‑hour product 17% too high. Pooling clusters helps but has limits: aggregating 13 clusters raised four‑hour “firmness” from 0.38 to 0.66, but positive covariance across clusters prevents the full independence gain. The production scheduler also adds almost no extra dependable capacity: newly deferrable arrivals average only 0.008 MW and provide zero 95%‑available capacity.