Operational cause of cross-cluster covariance

Identify the operational mechanism responsible for the persistent positive residual covariance among AI data-center clusters after accounting for calendar effects and capacity normalization.

Background

The paper finds that residual power fluctuations across clusters are positively correlated, causing portfolio aggregation to deliver less smoothing than an independent-cluster model predicts. Removing hour-of-day and weekday effects, followed by capacity normalization, explains only 34.4% of the excess covariance. Stable-capacity plateau analysis further shows that synchronized capacity expansion is not the dominant explanation.

Several possible causes—including common submissions, model releases, maintenance, or unobserved scheduler state—are mentioned elsewhere, but the available production trace cannot distinguish among them. Determining the operational source of this covariance is important for improving portfolio-level flexibility forecasts and avoiding optimistic independence assumptions in grid planning.

References

The stable-capacity test rules out expansion as the dominant covariance mechanism. Across six plateaus, raw cluster-to-fleet synchrony ranges from 1.645 to 1.900 and averages 1.765 (Fig.~\ref{fig:plateaus}). Capacity normalization within those same plateaus leaves a 1.567 average ratio. The result supports a portfolio correction while leaving its operational cause open.

Beyond Scalar Flexibility: From Eligible AI Workloads to Dependable Load Relief  (2609.05406 - Li, 4 Sep 2026) in Appendix, Section “Load and Aggregation Details,” discussion following Figure \ref{fig:plateaus}