- The paper finds that two cooling subloops serving 14 and 11 CDUs minimize Frontier’s seven-year life cycle impacts, reducing emissions by 50.2 metric tons and cost by $100,000 versus the as-built design.
- The framework evaluates all 611 feasible CDU partitions by combining validated operational control, equipment and piping embodied-carbon models, capital costs, and reliability estimates across 49,353 operating timesteps.
- The results show that higher subloop counts provide little operational benefit because piping embodied carbon drives the optimum, while availability requirements can justify larger configurations such as Frontier’s Tier IV-oriented four-subloop design.
Motivation and problem statement
Liquid-cooled exascale supercomputers such as Frontier at Oak Ridge National Laboratory dissipate tens of megawatts of waste heat through cooling plants organized as parallel subloops serving coolant distribution units (CDUs). The number of subloops K and the assignment of CDUs to them are fixed at construction, yet prior work has optimized only operational control at fixed topology. The paper argues that as grids decarbonize, embodied carbon becomes a dominant share of facility life cycle emissions—some estimates place it at 38–69% of total facility carbon over a 60-year horizon—and that cooling plant topology is therefore a consequential, largely unstudied design decision. No published study had systematically enumerated feasible subloop configurations for an exascale system and evaluated each jointly on operational energy, embodied carbon, capital cost, and reliability.
Three-layer co-design framework
The framework integrates three physics models and a validated control optimizer into a life cycle objective. The operational layer reuses an SLSQP-based flow fraction and supply temperature optimizer from earlier Stage 2 work (Jadhav et al., 15 May 2026), built on a Modelica digital twin of the Frontier high-temperature water (HTW) system validated to ASHRAE Guideline 14 (CV-RMSE within 2.67%, NMBE within ±2.5%) (Jadhav et al., 1 Mar 2026). At each of 49,353 ten-minute timesteps from 2023 data, the optimizer minimizes pump plus cooling tower fan power subject to return temperature, heat exchanger duty, hydraulic head, and flow bounds fk∈[0.05,0.95]. The embodied layer builds a bill of materials per partition: piping sized per ASME B36.10M/B31.9 with ICE database carbon factors, pumps via CIBSE TM65 factors, plate heat exchangers from LMTD sizing, and Turton cost correlations escalated to 2024. The reliability layer uses IEEE Std 493 component MTBFs with an 8 hr MTTR, mapped to Uptime Institute tier downtime targets rather than solved fault trees.
Layer 1 enumerates all 611 integer partitions of the 25 CDUs into K∈{2,…,6}; Layer 2 runs the SLSQP optimizer for each partition; Layer 3 integrates energy, embodied carbon, and downtime into life cycle cost NPV (7 yr central horizon, 5% discount), life cycle CO₂e under the TVA grid trajectory, and expected unplanned downtime.
Headline result
The life cycle cost and carbon optimum for Frontier is K=2 with partition (14,11), achieving **3,320.7 t CO₂e and $3.99M over a seven-year horizon**, versus 3,370.9 t and $4.087M for the as-built K=4 configuration (7,6,6,6)—a saving of 50.2 t CO₂e (1.5%) and $100k (2.5%). The penalty grows monotonically withK:K=6f_k \in [0.05, 0.95]$0133k relative to the optimum.
The decomposition is the analytically important finding. Operational carbon is nearly invariant to $f_k \in [0.05, 0.95]$1 (a ~16 t spread across $f_k \in [0.05, 0.95]$2, about 0.8%), because the SLSQP controller extracts essentially all available flow-control flexibility even at two subloops. Equipment embodied carbon is exactly invariant at 1,134 t, since each pump and heat exchanger is sized to its subloop's share of the fixed total flow. The only $f_k \in [0.05, 0.95]$3-sensitive term is piping, whose mass scales roughly as $f_k \in [0.05, 0.95]$4 because more loops add primary loop length. The result therefore reverses the conventional assumption that more subloops yield better operational efficiency: the marginal flexibility of higher $f_k \in [0.05, 0.95]$5 is not recovered in operations, and the optimum is an embodied-carbon effect driven by piping.
Robustness
A one-at-a-time sensitivity sweep across five dimensions (horizon, grid trajectory, discount rate, material cost, embodied carbon factor) produced 15 scenarios; the $f_k \in [0.05, 0.95]$6 $f_k \in [0.05, 0.95]$7 partition remains on the Pareto front in all of them. Embodied carbon factors are the largest swing (2,767–3,830 t around the 3,321 t center). A two-parameter test varying embodied factor (±40%) and grid intensity (factor of two) jointly keeps the advantage of $f_k \in [0.05, 0.95]$8 over the best $f_k \in [0.05, 0.95]$9 partition between 9 and 25 t CO₂e everywhere. Additional stress tests show the ranking is invariant to the operational strategy used (Strategies A/B/C) and robust to adversarial CDU assignment, with worst-case penalties bounded by 13 t at $K \in \{2,\dots,6\}0(under0.3K \in \{2,\dots,6\}$1104k at 25 years.
Generalization and reliability framing
A semi-analytical decision rule, obtained by differentiating the life cycle objective with respect to $K \in \{2,\dots,6\}$2 at fixed partition shape, maps total HTW peak flow $K \in \{2,\dots,6\}$3 and drop budget $K \in \{2,\dots,6\}$4 to a predicted $K \in \{2,\dots,6\}$5. Evaluated on a 40×40 grid ($K \in \{2,\dots,6\}$6 kg/s, drop budget 25–600 m), it predicts $K \in \{2,\dots,6\}$7 for Aurora (760 kg/s peak flow), $K \in \{2,\dots,6\}$8 for both Frontier and El Capitan (~370 kg/s), and $K \in \{2,\dots,6\}$9 for LUMI (75 kg/s). This converts a site-specific study into portable guidance, though boundary locations depend on local grid intensity and cost parameters.
When expected unplanned downtime enters as a third objective, the Pareto structure changes qualitatively. Frontier admits Pareto-optimal $K=2$0 spanning Uptime Institute Tiers II–IV; Aurora collapses to a single point at $K=2$1 (with a 72-fold downtime advantage over $K=2$2); El Capitan mirrors Frontier; LUMI admits the widest set including $K=2$3. The paper's recommended framing is that reliability should be treated as a hard constraint set by operations policy rather than a soft traded objective: pick the minimum $K=2$4 meeting the availability tier, then optimize the partition there. Under this view, the deployed Frontier $K=2$5 configuration—with optimal constrained partition $K=2$6—is consistent with Tier IV intent rather than a flawed design.
Limitations
Several limitations bear directly on the results. The reliability axis uses tier-level availability targets assigned by redundancy configuration, not a solved per-partition fault tree with explicit standby modeling; common cause failures at the CEP level are excluded, though they plausibly affect all $K=2$7 equally. Embodied carbon factors combine EPDs and the ICE database; systematic upward revision would strengthen the $K=2$8 case, while downward revision shifts the optimum toward higher $K=2$9 without crossing to $(14, 11)$0 within the explored range. The analysis assumes the 2023 heat load distribution is representative of the full life cycle, and heterogeneous future workloads could enlarge assignment penalties. Finally, the decision map is calibrated to TVA grid intensity and US cost structures, so operators elsewhere must recalibrate coefficients even if the functional form generalizes.
Conclusion
This paper reframes cooling subloop count as a life cycle co-design optimization rather than a convention-inherited constraint. For Frontier-class systems the cost and carbon optimum is two subloops holding 14 and 11 CDUs, with savings driven entirely by piping embodied carbon rather than operational efficiency—an inversion of historical design priorities that survives every sensitivity scenario tested. The semi-analytical decision rule extends the result to other operating points, predicting distinct optima for Aurora, El Capitan, and LUMI, while the three-objective analysis shows that availability policy can legitimately override the unconstrained optimum. Open questions include a full per-partition fault tree with explicit standby modeling, extension to immersion and two-phase cooling modalities, dispatch-level marginal grid carbon integration, and closed-loop deployment of the decision rule during operation.