Papers
Topics
Authors
Recent
Search
2000 character limit reached

EasyC: HPC Carbon Footprint Modeling

Updated 12 July 2026
  • The paper introduces EasyC as a lightweight yet sophisticated tool that models the carbon footprint of HPC systems using only a few key metrics.
  • It estimates operational and embodied emissions by leveraging publicly available data, statistical interpolation, and parameterized assumptions.
  • Applied to the Top500 list, EasyC quantifies 1,393.7 million MT CO2e operational and 1,881.8 million MT CO2e embodied emissions, supporting comparative and decision-making analyses.

Searching arXiv for the specified EasyC paper and closely related HPC carbon-accounting work. EasyC is a carbon-footprint modeling tool for high-performance computing systems that estimates both operational emissions and embodied emissions from a small set of high-level system metrics rather than from exhaustive GHG Protocol inventories. In "Modeling the Carbon Footprint of HPC: The Top 500 and EasyC," it is presented as the key enabler for Top500-scale carbon accounting under severe data scarcity, allowing the operational carbon of 391 HPC systems and the embodied carbon of 283 HPC systems to be modeled directly from disclosed data, with broader coverage obtained through public information and interpolation (Rao et al., 16 Sep 2025). The tool is explicitly designed for individual systems and collections of systems, and the paper uses it to derive the first full Top 500 carbon-footprint estimates: 1,393.71{,}393.7 million MT CO2e operational carbon for 1 year and 1,881.81{,}881.8 million MT CO2e embodied carbon (Rao et al., 16 Sep 2025).

1. Definition, purpose, and problem setting

EasyC is defined as a lightweight but technically sophisticated modeling tool for estimating operational carbon emissions, described as energy use over time and Scope 2 type, and embodied carbon emissions, described as manufacture and deployment of hardware and Scope 3 type. Its design target is broad, reproducible carbon accounting for HPC portfolios such as the Top500, national infrastructures, and research-computing sites, while keeping effort per system to less than $1$ person-hour per year (Rao et al., 16 Sep 2025).

The problem it addresses is the mismatch between conventional carbon-accounting methodologies and the data realities of HPC. The paper states that GHG Protocol and detailed LCA workflows require detailed inventories of CPUs, GPUs, memory, storage, networking, and cooling; exact embodied-carbon factors and supply-chain attributes; and time-series operational data. For a single large system, this can take weeks. For a collection of hundreds of systems, it is effectively infeasible. The paper further notes that there is no HPC-wide carbon reporting, that even the largest HPC sites do not do GHG Protocol reporting, and that none of the top 25 systems in the Top500 report their carbon footprint (Rao et al., 16 Sep 2025).

This establishes EasyC as a "few-metrics" modeling approach rather than a replacement for full GHG-compliant accounting. A common misconception is that such a model is intended to deliver exact inventory-grade values. The paper instead positions it as accurate enough for system-level and portfolio-level decisions, comparisons, annual reporting, and projections under constrained data availability (Rao et al., 16 Sep 2025).

2. Operational and embodied modeling framework

For operational carbon, EasyC uses a standard energy-to-CO2e pipeline with a compact input set: average power consumption PavgP_{avg}, operating time tt, optional power usage effectiveness, grid emission factor EFgridEF_{grid}, and optional utilization factor uu. The core relationships are

Eit=Pavg×tE_{it} = P_{avg} \times t

Etotal=Eit×PUEE_{total} = E_{it} \times PUE

Eop=Etotal×uE_{op} = E_{total} \times u

1,881.81{,}881.80

where 1,881.81{,}881.81 is IT energy, 1,881.81{,}881.82 includes facility overhead, 1,881.81{,}881.83 is operational energy actually consumed, and 1,881.81{,}881.84 is operational carbon emissions. In the Top500 setting, time is typically standardized to 1,881.81{,}881.85 year, grid emission factors use region-average ACI based on system location, and default PUE and utilization values are used when site-specific values are absent (Rao et al., 16 Sep 2025).

For embodied carbon, EasyC adopts a component-level statistical model that avoids exhaustive part enumeration. Instead of requiring every small part, it focuses on macro-configuration: number of compute nodes, number of CPUs, number of GPUs or accelerators, memory capacity, memory type, SSD capacity, and operation year or hardware generation. The paper describes internal models based on typical embodied carbon per CPU, per GPU, per TB of DRAM, and per TB of SSD, calibrated from published LCAs and industry data. Novel accelerators such as MI300A, A64FX, and SW26010 are approximately mapped to mainstream GPU or CPU carbon intensities when specific data are missing, with the explicit caveat that this can under-estimate embodied silicon carbon for very large dies (Rao et al., 16 Sep 2025).

A central design point is that EasyC needs just seven key data metrics, as contrasted with the hundreds of inputs required by detailed GHG-based methods. The paper also emphasizes that key assumptions are parameterized: system lifetime, PUE, utilization, location-based grid emission factor, and component intensities can all be tuned as improved information becomes available (Rao et al., 16 Sep 2025).

3. Data requirements, missingness, and coverage enhancement

EasyC is designed around public or easily obtainable information. From Top500.org it uses rank, system name, location, performance metrics such as 1,881.81{,}881.86 and 1,881.81{,}881.87, total cores, processor model, accelerator model where listed, and the power field when reported. Additional public sources include facility and system web pages, procurement specifications, press releases, vendor datasheets, media reports, technical blogs, and public configuration documentation, especially for open-science systems (Rao et al., 16 Sep 2025).

The paper’s coverage analysis shows why this design is necessary. Using Top500.org alone, the number of compute nodes and GPUs is missing for 209 systems, memory capacity and type are missing for virtually all systems, SSD capacity is missing for all 500 systems, and utilization and annual power consumed are missing for essentially all systems. With public information added, missing counts drop sharply, which increases direct modeling coverage from about 57% to almost full for operational emissions and from about 56% to about 81% for embodied emissions (Rao et al., 16 Sep 2025).

When systems still lack enough information after public-data enhancement, the paper introduces a rank-based interpolation method. For each missing system, EasyC takes the five systems just above and the five just below in Top500 rank space, computes the average carbon footprint of these peers, and assigns this average to the missing system. If some peers are also missing data, the neighborhood expands outward until a complete peer set is obtained. This procedure is applied only after coverage has been maximized through public information, on the assumption that systems near one another in rank are similar enough in performance scale and hardware generation to yield reasonable first-order estimates (Rao et al., 16 Sep 2025).

4. Application to the Top 500

Applied to the November 2024 Top500 list, EasyC models operational carbon for 391 of 500 systems and embodied carbon for 283 of 500 systems using only Top500.org data. With additional public information, operational coverage increases to 98% of systems and embodied coverage increases to 80.8% of systems, described in the paper as a 1,881.81{,}881.88 improvement over the embodied baseline. Interpolation is then used to obtain full Top500 coverage for both categories (Rao et al., 16 Sep 2025).

The resulting aggregate estimates are reported as

1,881.81{,}881.89

for total operational carbon over 1 year, and

$1$0

for total embodied carbon. Without interpolation, the paper reports $1$1 million MT CO2e operational carbon across 490 systems and $1$2 million MT CO2e embodied carbon across 404 systems. Interpolation therefore increases the operational estimate by only $1$3, but increases the embodied estimate by $1$4, which the paper attributes to originally missing embodied estimates for many large, accelerator-heavy systems (Rao et al., 16 Sep 2025).

The paper also emphasizes that systems adjacent in rank can nevertheless have large carbon differences. It cites a $1$5 operational CO2e difference between LUMI and Leonardo due to grid ACI, and a $1$6 embodied-carbon difference between Frontier and El Capitan due to accelerator and storage configuration differences. This suggests that rank is only a proxy for similarity and that location and architectural composition remain dominant explanatory variables even among ostensibly comparable systems (Rao et al., 16 Sep 2025).

5. Projections through 2030 and performance per carbon

Using observed Top500 turnover over two years, the paper reports an average of 48 new systems added to each list cycle. The inferred net effect is an operational-carbon increase of about $1$7 per cycle, equivalent to $1$8 per year, and an embodied-carbon increase of about $1$9 per cycle, equivalent to PavgP_{avg}0 per year (Rao et al., 16 Sep 2025).

The projection model is stated as

PavgP_{avg}1

and

PavgP_{avg}2

Under this parameterization, by 2030 operational carbon is projected to reach about PavgP_{avg}3 its 2024 level, while embodied carbon reaches about PavgP_{avg}4 its 2024 level. The interpretation offered in the paper is that overall computational performance continues to grow, post-Dennard efficiency improvements have slowed, and accelerator-rich architectures are expanding rapidly (Rao et al., 16 Sep 2025).

The paper also examines performance per unit carbon, or perf/CO2e, and reports that it increases only slowly, by PavgP_{avg}5 PFlop/s per million MT CO2e per year. This is stated to be far slower than historic Moore and Dennard scaling expectations. The consequence is that even if performance per carbon improves, total emissions can continue to rise because aggregate use of computing grows faster than perf/CO2e (Rao et al., 16 Sep 2025).

For an HPC center, the practical workflow described for EasyC is straightforward: gather the seven key metrics for each system, feed them into the tool, obtain per-system annual operational emissions, per-system embodied emissions, aggregated site or portfolio totals, and derived ratios such as perf/CO2e, then update annually as power data, utilization, and configurations change (Rao et al., 16 Sep 2025). This operationalizes system-level reporting with low effort and enables comparisons across clusters for procurement and upgrade decisions.

The paper is explicit about limitations. Operational results depend on location-based grid emission factors and can be skewed if those factors are wrong or outdated. Embodied intensities are drawn from generic LCAs, so unique accelerators and fabs may differ substantially. The approximation of novel accelerators with mainstream GPU intensities tends to under-estimate embodied silicon carbon. Rank-based interpolation assumes similarity that may not hold for all systems. Large storage systems and network infrastructure may be undercounted because the study focuses mainly on compute nodes. These are limitations of both input availability and model scope, not merely of implementation (Rao et al., 16 Sep 2025).

Validation remains partial because GHG-compliant reports for Top500 systems are largely absent. The paper therefore compares EasyC’s embodied estimates for Frontier, LUMI, and Perlmutter against detailed results from Li et al.; excluding storage to align scope, the reported embodied-carbon errors are in the PavgP_{avg}6 to PavgP_{avg}7 range. The authors argue that detailed methodologies also carry large inherent uncertainties and therefore regard EasyC as reasonable for large-scale assessments and comparative analysis, even if it does not achieve exact accounting fidelity (Rao et al., 16 Sep 2025).

EasyC should also be distinguished from EASEY, a separate framework for deploying containerized applications on future HPC systems. EASEY addresses middleware, container transformation, and job submission for exascale-oriented execution environments, whereas EasyC addresses carbon-footprint modeling; the similarity in naming can obscure that they solve different problems in the HPC stack (Höb et al., 2020). Within the carbon-accounting domain, the broader implication of EasyC is that it makes system-specific carbon analysis tractable enough to inform resource allocation, procurement, architectural choices, and possible changes to Top500 reporting practices, including better disclosure of actual power consumption and minimal hardware data aligned with the seven required metrics (Rao et al., 16 Sep 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to EasyC.