- The paper documents a preach–practice gap in coreset research: although efficiency motivates many methods, only six of ten report time savings and one reports energy savings.
- The paper estimates at least 46,179 ImageNet-1K training runs produced about 5.46 GWh of energy use and 2,430 tCO₂e from 2017–2025, excluding several lifecycle costs.
- The paper finds that pruning 25% of ImageNet-1K cuts training time by 24–40% and energy by 24–33% with little or no accuracy loss, while curated subsets can also reduce spurious bias.
This position paper argues that the ML community must close the gap between rhetorically endorsing data frugality and actually practising it. The authors—affiliated with the University of Copenhagen, the Technical University of Denmark, and Linköping University—support their position with three contributions: an audit of reporting practices in coreset research, indicative estimates of the energy and carbon costs of downstream ImageNet-1K use, and empirical evidence that coreset-based pruning reduces training time and energy with little accuracy loss while also mitigating dataset bias. They conclude with a stakeholder-differentiated call to action (2602.19789).
The preach–practice gap
The central diagnosis is a value-action gap: data reduction methods are frequently motivated by efficiency claims that are never empirically substantiated. In a representative sample of ten coreset methods applicable to deep learning, eight cite computational efficiency, three cite energy efficiency, and two cite storage as motivations, yet only six report time savings and only one evaluates energy savings. The authors note this omission is easily remedied with existing tools such as Carbontracker and CodeCarbon, which strengthens the claim that the gap is normative rather than technical. This finding implies that energy-related motivations in the coreset literature currently function as rhetoric rather than as evaluated claims.
Environmental cost of data
The paper situates data frugality within a lifecycle view that separates data from model lifecycles, arguing that environmental accounting has concentrated on model training and deployment while treating data as an effectively free input. Supporting evidence includes the BLOOM analysis, where non-training processes such as data processing and tokenization contributed 3.3 tCO₂e (5% of total emissions), and a Common Crawl/Tailpipe estimate of 326 kgCO₂e for crawling five billion web pages, 98% of it operational. The authors also flag attribution difficulties: network electricity is often overestimated by traffic-proportional models because network energy is driven by baseline capacity.
The paper's quantitative core is an estimate of aggregate downstream use of ImageNet-1K. Using OpenReview data on ICLR papers (2017–2022), LLM-assisted classification of how ImageNet was used, linear extrapolation to 2023–2025, and extrapolation of ICLR-derived training ratios onto a dimensions.ai keyword count, they estimate 46,179 training runs on ImageNet-1K between 2017 and 2025. Combining a measured 0.394 kWh/epoch (ResNet-50 on one NVIDIA A100, 300 epochs) yields roughly 5.46 GWh and approximately 2,430 tCO₂e at a global average carbon intensity of 445 gCO₂e/kWh—comparable to the annual carbon footprint of about 500 people. Storage of per-run local copies adds roughly 360 MWh/year.
Several assumptions warrant emphasis: the ICLR-derived training fraction is assumed representative of the wider literature; unpublished runs are excluded; only training and storage stages are counted; embodied emissions and pre-2017 usage are omitted. The authors state these are strict lower bounds, and Hugging Face Hub statistics support this: as of December 2025, 876 ImageNet-related datasets were hosted there (214 explicitly named imagenet-1k), with over 2.5 million downloads—a factor 55 above the estimated distinct training runs. The true footprint is therefore plausibly much larger than the headline figure.
Empirical gains from pruning
Drawing on prior SOTA results, the paper reports that Dyn-Unc prunes 25% of ImageNet-1K (and ImageNet-21K) with no drop in Top-1 accuracy, and InfoMax prunes up to 35% without performance loss; all coreset methods outperform random subsets at equal pruning ratios. Notably, these results span different architectures (Swin-T vs. ResNet-34) and are not compared against each other, so method rankings remain open.
The authors' own measurements quantify time and energy per epoch for full versus uniformly 25%-pruned ImageNet-1K:
| Model |
Time full → pruned (min/epoch) |
Time gain |
Energy full → pruned (kWh/epoch) |
Energy gain |
| ResNet-34 |
35.2 → 23.8 |
~32% |
0.2798 → 0.1989 |
~29% |
| ResNet-50 |
40.7 → 24.3 |
~40% |
0.3940 → 0.2645 |
~33% |
| Swin-T |
58.7 → 44.6 |
~24% |
0.7002 → 0.5300 |
~24% |
An important nuance is that a 25% dataset reduction does not yield proportional savings: observed gains range from 24–40%. The authors use this to argue against treating dataset size, wall-clock time, and energy as interchangeable proxies—a direct rebuttal to common efficiency claims. They also concede that coreset construction cost is neglected in these measurements, though it is amortizable as a one-time expense across reuses.
Bias mitigation via curated coresets
Beyond efficiency, the paper demonstrates on Colour-MNIST (with a 99% majority colour group) that balanced or reweighted coreset curation can recover "aligned" accuracy—performance not driven by the spurious colour feature—where random sampling under a conflicting bias does not. Gains grow with bias strength. This positions subset selection as an algorithmic remedy when the underlying collection cannot be corrected, though the demonstration remains a controlled toy setting rather than evidence on large-scale datasets.
Recommendations and counter-arguments
The call to action targets three levels: individuals (measure and report data frugality, adopt "Data-Pareto"/performance-per-data-point reporting, curate frugal datasets before release); platforms (mandatory resource reporting à la CVPR's compute form, resource-aware benchmarking with streamed canonical datasets, data-efficiency challenges such as BabyLM); and policy (standardized data-resource reporting, shared dataset infrastructure such as Sweden's Berzelius cluster, justification requirements and data sunset laws analogous to health-data governance).
The paper engages seriously with counter-positions. Scaling-law proponents argue continued expansion reliably drives progress and improves robustness to distribution shift and long-tail coverage; aggressive pruning may remove safety-critical rare samples. Rebound effects (Jevons paradox) may convert efficiency savings into more experiments rather than net reductions. And sustainability arguments can be misused to justify exclusion, a concern the authors explicitly reject, framing frugality as expanding participation. The paper concedes data frugality is neither universally applicable nor sufficient on its own.
Limitations and open questions
Beyond the conservative scope of the ImageNet estimate, the authors acknowledge that all experiments use image datasets, leaving generalization to tokenized or multimodal data unresolved. Methodologically, most coreset objectives target classification losses or gradient alignment and may discard rare modes and long-tail structure needed for generative modelling—an explicit open question given growing interest in condensed datasets for diffusion training. Coreset quality is sensitive to architecture, optimization settings, and augmentation choices, and the interaction between pruning and bias amplification at aggressive ratios is not fully characterized.
Conclusion
This position paper converts a widely voiced critique of data scaling into measurable terms: it documents systematic under-reporting of energy metrics in coreset research, quantifies a strict lower bound on the carbon cost of ImageNet-1K's downstream use, and shows that 25–35% pruning is achievable with negligible accuracy loss and 24–40% time savings. Its persuasive force rests on the argument that measurement is cheap and its absence therefore inexcusable. The main open problems it leaves are whether these findings transfer beyond image classification and whether efficiency gains can be prevented from being absorbed by increased experimental scale.