Frugal Learning: Efficient, Resource-Aware ML
- Frugal learning is a design principle that builds ML systems optimized to minimize resources—data, computation, and energy—while maintaining high accuracy.
- It categorizes resource use into input, learning-process, and model frugality, emphasizing selective data acquisition, efficient training, and compact architectures.
- Empirical studies demonstrate that frugal methods can achieve significant savings (e.g., 27% reduction in data consumption) with corresponding accuracy improvements.
Searching arXiv for recent and foundational papers on frugal learning and closely related formulations. Frugal learning, often termed frugal machine learning (FML), denotes the design of learning systems that aim to achieve useful or high predictive performance while minimizing consumption of scarce resources such as labeled data, features, computation, memory, time, energy, communication bandwidth, and deployment cost. In the broad formulation developed for wearables, embedded systems, IoT, and edge AI, frugality is not treated as mere runtime optimization, but as an explicit accuracy–resource trade-off spanning data acquisition, model construction, and inference-time operation (Evchenko et al., 2021). A later taxonomy organizes the field into input frugality, learning process frugality, and model frugality, emphasizing that frugal design can intervene at different stages of the ML lifecycle rather than only through model compression or acceleration (Violos et al., 2 Jun 2025).
1. Conceptual scope and taxonomy
Two complementary formulations recur across the literature. One defines frugal learning as building “the most accurate possible models using the least amount of resources,” with particular attention to wearable and privacy-sensitive settings where battery capacity, RAM, and on-device computation are binding constraints (Evchenko et al., 2021). Another defines FML as the practice of designing ML systems that are “efficient, cost-effective, and resource-aware,” with the explicit goal of minimizing resource use during both training and inference (Violos et al., 2 Jun 2025).
Within this literature, three categories are consistently distinguished. Input frugality concerns reducing the amount or cost of inputs, including fewer labeled examples, fewer features, or smaller selected subsets of data. Learning-process frugality concerns reducing computation, memory, or communication during training, including efficient hyperparameter search, continual updates, and avoiding repeated optimization from scratch. Model frugality concerns the size and execution cost of the final predictor, including parameter count, memory footprint, latency, and energy use (Evchenko et al., 2021, Violos et al., 2 Jun 2025).
This tripartite view is important because domain-specific work operationalizes “frugality” in different ways. In educational prediction, the central resource is data-source usage, so frugality appears as selective acquisition of additional student information only when confidence is low (Gagaoua et al., 10 Jan 2025). In software analytics, the dominant cost is manual labeling effort, so frugality appears as semi-supervised tuning with only 2.5% labeled data (Tu et al., 2021). In federated learning, the main burden may be iterative communication and local optimization, motivating one-shot analytical aggregation over frozen features (Dopico-Castro et al., 13 Feb 2026). In edge hardware, frugality is framed jointly in terms of energy efficiency, bounded-state computation, and interpretability (Shafik et al., 2023).
A plausible implication is that frugal learning is best understood not as a single method class, but as a resource-centered design principle whose concrete instantiation depends on what the application treats as scarce.
2. Resource objectives and quantitative formulations
Several papers make the accuracy–resource trade-off explicit through quantitative criteria. A general frugality score was proposed for comparing algorithms across datasets:
where is predictive performance, is resource consumption, and controls the strength of frugality in ranking algorithms (Evchenko et al., 2021). In the reported experiments, is mainly AUC and is mainly total CPU time for training plus prediction; varying changes the preferred learner as resource pressure increases.
Other formulations tailor the metric to the resource bottleneck. In early student failure prediction, data consumption is defined as
where is the number of data sources used for student ; this directly measures total source usage across students (Gagaoua et al., 10 Jan 2025). In the same setting, temporal utility is quantified by earliness, stability, and the Earliness Stability Score (ESS):
0
1
2
These measures capture whether predictions become correct early and remain correct over time, rather than only whether final classification is accurate (Gagaoua et al., 10 Jan 2025).
Length-sensitive frugality appears in RLVR for mathematical reasoning, where verbosity is itself a resource burden. The paper introduces Efficiency-Adjusted Accuracy (EAA),
3
to evaluate correctness jointly with output length, while emphasizing that the method does not train with an explicit length penalty (Bounhar et al., 2 Nov 2025). In ecological macroeconomic modeling, frugality is tied to simple classifiers and tabular reinforcement learning that run in seconds to minutes on a laptop rather than large-scale optimization (Vrizzi et al., 1 Dec 2025).
The field therefore does not rely on one universal frugality metric. Instead, each formulation measures the relevant scarce quantity: labels, source usage, CPU time, output tokens, energy, communication rounds, or memory footprint. This suggests that frugality is a multi-objective notion whose operational definition is deployment-specific rather than purely algorithmic.
3. Core methodological patterns
A recurring pattern is selective use of expensive resources rather than universal minimization. This is clearest in the Frugal Early Prediction (FEP) model for student failure prediction. FEP starts from a primary source, estimates student-level predictive confidence using a dataset-dependent heuristic, and retrieves additional sources only for students whose initial prediction is insufficiently confident. In the reported experiment on OULAD course 2014J-CCC, only 323 out of 694 students required extra data, so FEP used 1.46n data units rather than 2n for systematic multi-source use, corresponding to a 27% reduction in data consumption, while also achieving an average 7.3% accuracy gain over traditional approaches (Gagaoua et al., 10 Jan 2025). The central point is that “more data for everyone” is not guaranteed to improve prediction.
A second pattern is small labeled seeds with semi-supervised structure. In software analytics, FRUGAL tunes an unsupervised learner family over three modes—CLA, CLA+ML, and CLAFI+ML—and over percentile thresholds 4, validating only on 2.5% labeled data. The method is motivated by “label famine” and reduces labeling effort by a factor of 40 relative to fully labeled baselines, while matching or outperforming state-of-the-art methods on the studied tasks (Tu et al., 2021). The paper explicitly argues that complex and expensive learners should be baselined against simpler and cheaper alternatives.
A third pattern is resource-aware architecture restriction. FedHENet freezes a pre-trained feature extractor, learns only a single output layer, and computes that layer analytically in a single encrypted aggregation round. This removes repeated local fine-tuning and hyperparameter search, yielding a federated system described as hyperparameter-free and reporting up to 70\% better energy efficiency than iterative baselines, together with competitive accuracy and superior stability under heterogeneity (Dopico-Castro et al., 13 Feb 2026). A related continual-learning example is the replay-free conditional VAE framework for incremental generative modeling, which relies on a multimodal latent space and null-space gradient projection while keeping parameter growth static or tightly controlled; it is reported to be at least an order of magnitude more memory-frugal than closely related work (Enescu et al., 28 May 2025).
A fourth pattern is frugal transfer and pseudo-label reuse. In camouflaged human detection, fine-tuning pre-trained camouflaged object detectors with only 30 target images on CPD1K is identified as a practical “sweet spot,” and optimized GSAM pseudo-labeling achieves 5, slightly above the supervised frugal baseline at 6, though still below the full-data model at 7 (Pijarowski et al., 2024). This is a case where frugality reduces annotation but inherits failure modes from the pseudo-label generator.
These patterns share a common logic: spend resources only where they change the learned decision function materially.
4. Active, cost-aware, and reinforcement-guided querying
A major strand of frugal learning is active learning under nonuniform query cost. In cost-aware model exploration, query cost is incorporated directly into acquisition by dividing uncertainty scores by cost when costs are known; however, the paper shows that purely frugal selection can be too conservative, because it may avoid expensive but informative regions near the decision boundary. To address this, it introduces the 8-frugal learner, which usually prefers low-cost samples but, with probability 9, deliberately explores costly samples. On the Breast Cancer Wisconsin dataset with synthetic heterogeneous costs, the 0-frugal learner reaches 99.3% accuracy while maintaining reduced cumulative cost, outperforming both the known-cost frugal learner and random sampling (Stillman et al., 2020).
The same tension appears in automated algorithm selection. There, the expensive labels are algorithm runtimes on training instances, and many evaluations end in timeout. The frugal algorithm selection framework combines pool-based active learning with timeout predictors and dynamic timeouts. Across six ASLib datasets, the main conclusion is that frugal methods often retain 100% of the predictive power of passive learning while using less than 60% of the total labeling cost, and in some cases as little as 10%. The strongest savings come from dynamic timeouts rather than from uncertainty-based querying, because uncertainty sampling ignores heterogeneous labeling cost (Kuş et al., 2024). This directly challenges the misconception that active learning is automatically frugal when it is purely uncertainty-driven.
Reinforcement learning has also been used to adapt query criteria over active-learning rounds. In satellite image change detection, the proposed method assigns each unlabeled sample a relevance weight 1 under constraints 2 and 3, minimizing an objective that combines representativity, diversity, ambiguity, and an entropy regularizer. A Q-learning controller then selects which subset of criteria to emphasize at each iteration, because diversity is most useful early while ambiguity becomes more important later. On the Jefferson dataset, which contains only 39 positive patch pairs out of 2,200, the RL-based strategy achieves the lowest final EER and the best area under the EER curve among the compared methods (Deschamps et al., 2022). A related framework for image classification uses stateless Q-learning to adapt the weights 4 over diversity, representativity, and uncertainty, outperforming random, uncertainty-only, and fixed-weight baselines on Object-DOTA and other remote-sensing tasks (Deschamps et al., 2022).
These results collectively indicate that frugality in querying is not equivalent to greedily minimizing immediate cost. The cited work repeatedly shows that selective exploration of expensive or diverse cases can reduce total cost downstream by improving the model more quickly.
5. Deployment contexts and application domains
The deployment settings that motivate frugal learning are unusually diverse, but they are unified by explicit constraints on resource use.
In wearables and embedded sensing, frugality is tied to on-device learning, privacy, RAM, and battery life. A large empirical study over 103 WEKA algorithms and 517 OpenML classification datasets shows that algorithm rankings change sharply as frugality weight 5 increases, with methods such as A1DE, Naïve Bayes variants, and boosted stumps becoming preferable under stronger resource constraints. Deployment experiments on an LG URBANE smartwatch further show that Random Forest drains the battery faster than simpler learners such as HyperPipes and Naïve Bayes (Evchenko et al., 2021).
In microedge AI hardware, frugality is pursued through Tsetlin-machine design based on finite-state automata rather than arithmetic-heavy neural accelerators. Energy–performance trade-offs are controlled by the number of clauses, automaton states, feedback threshold 6, learning sensitivity 7, and the precision of pseudorandom number generation. The paper reports, for example, that on Iris, 50 clauses at 8 can match the accuracy otherwise requiring 150 clauses at 9, and that an 8-bit LFSR maintains accuracy comparable to a 64-bit PCG baseline, while lower than 7 bits causes a sharp drop in accuracy (Shafik et al., 2023).
In federated and distributed learning, frugality often concerns communication rounds, local optimization, and adaptation under non-IID data. FedHENet replaces multi-round gradient exchange with one-shot analytical aggregation under CKKS homomorphic encryption (Dopico-Castro et al., 13 Feb 2026). A separate study on violence detection compares LoRA-tuned LLaVA-7B with a 65.8M-parameter personalized CNN3D in a 10-client non-IID federation. Both exceed 90% accuracy, but the CNN3D achieves the best balance of energy, calibration, and ROC AUC among the trained models, while the paper recommends a hybrid design in which lightweight CNNs handle routine screening and VLMs are activated selectively for cases needing richer contextual reasoning (Thuau et al., 20 Oct 2025).
In education, frugality appears both at the analytics level and at the pedagogical level. FEP is a frugal prediction system over multiple student data sources (Gagaoua et al., 10 Jan 2025). Separately, a low-cost self-training scheme based on optional multiple-choice homework, computer-based grading through Auto-Multiple-Choice, and a sigmoidal bonus weighting
0
with 1, is presented as an inexpensive, scalable mechanism for improving mathematical skills in large physics classes (Lippi, 2018). The reported data show that homework participants improved by +0.13 in normalized final grade, while non-participants changed by -0.02, and the average direct homework bonus itself was only about 0.03 (Lippi, 2018).
In combinatorial scientific domains, spectral 2-regularization over the Walsh–Hadamard representation is used to enable data-frugal learning of pseudo-Boolean functions from few labels. The paper proves statistically optimal 3-type guarantees under restricted secant or quadratic-growth conditions and reports improved generalization in biological and physical tasks under label scarcity (Aghazadeh et al., 2022).
A condensed summary of these recurring deployment patterns is useful:
| Frugality type | Representative mechanism | Example paper |
|---|---|---|
| Input frugality | Query only critical samples or use few labels | (Stillman et al., 2020, Deschamps et al., 2022, Tu et al., 2021) |
| Learning-process frugality | Avoid repeated optimization or expensive search | (Kuş et al., 2024, Dopico-Castro et al., 13 Feb 2026, Enescu et al., 28 May 2025) |
| Model frugality | Restrict model size, clauses, layers, or active components | (Evchenko et al., 2021, Shafik et al., 2023, Thuau et al., 20 Oct 2025) |
This breadth suggests that the field is less a narrow subdiscipline than a cross-cutting methodology for building resource-aware ML systems.
6. Misconceptions, limitations, and open problems
A recurrent misconception is that frugality is synonymous with using less of everything at all times. Multiple papers explicitly reject that view. In cost-aware active learning, minimizing immediate query cost alone degrades accuracy because the learner may never visit high-cost boundary regions (Stillman et al., 2020). In algorithm selection, uncertainty-based active learning alone is not enough when labeling costs are highly non-uniform; timeout-aware mechanisms are more consequential (Kuş et al., 2024). In RLVR, the problem is not merely excessive length but a training distribution that makes the model conflate “thinking longer” with “thinking better”; the proposed remedy is to retain and modestly up-weight moderately easy problems, not to impose an explicit penalty on length (Bounhar et al., 2 Nov 2025).
Another misconception is that more data or larger models are automatically superior. FEP shows that systematic multi-source use is not consistently better than a single-source baseline and can be wasteful (Gagaoua et al., 10 Jan 2025). In violence detection under federated constraints, a much smaller CNN3D can outperform LoRA-tuned VLMs on ROC AUC and log loss while using less energy (Thuau et al., 20 Oct 2025). The wearable study similarly shows that high-accuracy ensembles become poor choices as runtime or battery limitations tighten (Evchenko et al., 2021).
The literature also states clear limitations. FRUGAL for software analytics relies on a monotonicity assumption embedded in CLA/CLAFI-style percentile logic and is validated only on two domains, with strongest results in intrinsically low-dimensional datasets (Tu et al., 2021). The self-supervised camouflaged human detector based on GSAM performs poorly on background-only images, with FPR = 0.680 and TNR = 0.319, indicating severe hallucination when no target is present (Pijarowski et al., 2024). FedHENet transmits 4 and 5 in plaintext and encrypts only 6, so its privacy story is specific to the sufficient statistic actually protected (Dopico-Castro et al., 13 Feb 2026). The ecological macroeconomics proof-of-concept uses a toy two-indicator model and explicitly notes that scalar Doughnut scoring is only an approximation to strong sustainability (Vrizzi et al., 1 Dec 2025). The spectral-regularization theory depends on nontrivial assumptions such as RSI or QG for the empirical error landscape (Aghazadeh et al., 2022).
At the survey level, the field still lacks standardized benchmarks that jointly measure accuracy, latency, memory, energy, and carbon footprint, and there is limited consensus on end-to-end deployment pipelines that combine data frugality, process frugality, compression, and hardware-aware execution (Violos et al., 2 Jun 2025). This suggests that a mature theory of frugal learning will likely require not only new algorithms, but also clearer evaluation protocols that treat resource use as a first-class outcome rather than an afterthought.