Papers
Topics
Authors
Recent
Search
2000 character limit reached

Model Scaling Laws in ML

Updated 24 March 2026
  • Model scaling laws in machine learning are power-law equations linking model size, dataset size, and compute to test loss, enabling performance prediction at large scales.
  • They are derived through systematic log-linear regression on small models and validated across domains such as language, vision, and code.
  • These laws inform compute-optimal resource allocation and model design, though they may break down in data-limited or high-noise regimes.

Model scaling laws in machine learning formalize the empirical observation that as one increases model size, dataset size, or compute resources, loss typically decreases according to simple power-law relationships, subject to irreducible task and noise floors. These quantitative laws enable practical prediction of model performance at resource scales orders of magnitude beyond initial experiments, and underpin resource allocation in the development of large-scale neural networks, including state-of-the-art LLMs. This article provides a comprehensive survey of the mathematical forms, methods of estimation, empirical validation, theoretical foundations, domain-specific variants, and known breakdowns of scaling laws, primarily citing technical results from “Unraveling the Mystery of Scaling Laws: Part I” (Su et al., 2024) and supporting contemporary literature.

1. Canonical Power Laws: Formulation and Core Relations

Model scaling laws specify the asymptotic relationship between test (cross-entropy) loss LL and the key resource variables: model size (NN), dataset size (DD), optimization steps (SS), and compute (CC). Under fixed architecture, data distribution, and well-tuned hyperparameters, the empirically robust forms are:

  • Model-size scaling (infinite data/compute):

L(N)=(Nc/N)αNL(N) = (N_c / N)^{\alpha_N}

where NN = number of (non-embedding) parameters, αN\alpha_N = scaling exponent, NcN_c = characteristic scale.

  • Data-size scaling (infinite model/compute):

L(D)=(Dc/D)αDL(D) = (D_c / D)^{\alpha_D}

NN0 = dataset size (tokens), NN1 = data exponent (NN2 commonly), NN3 = characteristic token count.

  • Combined regime (infinite compute):

NN4

This form captures trade-offs when both NN5 and NN6 are varied (Su et al., 2024).

  • Compute-limited regime:

NN7

NN8: minimal steps to reach a given NN9 at infinite batch size, DD0 = step exponent, DD1 = step constant.

DD2

DD3 and DD4 tuned per experiment.

  • Finite-batch trajectory (implicit in DD5):

DD6

This formula predicts the full time/loss trajectory with arbitrary batch size.

All exponents DD7 and prefactors are determined via log-linear regression on experiments with small-scale models (DD8–DD9M parameters) and are then used to extrapolate to models up to SS0B parameters, as confirmed empirically (Su et al., 2024).

2. Practical Estimation of Scaling Parameters

Accurate scaling-law predictions hinge on systematic small-scale experimentation:

  • Estimating SS1:

Train SS2 decoder-only Transformers (SS3M–SS4M) to convergence on massive corpora (negligible data-limited effects), then fit SS5.

  • Estimating SS6:

For fixed SS7, use very large batch sizes to minimize step noise, record SS8 at many SS9, then fit CC0 vs. CC1.

  • Estimating CC2:

For fixed CC3, train short runs at various CC4, compute CC5 from CC6 contours of constant CC7, then fit CC8.

Example fitted values: | Corpus/Context | CC9 | L(N)=(Nc/N)αNL(N) = (N_c / N)^{\alpha_N}0 | L(N)=(Nc/N)αNL(N) = (N_c / N)^{\alpha_N}1 | L(N)=(Nc/N)αNL(N) = (N_c / N)^{\alpha_N}2 | L(N)=(Nc/N)αNL(N) = (N_c / N)^{\alpha_N}3 | L(N)=(Nc/N)αNL(N) = (N_c / N)^{\alpha_N}4 | |-----------------|:----------:|:-------------:|:----------:|:---------:|:----------:|:-----------:| | C4/1024 | 0.076 | L(N)=(Nc/N)αNL(N) = (N_c / N)^{\alpha_N}5 | 0.67 | L(N)=(Nc/N)αNL(N) = (N_c / N)^{\alpha_N}6 | 0.205 | L(N)=(Nc/N)αNL(N) = (N_c / N)^{\alpha_N}7 | | 3T-mix/4096 | 0.0615 | L(N)=(Nc/N)αNL(N) = (N_c / N)^{\alpha_N}8 | 0.672 | L(N)=(Nc/N)αNL(N) = (N_c / N)^{\alpha_N}9 | 0.139 | NN0 |

Constants depend sensitively on context length, tokenization, and data specifics (Su et al., 2024).

3. Empirical Validation and Cross-Domain Generality

Scaling laws have been robustly validated across language modeling, vision, code understanding, recommendation, acoustic modeling, and more.

  • Language and code models: Exponents are typically NN1 for cross-entropy loss, both in large-scale NLP (Su et al., 2024) and masked LLMs for code (Lin et al., 2024).
  • Vision/TinyML: For models below 20M parameters, exponents are substantially steeper, NN2 for error rate in ConvNets (Alnemari et al., 7 Mar 2026), but local exponents decay and saturate at scale.
  • Multi-output and kernel regression: Theoretical results confirm two-term power-law expansions: NN3, where NN4 reflects the spectrum of the data covariance; these predict larger exponents than found in large-scale deep models, supporting the universality-but-nonuniversality hypothesis (Chen et al., 3 Mar 2025).
  • Cases where scaling breaks: In small data or "critical size" regimes scaling laws break down; tasks with less than 10K examples, mismatch between pretrain and downstream tasks, or tasks dominated by irreducible noise often do not exhibit simple scaling (Ivgi et al., 2022, Alnemari et al., 7 Mar 2026).
  • Composition bias and architecture sensitivity: For NMT, scaling exponents differ notably for encoder and decoder, and composition bias in train/test data can dominate the scaling phase and even suppress BLEU improvements beyond a threshold (Ghorbani et al., 2021).

4. Theoretical Foundations: Mechanisms and Universality

Scaling-law phenomena originate from a combination of statistical and dynamical effects in high-dimensional learning:

  • Polynomial-spectrum mechanism: When the eigenvalues of the data covariance (or kernel) decay as a power law, optimal generalization yields excess risk falling off as NN5 with NN6, where NN7 quantifies target smoothness and NN8 is the redundancy index (Bi et al., 25 Sep 2025).
  • Random-feature and kernel regimes: Solvable models (random-feature ridge regression, NTK, field-theory dualities) display exact NN9 symmetry and identical scaling exponents for model and sample size, with breakdown or plateau when αN\alpha_N0 or αN\alpha_N1 approaches data intrinsic dimension (Maloney et al., 2022, Zhang, 2024).
  • SGD implicit regularization: In linear settings with power-law spectrum and Gaussian prior, one-pass SGD suppresses variance terms, yielding the empirical scaling law αN\alpha_N2, in sharp contrast to classical variance-limited bias-variance trade-offs (Lin et al., 2024).
  • Capacity and redundancy: The sharpness of the spectrum—i.e., degree of redundancy—directly modulates scaling exponents; flatter spectra (higher redundancy) slow down returns-to-scale, motivating investigations of representation learning and spectrum regularization to accelerate power-law decay (Bi et al., 25 Sep 2025).

5. Extensions: Compressed Models and Composite Laws

Unified scaling laws extend to models trained or deployed under compression (quantization, sparsity, etc.):

  • Unified capacity law: All compressed formats obey

αN\alpha_N3

where αN\alpha_N4 is the "dense-equivalent" capacity of format αN\alpha_N5, measured by the per-dimension Gaussian MSE of αN\alpha_N6 (Panferov et al., 2 Jun 2025).

  • Compositionality: For combinations (e.g., sparse+quantized), αN\alpha_N7 factors multiplicatively, greatly simplifying cross-format scaling prediction.
  • Empirical validation: For Llama-style Transformers, αN\alpha_N8, αN\alpha_N9 remain stable across INT, FP formats, and sparse/quantized hybrids. RMSE-injection predicts scaling-law fit parameter efficiency directly from format statistics without retraining (Panferov et al., 2 Jun 2025).

6. Methodological Best Practices for Scaling Law Estimation

Robust scaling law estimation requires precise methodology:

  • Data: Preferably train 5–10 small to moderate models spanning NcN_c0–NcN_c1 range beneath the target, holding architecture and training protocol constant (Choshen et al., 2024).
  • Checkpoint inclusion: Always include intermediate training checkpoints (excluding the first 10% of steps), substantially improving predictive accuracy.
  • Goodness-of-fit and extrapolation: Only trust extrapolations when the fit achieves NcN_c2; out-of-family extrapolation requires careful cross-family parameter transfer with fixed exponents.
  • Uncertainty quantification: Use at least 5 seeds per scale, hierarchical bootstrap (≥1000 resamples) for confidence intervals on scaling parameters and predictions.
  • Small-scale protocol efficiency: Training only small models to fit power laws can forecast attributes (loss, steps-to-loss, compute, optimal batch size) of 10B+ models in advance—achieving up to NcN_c3 compute savings in model design (Su et al., 2024, Choshen et al., 2024).

7. Implications, Limitations, and Open Directions

Scaling laws are central to pretraining strategy, compute budgeting, and architecture search for frontier models.

  • Compute-optimal allocation: For two-term power law NcN_c4 with fixed budget NcN_c5, optimal scaling is NcN_c6, aligning with the empirical "Chinchilla-optimal" prescription and spectral-theory predictions (Su et al., 2024, Lin et al., 2024).
  • Limits and breakdowns: When approaching irreducible task noise, as in acoustic modeling or in translationese-dominated NMT, returns from further scale diminish rapidly. Architecture, domain, training regime, and data distribution can all induce deviations, saturation, or even reversal (e.g., systematic error redistribution in TinyML regimes or plateau if source/data spectrum is exhausted) (Alnemari et al., 7 Mar 2026, Ghorbani et al., 2021).
  • Future work: Open technical fronts include inference-aware scaling laws (test-time compute, per-query adaptivity), multi-objective or fairness-aware laws, spectrum-specific and compositional architectures, and predictive scaling in safety-critical, data-limited, or highly non-i.i.d. settings (Sengupta et al., 17 Feb 2025).

Scaling laws provide a precise predictive lens but must be contextualized to architecture, domain, data spectrum, and operational constraints for optimal practical use.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Model Scaling Laws in Machine Learning.