Scaling Laws: Definitions, Models, and Applications
- Scaling laws are quantitative relationships, often power laws, that show how an observable changes with scale and help identify structure across systems such as galaxies, texts, and neural networks.
- Their exponents often reflect underlying mechanisms, including constituent counts in particle physics, energy balances in materials, spectral decay in regression, and competition among skills in agent systems.
- Scaling laws are regime-dependent rather than universally valid, so researchers should test log-log fits, identify crossovers and saturation, and validate predictions on held-out data or new experiments.
Scaling laws are quantitative relationships describing how an observable changes under variation of scale, resource, geometry, or structural complexity. In their canonical form, they are power-law relations such as , for which dilation gives . On a double-logarithmic plot, , so the exponent is the slope. Scaling laws are used across particle physics, astrophysics, linguistics, machine learning, materials science, fluid dynamics, information retrieval, and agent systems because they can expose robust dynamical or statistical structure without requiring a complete solution of the underlying theory. Their validity is generally regime-dependent: crossovers, saturation, finite-size effects, spectral structure, or changes in the governing mechanism can produce multiple successive scaling laws rather than one universal relation.
1. Mathematical meaning and methodological role
A scaling law is commonly represented by a homogeneous power law,
where is the scaling exponent and supplies the dimensions required for consistency. It satisfies Euler’s homogeneity equation,
Under a dilation,
so the functional form remains invariant apart from an overall multiplicative factor. With logarithmic coordinates,
the relation becomes
0
The exponent is therefore the slope of a log-log plot, while the coefficient determines its vertical position.
The methodological significance of an exponent depends on the mechanism that generates it. In particle physics, an exponent may encode the number of active constituents, as in dimensional quark-counting rules. In a geometric variational problem, it may arise from balancing surface and elastic energies. In statistical learning, it may be controlled by a covariance-spectrum tail or an effective dimension. In language modeling, it may describe the tradeoff between model capacity, data, and privacy noise. A power law is therefore not merely a curve-fitting convenience when its variables and exponent have a theoretically motivated interpretation.
Scaling relations are usually approximate rather than exact. A single law may describe only an intermediate regime, with earlier or later behavior controlled by different mechanisms. “Radical scaling” emphasizes that two intersecting laws can provide objective characteristic units, while three connected laws can also define an objective dimensionless radix for representing regime crossovers (Fardin et al., 3 Jul 2025). The same work distinguishes dimensional units, nondimensional variables, and the choice of logarithmic base: the latter changes the representation of multiplicative separations without changing the underlying physical data.
2. Physical scaling laws in particle physics and astrophysics
Particle-physics scaling
Bjorken scaling emerged from deep-inelastic electron–proton scattering at SLAC in 1969. For
1
the structure functions were approximately independent of the absolute momentum-transfer scale when expressed in terms of the dimensionless Bjorken variable
2
where 3. Schematically,
4
Its interpretation was that the proton contains effectively point-like constituents—partons, later identified with quarks and gluons. Exact scaling is replaced in QCD by calculable 5 evolution due to gluon radiation and quark–gluon interactions, but the approximate experimental scaling was decisive evidence for physical proton substructure (Muradyan, 2011).
Dimensional quark-counting rules relate the asymptotic electromagnetic form factor of a composite object 6 to the number 7 of elementary constituents:
8
The constituent counts discussed for the pion, nucleon, deuteron, 9, and 0 are respectively 1, 2, 3, 4, and 5. Thus, in schematic form,
6
For exclusive reactions,
7
the corresponding cross-section law is
8
At fixed center-of-mass angle and large 9, the examples include
0
1
and
2
These exponents have a structural interpretation as counts of active elementary degrees of freedom. The rules are nevertheless restricted to hard-scattering regimes and can be modified by logarithmic QCD corrections, helicity-selection effects, endpoint configurations, higher-twist contributions, and soft rescattering.
Regge and astrophysical angular-momentum relations
The Chew–Frautschi relation places hadrons approximately on Regge trajectories,
3
Neglecting the intercept gives 4, or, in the normalization used in the paper,
5
The relation is associated with Regge phenomenology and string-like or rotating extended objects. It is not an exact universal trajectory: flavor families, parity sectors, radial excitations, and intercepts produce distinct trajectories.
An astrophysical generalization assigns an effective geometric dimension 6:
7
The cases are interpreted as string-like, disk-like, and ball-like objects:
8
The disk relation is associated with galaxies and the ball relation with stars and planets. The paper gives example angular momenta of approximately 9 for Andromeda and 0 for the Milky Way, using estimated masses including dark matter.
The extremal Kerr bound for a rotating black hole is
1
with dimensionless spin parameter
2
Using the Planck mass,
3
this becomes
4
The Kerr and hadronic relations are parallel on a log-log plot because both have exponent 5, but their coefficients differ by approximately 6. Parallel exponents do not establish a common mechanism: the Kerr relation follows from general-relativistic black-hole mechanics, whereas the hadronic relation is an approximate empirical relation for strong-interaction bound states.
3. Scaling in statistical systems and continuum mechanics
Linguistic scaling
Zipf’s law relates word frequency 7 and rank 8 through
9
or, for the distribution of types by absolute frequency,
0
A finite-text scaling formulation instead defines relative frequency
1
and assumes that rank depends on frequency and text length through their ratio:
2
Differentiation yields
3
where 4 is vocabulary size. Consequently, plotting 5 against 6 should collapse distributions from different text lengths onto a length-independent curve. The invariant is the shape in relative-frequency coordinates, whereas absolute crossover frequencies increase linearly with 7 (Font-Clos et al., 2013).
For lemmatized texts, the proposed scaling function is
8
It has a low-frequency regime 9 and a high-frequency regime 0. The high-frequency exponent is close to 1, corresponding to the conventional Zipf exponent 2. The crossover occurs at
3
so its absolute location grows linearly with text length while its relative location remains fixed.
The vocabulary relation follows from
4
For the double-power-law form, the continuous approximation gives
5
Thus the asymptotic growth is logarithmic rather than a pure Heaps law. The paper emphasizes that discreteness at low frequencies is essential: the discrete expression tracks empirical vocabulary growth better than the continuous approximation.
Elastic films with dislocations
A variational model for a two-dimensional epitaxial film includes surface energy, elastic misfit energy, and dislocation nucleation energy. The total energy is
6
Here 7 is surface tension, 8 is lattice mismatch, 9 is the Burgers-vector magnitude, 0 is the dislocation core radius, and 1 is the number of dislocations. Under the stated assumptions, the infimal energy obeys, up to multiplicative constants,
2
The first branch is the coherent or defect-free island scale. For an island of width 3,
4
and balancing the last two terms gives
5
The dislocation branch follows from a spacing scale
6
with approximately 7 dislocations in an island of width 8. Its energy is
9
leading to
0
The logarithm is the signature of two-dimensional dislocation self-energy. The lower bound uses a ball construction, Korn inequalities for fields with nonzero curl, and local estimates distinguishing elastic mismatch from defect-mediated relaxation (Abel et al., 2024).
Martensitic transformations and geometry
For two-dimensional martensitic phase transformations, the geometrically nonlinear energy is
1
while the linearized model is
2
If the boundary tangent or normal directions are incompatible with the martensitic wells in the sense of the Hadamard jump condition, the minimum energy has the generic scaling
3
For specially adapted polygonal domains with compatible geometry,
4
The logarithmic loss arises from a slice estimate of order 5 integrated across boundary distances:
6
Compatible polygonal domains support finite-scale, stress-free or nearly stress-free microstructures, whereas generic domains require increasingly fine branching toward incompatible boundaries (Ginster et al., 2024). The scaling law thus acts as a selection principle for domain geometry: wells and symmetry determine preferred polygonal morphologies.
Explosions and capillary phenomena
For explosions, three regimes are identified:
7
The first is finite-speed initial ejection, the second is the Taylor–Sedov blast regime, and the third is acoustic propagation. The characteristic crossovers are
8
The dimensionless initial Mach number is
9
Using the radix 0 and Hopkinson–Cranz variables gives the idealized master curve
1
The framework has been applied to nuclear, chemical, underwater, laser-induced, and vapor-cloud explosions.
For capillarity-driven processes, the proposed regimes are
2
representing viscous, inertio-capillary, and geometric-saturation regimes. Their crossover structure is controlled by an inverse-Ohnesorge-type quantity
3
The construction has been applied to drop and bubble spreading, coalescence, and pinch-off. For completely wetting spreading, the terminal plateau is only approximate because Tanner’s law gives a later 4 regime (Fardin et al., 3 Jul 2025).
4. Scaling laws in learning systems
Dense-model and pruning laws
For a fixed dataset size 5, model error is approximated by
6
while for fixed model size 7,
8
A joint approximation is
9
The terms represent data-limited error, model-limited error, and an asymptotic lower error. The thesis also introduces a rational complex-envelope transition to model the random-guess regime, intermediate power-law region, and asymptotic plateau. These forms are phenomenological approximations, with exponents fitted per task and scaling policy (Rosenfeld, 2021).
Training compute is approximated by
00
In the power-law regime, minimizing 01 at fixed error leads to the allocation condition
02
Thus model and data should be increased in a balanced manner rather than independently.
Iterative magnitude pruning exhibits three density regimes: a low-error plateau, an intermediate power-law region, and a high-error plateau. The intermediate behavior is
03
where 04 is retained density. Depth 05, width 06, and density can compensate through an approximate invariant
07
This invariant produces a joint pruning law in which architecture and sparsity preserve approximately constant error. Its validity is established for iterative magnitude pruning with weight rewinding and is not automatically transferable to structured pruning, quantization, distillation, or other compression procedures.
Linear, kernel, and quadratic regression
For sketched infinite-dimensional linear regression with covariance eigenvalues
08
one-pass SGD produces reducible error of the form
09
The first term is sketch-induced approximation error; the second is finite-training bias. SGD’s implicit regularization keeps variance lower order under the stated one-pass schedule, even though variance can increase with model size (Lin et al., 2024). Balancing the two terms under compute 10 gives, up to logarithmic factors,
11
Related scaling extends to multiple regression and kernel regression, where a Gaussian sketch dimension acts as effective model size and the covariance or kernel spectrum determines the exponent (Chen et al., 3 Mar 2025).
A quadratically parameterized linear model uses
12
Its SGD update contains a multiplicative factor 13, producing adaptive feature selection. Under covariance decay 14 and source decay 15, the dynamically learned effective dimension is
16
In the aligned regime 17, the sample-size exponent is
18
whereas in the opposing regime 19, quadratic SGD has exponent
20
compared with 21 for ordinary linear SGD (Ding et al., 13 Feb 2025).
Redundancy and spectral scaling
A kernel covariance spectrum
22
has effective dimension
23
For a target satisfying a source condition with smoothness 24, kernel ridge regression has bias–variance structure
25
Optimizing 26 gives
27
and
28
Here 29 is termed the redundancy index: a flatter spectrum has more statistically active directions and slower scaling. Boundedly invertible representations preserve the spectral tail and exponent, while mixtures are controlled asymptotically by the component with the smallest 30 (Bi et al., 25 Sep 2025).
Time-series forecasting
Time-series forecasting introduces look-back horizon 31 as an additional scaling variable. The loss decomposes into Bayesian error and approximation error,
32
Longer histories reduce truncation uncertainty, but increase intrinsic dimension 33 and reduce the effective number of independent sliding-window samples, approximately
34
In data-sufficient settings, model partitioning error decreases with complexity while estimation error increases with complexity. In data-scarce settings, nearest-neighbor behavior yields approximation error proportional to
35
The resulting tradeoff predicts an interior optimal horizon: more history is beneficial initially but can harm performance when added dimensions exceed the available data and model capacity. Empirical results on ETTh1, ETTh2, ETTm1, ETTm2, Exchange, Weather, ECL, and Traffic support dataset-size scaling, model-width scaling, saturation, and horizon optima (Shi et al., 2024).
Privacy-constrained and multilingual scaling
Differential privacy adds gradient noise as a new scaling dimension. The compute model is
36
where 37 is parameter count, 38 batch size, 39 sequence length, and 40 iterations. Utility is modeled as a surface
41
where 42 is the noise-batch ratio. Under moderate privacy, compute-optimal private training favors substantially smaller models, much larger batches, and more examples per parameter than non-private training. The paper reports optimal batches often in the range 43–44 and token-to-model ratios often between 45 and 46 for 47 (McKenna et al., 31 Jan 2025).
For multilingual LLMs, language-family loss is modeled as
48
where 49 is the family’s sampling ratio. The family-independence hypothesis states that a family’s loss depends primarily on its own sampling ratio, while within-family transfer remains important. Optimizing a weighted mixture gives
50
When 51 and 52 is approximately scale-invariant,
53
Experiments involving 23 languages and five families found that ratios optimized on 85-million-parameter models transferred effectively to models up to 1.2 billion parameters (He et al., 2024).
5. Scaling laws for ranking, retrieval, and agent systems
Rank-based language-model scaling
Relative-Based Probability is defined by the rank 54 of the correct token:
55
The proposed law is
56
where 57 is the number of non-embedding parameters. It complements cross-entropy, which measures the absolute probability assigned to the correct token. RBP instead measures whether the correct token is ranked first or within a top-58 candidate set. Across Pythia, GPT-2, OPT, and Qwen families, the reported fits have high 59, approximately 60 for 61 and at least 62 for many moderate-63 experiments. The law becomes unstable when 64 approaches vocabulary size because 65 saturates near one (Yue et al., 23 Oct 2025).
A sequence requiring 66 consecutive correct top-67 decisions has, under independence,
68
Thus a smooth token-level scaling law can produce a sharp sequence-level transition without requiring a discontinuous underlying capability transition.
Reranking
For cross-encoder rerankers, quality is modeled with saturating laws such as
69
70
and
71
Experiments across pointwise, pairwise, and listwise reranking use the Ettin series from 17 million to 1 billion parameters, with BM25 top-100 candidate sets from MS MARCO. NDCG@10 and MAP exhibit comparatively reliable scaling, whereas Contrastive Entropy is often unstable and MRR is dataset-dependent. Models up to 400 million parameters were used to forecast the NDCG of a 1-billion-parameter reranker with low reported RMSE (Seetharaman et al., 5 Mar 2026).
Agent skill-library scaling
In LLM-agent systems, the scaling variable can be the number of exposed skills rather than parameters or training data. The empirical routing law is
72
where 73 is the number of simultaneously exposed skills and 74 is the logarithmic decay slope. Across 15 frontier models and 1,141 skills, all reported fits have 75 over the tested range 76. Errors progress from local competition to cross-family drift and capture by overly general “black-hole skills.” Local semantic competition predicts errors better than global library size alone.
For multi-step route-only pipelines,
77
Before execution artifacts are realized, downstream routing is approximately multiplicative. After correct upstream execution, a concrete artifact can rescue downstream decisions:
78
with pooled estimate 79. A law-guided library intervention increased held-out routing accuracy from 80 to 81 and reduced in-library hijack from 82 to 83 (Chen et al., 15 May 2026).
Automated scaling-law discovery
EvoSLD formulates scaling-law discovery as grouped symbolic learning. Scaling variables 84 are varied within an experiment, while control variables 85 define groups with separate coefficients. The goal is a shared symbolic structure
86
with group-specific parameter values. The method co-evolves symbolic expressions and their coefficient-optimization routines using LLM-guided mutation and evolutionary selection. It uses group-wise normalized mean squared error and held-out evaluation.
Across five scenarios—vocabulary scaling, supervised fine-tuning, domain-mixture scaling, mixture-of-experts scaling, and data-constrained pretraining—the method exactly rediscovered two human-derived forms and outperformed published forms in the other three scenarios. Its function is hypothesis generation and validation, not proof of causal or universal laws (Lin et al., 27 Jul 2025).
6. Limitations, controversies, and interpretation
Scaling laws have several recurring limitations.
Finite regimes: A relation such as 87 or 88 is generally a local approximation. It can fail outside the measured range, at crossovers, near saturation, or when a new mechanism appears.
Metric dependence: Exponents depend on the observable. Cross-entropy, classification error, NDCG, MAP, MRR, calibration error, and rank-based probability can have different scaling behavior. Comparisons between exponents fitted to different metrics are therefore approximate.
Architecture and protocol dependence: Deep-learning scaling laws depend on model family, optimizer, hyperparameters, data distribution, training schedule, compression method, and evaluation protocol. The reported laws are not universal equations for all neural networks.
Spectral and geometric dependence: In regression and kernel models, exponents arise from covariance-spectrum tails and target smoothness. In nearest-neighbor classification, fast or slow rates depend on positive or negative dominance and high-dimensional neighborhood geometry (Yang et al., 2023). In martensitic materials, the exponent depends on compatibility between domain geometry and energy wells. Similar exponents can consequently arise from different mechanisms.
Overfitting and extrapolation: High 89 establishes regularity within sampled data but does not guarantee reliable extrapolation. Symbolic forms can fit passive datasets while reflecting correlations, omitted variables, or selection effects. EvoSLD and conventional scaling analyses therefore require held-out validation and, ideally, new controlled experiments.
Interpretive status: Some relations are strongly established within explicit regimes, such as Bjorken scaling, quark-counting behavior, or the variational energy bounds for specified film models. Others are phenomenological, such as dense-model scaling and reranking laws. Still others are conjectural, including Nyquist learners for deep learning, the extension of hadronic spin–mass relations to astronomical objects, and the claim that neural scaling exponents universally represent redundancy.
The most general interpretation is that scaling laws expose low-dimensional structure in systems whose microscopic or operational details are high-dimensional. A useful law identifies the relevant scale variables, the observable, the regime of validity, the exponent, and the mechanisms responsible for deviations. Across the surveyed domains, the recurring chain is
90
Scaling laws are therefore best treated as theoretically interpretable, empirically testable regime descriptions rather than unconditional universal formulas.