---
title: 'Scaling Laws: Definitions, Models, and Applications'
url: https://www.emergentmind.com/topics/skaling-law
type: topic
---

# Scaling Laws: Definitions, Models, and Applications

Scaling laws are quantitative relationships describing how an observable changes under variation of scale, resource, geometry, or structural complexity. In their canonical form, they are power-law relations such as $f(x)=cx^n$, for which dilation $x\rightarrow\lambda x$ gives $f(\lambda x)=\lambda^n f(x)$. On a double-logarithmic plot, $\log f=\log c+n\log x$, so the exponent is the slope. Scaling laws are used across particle physics, astrophysics, linguistics, machine learning, materials science, fluid dynamics, information retrieval, and agent systems because they can expose robust dynamical or statistical structure without requiring a complete solution of the underlying theory. Their validity is generally regime-dependent: crossovers, saturation, finite-size effects, spectral structure, or changes in the governing mechanism can produce multiple successive scaling laws rather than one universal relation.

## 1. Mathematical meaning and methodological role

A scaling law is commonly represented by a homogeneous power law,

$$
f(x)=cx^n,
$$

where $n$ is the scaling exponent and $c$ supplies the dimensions required for consistency. It satisfies Euler’s homogeneity equation,

$$
xf'(x)=nf(x).
$$

Under a dilation,

$$
f(\lambda x)=\lambda^n f(x),
$$

so the functional form remains invariant apart from an overall multiplicative factor. With logarithmic coordinates,

$$
Y=\log y,\qquad X=\log x,
$$

the relation becomes

$$
Y=\log c+nX.
$$

The exponent is therefore the slope of a log-log plot, while the coefficient determines its vertical position.

The methodological significance of an exponent depends on the mechanism that generates it. In particle physics, an exponent may encode the number of active constituents, as in dimensional quark-counting rules. In a geometric variational problem, it may arise from balancing surface and elastic energies. In statistical learning, it may be controlled by a covariance-spectrum tail or an effective dimension. In language modeling, it may describe the tradeoff between model capacity, data, and privacy noise. A power law is therefore not merely a curve-fitting convenience when its variables and exponent have a theoretically motivated interpretation.

Scaling relations are usually approximate rather than exact. A single law may describe only an intermediate regime, with earlier or later behavior controlled by different mechanisms. “Radical scaling” emphasizes that two intersecting laws can provide objective characteristic units, while three connected laws can also define an objective dimensionless radix for representing regime crossovers [2507.02631]. The same work distinguishes dimensional units, nondimensional variables, and the choice of logarithmic base: the latter changes the representation of multiplicative separations without changing the underlying physical data.

## 2. Physical scaling laws in particle physics and astrophysics

### Particle-physics scaling

Bjorken scaling emerged from deep-inelastic electron–proton scattering at SLAC in 1969. For

$$
e+p\rightarrow e+\text{anything},
$$

the structure functions were approximately independent of the absolute momentum-transfer scale when expressed in terms of the dimensionless Bjorken variable

$$
x_B=\frac{Q^2}{2p\cdot q},
$$

where $Q^2=-q^2$. Schematically,

$$
F_i(x_B,Q^2)\simeq F_i(x_B).
$$

Its interpretation was that the proton contains effectively point-like constituents—partons, later identified with quarks and gluons. Exact scaling is replaced in QCD by calculable $Q^2$ evolution due to gluon radiation and quark–gluon interactions, but the approximate experimental scaling was decisive evidence for physical proton substructure [1106.1270].

Dimensional quark-counting rules relate the asymptotic electromagnetic form factor of a composite object $a$ to the number $n_a$ of elementary constituents:

$$
F_a(t)\sim t^{-(n_a-1)}.
$$

The constituent counts discussed for the pion, nucleon, deuteron, ${}^3\mathrm{He}$, and ${}^4\mathrm{He}$ are respectively $2$, $3$, $6$, $9$, and $12$. Thus, in schematic form,

$$
F_\pi(t)\sim t^{-1},\qquad
F_N(t)\sim t^{-2},\qquad
F_d(t)\sim t^{-5}.
$$

For exclusive reactions,

$$
a+b\rightarrow c+d,
$$

the corresponding cross-section law is

$$
\frac{d\sigma}{dt}(ab\rightarrow cd)
=
\frac{f(t/s)}{s^{n-2}},
\qquad
n=n_a+n_b+n_c+n_d.
$$

At fixed center-of-mass angle and large $s$, the examples include

$$
\pi+p\rightarrow\pi+p:\quad \frac{d\sigma}{dt}\sim s^{-8},
$$

$$
p+p\rightarrow p+p:\quad \frac{d\sigma}{dt}\sim s^{-10},
$$

and

$$
\gamma+d\rightarrow n+p:\quad \frac{d\sigma}{dt}\sim s^{-11}.
$$

These exponents have a structural interpretation as counts of active elementary degrees of freedom. The rules are nevertheless restricted to hard-scattering regimes and can be modified by logarithmic QCD corrections, helicity-selection effects, endpoint configurations, higher-twist contributions, and soft rescattering.

### Regge and astrophysical angular-momentum relations

The Chew–Frautschi relation places hadrons approximately on Regge trajectories,

$$
J=\alpha(0)+\alpha' M^2.
$$

Neglecting the intercept gives $J\propto M^2$, or, in the normalization used in the paper,

$$
J\simeq \hbar\left(\frac{m}{m_p}\right)^2.
$$

The relation is associated with Regge phenomenology and string-like or rotating extended objects. It is not an exact universal trajectory: flavor families, parity sectors, radial excitations, and intercepts produce distinct trajectories.

An astrophysical generalization assigns an effective geometric dimension $n$:

$$
J_n=\hbar\left(\frac{m}{m_p}\right)^{1+1/n},
\qquad n=1,2,3.
$$

The cases are interpreted as string-like, disk-like, and ball-like objects:

$$
J_{\rm string}\propto m^2,\qquad
J_{\rm disk}\propto m^{3/2},\qquad
J_{\rm ball}\propto m^{4/3}.
$$

The disk relation is associated with galaxies and the ball relation with stars and planets. The paper gives example angular momenta of approximately $5.79\times10^{69}\,\mathrm{J\,s}$ for Andromeda and $1.11\times10^{70}\,\mathrm{J\,s}$ for the Milky Way, using estimated masses including dark matter.

The extremal Kerr bound for a rotating black hole is

$$
J_{\rm Kerr}\sim\frac{Gm^2}{c},
$$

with dimensionless spin parameter

$$
\chi=\frac{cJ}{Gm^2}\leq1.
$$

Using the Planck mass,

$$
m_{\rm Pl}=\sqrt{\frac{\hbar c}{G}},
$$

this becomes

$$
J_{\rm Kerr}=\hbar\left(\frac{m}{m_{\rm Pl}}\right)^2.
$$

The Kerr and hadronic relations are parallel on a log-log plot because both have exponent $2$, but their coefficients differ by approximately $1.7\times10^{38}$. Parallel exponents do not establish a common mechanism: the Kerr relation follows from general-relativistic black-hole mechanics, whereas the hadronic relation is an approximate empirical relation for strong-interaction bound states.

## 3. Scaling in statistical systems and continuum mechanics

### Linguistic scaling

Zipf’s law relates word frequency $n$ and rank $r$ through

$$
n(r)\propto r^{-\beta},
$$

or, for the distribution of types by absolute frequency,

$$
D(n)\propto n^{-\gamma},
\qquad
\gamma=1+\frac{1}{\beta}.
$$

A finite-text scaling formulation instead defines relative frequency

$$
x=\frac{n}{L},
$$

and assumes that rank depends on frequency and text length through their ratio:

$$
r=G\left(\frac{n}{L}\right).
$$

Differentiation yields

$$
D_L(n)=\frac{g(n/L)}{LV_L},
\qquad
g(x)=-G'(x),
$$

where $V_L$ is vocabulary size. Consequently, plotting $LV_LD_L(n)$ against $n/L$ should collapse distributions from different text lengths onto a length-independent curve. The invariant is the shape in relative-frequency coordinates, whereas absolute crossover frequencies increase linearly with $L$ [1303.0705].

For lemmatized texts, the proposed scaling function is

$$
g(x)=\frac{k}{x(a+x^{\gamma-1})}.
$$

It has a low-frequency regime $g(x)\sim x^{-1}$ and a high-frequency regime $g(x)\sim x^{-\gamma}$. The high-frequency exponent is close to $\gamma\simeq2$, corresponding to the conventional Zipf exponent $\beta\simeq1$. The crossover occurs at

$$
n_a=La^{1/(\gamma-1)},
$$

so its absolute location grows linearly with text length while its relative location remains fixed.

The vocabulary relation follows from

$$
V_L=G(1/L).
$$

For the double-power-law form, the continuous approximation gives

$$
V_L=
\frac{k}{a(\gamma-1)}
\ln\!\left(aL^{\gamma-1}+1\right).
$$

Thus the asymptotic growth is logarithmic rather than a pure Heaps law. The paper emphasizes that discreteness at low frequencies is essential: the discrete expression tracks empirical vocabulary growth better than the continuous approximation.

### Elastic films with dislocations

A variational model for a two-dimensional epitaxial film includes surface energy, elastic misfit energy, and dislocation nucleation energy. The total energy is

$$
\mathcal F(h,H,\sigma)
=
\gamma\int_0^1\sqrt{1+|h'(x)|^2}\,dx
+
\int_{\Omega_h}W(H)\,dx\,dy
+
c_0kb^2.
$$

Here $\gamma$ is surface tension, $e_0$ is lattice mismatch, $b$ is the Burgers-vector magnitude, $r_0$ is the dislocation core radius, and $k$ is the number of dislocations. Under the stated assumptions, the infimal energy obeys, up to multiplicative constants,

$$
\inf\mathcal F
\asymp
\gamma(1+d)
+
\min\left\{
(\gamma e_0d)^{2/3},
\left[
\gamma e_0bd
\left(1+\log\frac{b}{e_0r_0}\right)
\right]^{1/2}
\right\}.
$$

The first branch is the coherent or defect-free island scale. For an island of width $L$,

$$
\mathcal F_{\rm island}(L)
\sim
\gamma+\frac{\gamma d}{L}+e_0^2L^2,
$$

and balancing the last two terms gives

$$
L_{\rm el}\sim e_0^{-2/3}(\gamma d)^{1/3}.
$$

The dislocation branch follows from a spacing scale

$$
\ell_{\rm dis}\sim\frac{b}{e_0},
$$

with approximately $Le_0/b$ dislocations in an island of width $L$. Its energy is

$$
\mathcal F_{\rm dis}(L)
\sim
\gamma+\frac{\gamma d}{L}
+
Le_0b\left(1+\log\frac{b}{e_0r_0}\right),
$$

leading to

$$
L_{\rm dis}
\sim
(\gamma d)^{1/2}
\left[
e_0b\left(1+\log\frac{b}{e_0r_0}\right)
\right]^{-1/2}.
$$

The logarithm is the signature of two-dimensional dislocation self-energy. The lower bound uses a ball construction, Korn inequalities for fields with nonzero curl, and local estimates distinguishing elastic mismatch from defect-mediated relaxation [2403.13646].

### Martensitic transformations and geometry

For two-dimensional martensitic phase transformations, the geometrically nonlinear energy is

$$
E_\epsilon(u,\Omega)
=
\int_\Omega \operatorname{dist}^2(\nabla u,K)\,dx
+
\epsilon |D^2u|(\Omega),
$$

while the linearized model is

$$
F_\epsilon(v,\Omega)
=
\int_\Omega \operatorname{dist}^2(e(v),\widetilde K)\,dx
+
\epsilon |D^2v|(\Omega).
$$

If the boundary tangent or normal directions are incompatible with the martensitic wells in the sense of the Hadamard jump condition, the minimum energy has the generic scaling

$$
E_{\min}(\epsilon)\asymp
\epsilon\bigl(|\log\epsilon|+1\bigr).
$$

For specially adapted polygonal domains with compatible geometry,

$$
E_{\min}(\epsilon)\asymp\epsilon.
$$

The logarithmic loss arises from a slice estimate of order $\epsilon/x$ integrated across boundary distances:

$$
\int_\epsilon^1\frac{\epsilon}{x}\,dx
=
\epsilon|\log\epsilon|.
$$

Compatible polygonal domains support finite-scale, stress-free or nearly stress-free microstructures, whereas generic domains require increasingly fine branching toward incompatible boundaries [2405.05927]. The scaling law thus acts as a selection principle for domain geometry: wells and symmetry determine preferred polygonal morphologies.

### Explosions and capillary phenomena

For explosions, three regimes are identified:

$$
d\simeq c_0t,
\qquad
d\simeq Kt^{2/5},
\qquad
d\simeq c_st.
$$

The first is finite-speed initial ejection, the second is the Taylor–Sedov blast regime, and the third is acoustic propagation. The characteristic crossovers are

$$
t_0\simeq\left(\frac{K}{c_0}\right)^{5/3},
\qquad
t_*\simeq\left(\frac{K}{c_s}\right)^{5/3}.
$$

The dimensionless initial Mach number is

$$
\mathcal N=\frac{c_0}{c_s}.
$$

Using the radix $R=\mathcal N^{1/3}$ and Hopkinson–Cranz variables gives the idealized master curve

$$
y=
\begin{cases}
x+3,&x<-5,\\[2pt]
\frac25x,&-5<x<0,\\[2pt]
x,&x>0.
\end{cases}
$$

The framework has been applied to nuclear, chemical, underwater, laser-induced, and vapor-cloud explosions.

For capillarity-driven processes, the proposed regimes are

$$
d\simeq c_vt,\qquad
d\simeq K_it^{2/3},\qquad
d\simeq D,
$$

representing viscous, inertio-capillary, and geometric-saturation regimes. Their crossover structure is controlled by an inverse-Ohnesorge-type quantity

$$
\mathcal N\equiv \frac{c_vD^{1/2}}{K_i^{3/2}}.
$$

The construction has been applied to drop and bubble spreading, coalescence, and pinch-off. For completely wetting spreading, the terminal plateau is only approximate because Tanner’s law gives a later $t^{1/10}$ regime [2507.02631].

## 4. Scaling laws in learning systems

### Dense-model and pruning laws

For a fixed dataset size $n$, model error is approximated by

$$
\epsilon(m,n)\approx b(n)m^{-\beta(n)}+c_m(n),
$$

while for fixed model size $m$,

$$
\epsilon(m,n)\approx a(m)n^{-\alpha(m)}+c_n(m).
$$

A joint approximation is

$$
\epsilon(m,n)
\approx
a n^{-\alpha}
+
b m^{-\beta}
+
c_\infty.
$$

The terms represent data-limited error, model-limited error, and an asymptotic lower error. The thesis also introduces a rational complex-envelope transition to model the random-guess regime, intermediate power-law region, and asymptotic plateau. These forms are phenomenological approximations, with exponents fitted per task and scaling policy [2108.07686].

Training compute is approximated by

$$
C_{\mathrm{train}}\propto mn.
$$

In the power-law regime, minimizing $mn$ at fixed error leads to the allocation condition

$$
\frac{b\beta}{\alpha}\frac{n^\alpha}{m^\beta}=1.
$$

Thus model and data should be increased in a balanced manner rather than independently.

Iterative magnitude pruning exhibits three density regimes: a low-error plateau, an intermediate power-law region, and a high-error plateau. The intermediate behavior is

$$
\epsilon(d,w)\approx cd^{-\gamma},
$$

where $d$ is retained density. Depth $l$, width $w$, and density can compensate through an approximate invariant

$$
m^*=l^\phi w^\psi d.
$$

This invariant produces a joint pruning law in which architecture and sparsity preserve approximately constant error. Its validity is established for iterative magnitude pruning with weight rewinding and is not automatically transferable to structured pruning, quantization, distillation, or other compression procedures.

### Linear, kernel, and quadratic regression

For sketched infinite-dimensional linear regression with covariance eigenvalues

$$
\lambda_i\asymp i^{-a},
\qquad a>1,
$$

one-pass SGD produces reducible error of the form

$$
\Theta\!\left(M^{-(a-1)}+N_{\mathrm{eff}}^{-(a-1)/a}\right),
\qquad
N_{\mathrm{eff}}=\frac{N}{\log N}.
$$

The first term is sketch-induced approximation error; the second is finite-training bias. SGD’s implicit regularization keeps variance lower order under the stated one-pass schedule, even though variance can increase with model size [2406.08466]. Balancing the two terms under compute $C=MN$ gives, up to logarithmic factors,

$$
M=\widetilde\Theta\!\left(C^{1/(a+1)}\right),
\qquad
N=\widetilde\Theta\!\left(C^{a/(a+1)}\right).
$$

Related scaling extends to multiple regression and kernel regression, where a Gaussian sketch dimension acts as effective model size and the covariance or kernel spectrum determines the exponent [2503.01314].

A quadratically parameterized linear model uses

$$
f_{\mathbf v}(\mathbf x)
=
\langle\mathbf x,\mathbf v^{\odot2}\rangle.
$$

Its SGD update contains a multiplicative factor $v_i$, producing adaptive feature selection. Under covariance decay $\lambda_i\asymp i^{-\alpha}$ and source decay $\lambda_i(v_i^*)^4\asymp i^{-\beta}$, the dynamically learned effective dimension is

$$
D\asymp
\min\left\{
T^{1/\max\{\beta,(\alpha+\beta)/2\}},
M
\right\}.
$$

In the aligned regime $\alpha\leq\beta$, the sample-size exponent is

$$
\frac{\beta-1}{\beta},
$$

whereas in the opposing regime $\alpha>\beta$, quadratic SGD has exponent

$$
\frac{2\beta-2}{\alpha+\beta},
$$

compared with $(\beta-1)/\alpha$ for ordinary linear SGD [2502.09106].

### Redundancy and spectral scaling

A kernel covariance spectrum

$$
\lambda_i\asymp i^{-1/\beta},
\qquad \beta>1,
$$

has effective dimension

$$
N_{\mathrm{eff}}(\lambda)
=
\sum_i\frac{\lambda_i}{\lambda_i+\lambda}
\asymp
\lambda^{-1/\beta}.
$$

For a target satisfying a source condition with smoothness $s$, kernel ridge regression has bias–variance structure

$$
\mathbb E\,\mathcal E(f_{\lambda,n})
\lesssim
\lambda^{2s}
+
\frac{\sigma^2}{n}\lambda^{-1/\beta}.
$$

Optimizing $\lambda$ gives

$$
\lambda^\star\asymp n^{-1/(2s+1/\beta)}
$$

and

$$
\mathbb E\,\mathcal E(f_{\lambda^\star,n})
\asymp
n^{-\alpha},
\qquad
\alpha=
\frac{2s}{2s+1/\beta}.
$$

Here $1/\beta$ is termed the redundancy index: a flatter spectrum has more statistically active directions and slower scaling. Boundedly invertible representations preserve the spectral tail and exponent, while mixtures are controlled asymptotically by the component with the smallest $\beta$ [2509.20721].

### Time-series forecasting

Time-series forecasting introduces look-back horizon $H$ as an additional scaling variable. The loss decomposes into Bayesian error and approximation error,

$$
L=L_{\mathrm{Bayesian}+L_{\mathrm{approx}}.
$$

Longer histories reduce truncation uncertainty, but increase intrinsic dimension $d_I(H)$ and reduce the effective number of independent sliding-window samples, approximately

$$
D_{\mathrm{eff}}\propto\frac{D}{H}.
$$

In data-sufficient settings, model partitioning error decreases with complexity while estimation error increases with complexity. In data-scarce settings, nearest-neighbor behavior yields approximation error proportional to

$$
d_I(H)D^{-2/d_I(H)}.
$$

The resulting tradeoff predicts an interior optimal horizon: more history is beneficial initially but can harm performance when added dimensions exceed the available data and model capacity. Empirical results on ETTh1, ETTh2, ETTm1, ETTm2, Exchange, Weather, ECL, and Traffic support dataset-size scaling, model-width scaling, saturation, and horizon optima [2405.15124].

### Privacy-constrained and multilingual scaling

Differential privacy adds gradient noise as a new scaling dimension. The compute model is

$$
C\approx6MBST,
$$

where $M$ is parameter count, $B$ batch size, $S$ sequence length, and $T$ iterations. Utility is modeled as a surface

$$
L(M,T,\bar\sigma),
$$

where $\bar\sigma$ is the noise-batch ratio. Under moderate privacy, compute-optimal private training favors substantially smaller models, much larger batches, and more examples per parameter than non-private training. The paper reports optimal batches often in the range $10\mathrm{K}$–$100\mathrm{K}$ and token-to-model ratios often between $10^3$ and $10^5$ for $\epsilon\in[1,10]$ [2501.18914].

For multilingual language models, language-family loss is modeled as

$$
L_i(N,D,p_i)
=
\left(
E_i+
\frac{A_i}{N^{\alpha_i}+B_i/D^{\beta_i}}
\right)p_i^{-\gamma_i},
$$

where $p_i$ is the family’s sampling ratio. The family-independence hypothesis states that a family’s loss depends primarily on its own sampling ratio, while within-family transfer remains important. Optimizing a weighted mixture gives

$$
p_i
\propto
(w_iL_i^\star\gamma_i)^{1/(1+\gamma_i)}.
$$

When $w_i=1/L_i^\star$ and $\gamma_i$ is approximately scale-invariant,

$$
p_i^\star\approx
\frac{\gamma_i}{\sum_j\gamma_j}.
$$

Experiments involving 23 languages and five families found that ratios optimized on 85-million-parameter models transferred effectively to models up to 1.2 billion parameters [2410.12883].

## 5. Scaling laws for ranking, retrieval, and agent systems

### Rank-based language-model scaling

Relative-Based Probability is defined by the rank $R$ of the correct token:

$$
\mathrm{RBP}_k=\Pr(R\leq k).
$$

The proposed law is

$$
-\log\mathrm{RBP}_k\propto S^{-\alpha},
\qquad k\ll|\mathcal V|,
$$

where $S$ is the number of non-embedding parameters. It complements cross-entropy, which measures the absolute probability assigned to the correct token. RBP instead measures whether the correct token is ranked first or within a top-$k$ candidate set. Across Pythia, GPT-2, OPT, and Qwen families, the reported fits have high $R^2$, approximately $0.99$ for $k=1$ and at least $0.97$ for many moderate-$k$ experiments. The law becomes unstable when $k$ approaches vocabulary size because $\mathrm{RBP}_k$ saturates near one [2510.20387].

A sequence requiring $N$ consecutive correct top-$k$ decisions has, under independence,

$$
p_{N,k}=(\mathrm{RBP}_k)^N.
$$

Thus a smooth token-level scaling law can produce a sharp sequence-level transition without requiring a discontinuous underlying capability transition.

### Reranking

For cross-encoder rerankers, quality is modeled with saturating laws such as

$$
\mathcal M(M)=a-bM^{-c},
$$

$$
\mathcal M(S)=a-bS^{-c},
$$

and

$$
\mathcal M(M,S)=a-bM^{-\alpha}-cS^{-\beta}.
$$

Experiments across pointwise, pairwise, and listwise reranking use the Ettin series from 17 million to 1 billion parameters, with BM25 top-100 candidate sets from MS MARCO. NDCG@10 and MAP exhibit comparatively reliable scaling, whereas Contrastive Entropy is often unstable and MRR is dataset-dependent. Models up to 400 million parameters were used to forecast the NDCG of a 1-billion-parameter reranker with low reported RMSE [2603.04816].

### Agent skill-library scaling

In LLM-agent systems, the scaling variable can be the number of exposed skills rather than parameters or training data. The empirical routing law is

$$
Acc(N)=a-b\ln N,
$$

where $N$ is the number of simultaneously exposed skills and $b$ is the logarithmic decay slope. Across 15 frontier models and 1,141 skills, all reported fits have $R^2>0.97$ over the tested range $N\in\{10,20,50,100,200,500\}$. Errors progress from local competition to cross-family drift and capture by overly general “black-hole skills.” Local semantic competition predicts errors better than global library size alone.

For multi-step route-only pipelines,

$$
Acc(N,K)\approx(a-b\ln N)^{\gamma K},
\qquad
\gamma=6.7b+1.09.
$$

Before execution artifacts are realized, downstream routing is approximately multiplicative. After correct upstream execution, a concrete artifact can rescue downstream decisions:

$$
\Delta P(B\mid A)
=
2\alpha(1-P(B))P(A),
$$

with pooled estimate $2\alpha\approx0.76$. A law-guided library intervention increased held-out routing accuracy from $71.3\%$ to $91.7\%$ and reduced in-library hijack from $22.4\%$ to $4.1\%$ [2605.16508].

### Automated scaling-law discovery

EvoSLD formulates scaling-law discovery as grouped symbolic learning. Scaling variables $\mathbf x$ are varied within an experiment, while control variables $\mathbf c$ define groups with separate coefficients. The goal is a shared symbolic structure

$$
f(\mathbf x;\boldsymbol\theta_{\mathbf c})
$$

with group-specific parameter values. The method co-evolves symbolic expressions and their coefficient-optimization routines using LLM-guided mutation and evolutionary selection. It uses group-wise normalized mean squared error and held-out evaluation.

Across five scenarios—vocabulary scaling, supervised fine-tuning, domain-mixture scaling, mixture-of-experts scaling, and data-constrained pretraining—the method exactly rediscovered two human-derived forms and outperformed published forms in the other three scenarios. Its function is hypothesis generation and validation, not proof of causal or universal laws [2507.21184].

## 6. Limitations, controversies, and interpretation

Scaling laws have several recurring limitations.

**Finite regimes**: A relation such as $a-b\ln N$ or $a-bM^{-c}$ is generally a local approximation. It can fail outside the measured range, at crossovers, near saturation, or when a new mechanism appears.

**Metric dependence**: Exponents depend on the observable. Cross-entropy, classification error, NDCG, MAP, MRR, calibration error, and rank-based probability can have different scaling behavior. Comparisons between exponents fitted to different metrics are therefore approximate.

**Architecture and protocol dependence**: Deep-learning scaling laws depend on model family, optimizer, hyperparameters, data distribution, training schedule, compression method, and evaluation protocol. The reported laws are not universal equations for all neural networks.

**Spectral and geometric dependence**: In regression and kernel models, exponents arise from covariance-spectrum tails and target smoothness. In nearest-neighbor classification, fast or slow rates depend on positive or negative dominance and high-dimensional neighborhood geometry [2308.08247]. In martensitic materials, the exponent depends on compatibility between domain geometry and energy wells. Similar exponents can consequently arise from different mechanisms.

**Overfitting and extrapolation**: High $R^2$ establishes regularity within sampled data but does not guarantee reliable extrapolation. Symbolic forms can fit passive datasets while reflecting correlations, omitted variables, or selection effects. EvoSLD and conventional scaling analyses therefore require held-out validation and, ideally, new controlled experiments.

**Interpretive status**: Some relations are strongly established within explicit regimes, such as Bjorken scaling, quark-counting behavior, or the variational energy bounds for specified film models. Others are phenomenological, such as dense-model scaling and reranking laws. Still others are conjectural, including Nyquist learners for deep learning, the extension of hadronic spin–mass relations to astronomical objects, and the claim that neural scaling exponents universally represent redundancy.

The most general interpretation is that scaling laws expose low-dimensional structure in systems whose microscopic or operational details are high-dimensional. A useful law identifies the relevant scale variables, the observable, the regime of validity, the exponent, and the mechanisms responsible for deviations. Across the surveyed domains, the recurring chain is

$$
\text{structure or geometry}
\longrightarrow
\text{effective dimension or competition}
\longrightarrow
\text{bias--variance or energy balance}
\longrightarrow
\text{power-law scaling}.
$$

Scaling laws are therefore best treated as theoretically interpretable, empirically testable regime descriptions rather than unconditional universal formulas.

Source: https://www.emergentmind.com/topics/skaling-law