Papers
Topics
Authors
Recent
Search
2000 character limit reached

Asymmetric Scaling Formula

Updated 10 July 2026
  • Asymmetric scaling formula is a framework where different variables or regimes are scaled non-uniformly with distinct exponents and prefactors.
  • It appears in diverse applications such as knowledge distillation, sparse-feature modeling, plasma turbulence, and gravitational thermodynamics.
  • The method captures intricate system behaviors by encoding directional, channel-dependent, or regime-specific asymmetries in its scaling relations.

An asymmetric scaling formula is a scaling relation in which distinct variables, channels, geometrical sectors, or fluctuation regimes are not rescaled uniformly. In the arXiv literature, the term appears in several unrelated technical settings, including knowledge distillation, sparse-feature neural scaling laws, tokamak momentum transport, charged Lifshitz black-hole thermodynamics, finite-size percolation, asymmetric nuclear response modeling, and large-deviation theory. Across these settings, asymmetry typically means that one side of a relation is controlled by a different exponent, prefactor, temperature, or charge coefficient than another, or that different phases obey genuinely different asymptotic laws (Li et al., 2022, Sous et al., 22 May 2026, Ball et al., 2016, Bravo-Gaete et al., 2015, Žeželj et al., 2011, Martinez-Consentino et al., 10 Sep 2025, Monthus, 2019).

1. General forms of asymmetric scaling

The literature exhibits several recurrent mathematical forms. One form assigns different scaling parameters to different components of the same object. In knowledge distillation, asymmetric temperature scaling uses

pc(τ1,τ2)=exp(fc/τc)j[C]exp(fj/τj),τi=I{i=y}τ1+I{iy}τ2,\mathbf{p}_c(\tau_1,\tau_2) = \frac{\exp(\mathbf{f}_c/\tau_c)} {\sum_{j\in[C]}\exp(\mathbf{f}_j/\tau_j)}, \qquad \tau_i=\mathcal{I}\{i=y\}\tau_1+\mathcal{I}\{i\neq y\}\tau_2,

with the recommendation τ1>τ2>0\tau_1 > \tau_2 > 0, so the correct class and wrong classes are softened differently (Li et al., 2022). A second form assigns different exponents to different asymptotic regimes. In sparse random-feature models, the summary law

Bayes(N,D)N(α1+α2+1)+D(α1+α2+1)/(α1+1)\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-(\alpha_1+\alpha_2+1)}+ D^{-(\alpha_1+\alpha_2+1)/(\alpha_1+1)}

gives one exponent in the underparameterized regime and another in the overparameterized regime (Sous et al., 22 May 2026). A third form distinguishes symmetry classes or channels. In tokamaks, non-mirror symmetric shaping yields a power law, whereas mirror-symmetric shaping yields exponential suppression (Ball et al., 2016). In asymmetric nuclear 2p2h response modeling, the scaling is channel dependent and uses separate proton and neutron Fermi momenta (Martinez-Consentino et al., 10 Sep 2025). In finite-size percolation, asymmetry enters through aspect-ratio-dependent prefactors ai(r)a_i(r) and bi(r)b_i(r) in

nL,r=nc+L1/νi=0ai(r)Lθi,ΔL,r2=L2/νi=0bi(r)Lθi.\langle n \rangle_{L,r} = n_c + L^{-1/\nu}\sum_{i=0}^\infty a_i(r) L^{-\theta_i}, \qquad \Delta^2_{L,r} = L^{-2/\nu}\sum_{i=0}^\infty b_i(r) L^{-\theta_i}.

(Žeželj et al., 2011)

Domain Representative formula Asymmetry source
Knowledge distillation pc(τ1,τ2)\mathbf{p}_c(\tau_1,\tau_2) Correct vs wrong classes
Sparse-feature scaling laws Bayes(N,D)NαN+DαD\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-\alpha_N}+D^{-\alpha_D} Model-limited vs data-limited regimes
Tokamak momentum flux power law vs Πsβeγm\Pi_s \sim -\beta e^{-\gamma m} Non-mirror symmetric vs mirror symmetric shaping
2p2h MEC response RK(X)RK(12C)×(factor)×C(X)R^K(X)\approx R^K({}^{12}\mathrm{C})\times(\text{factor})\times\mathcal C(X) Proton-neutron imbalance and channel dependence
Finite-size percolation τ1>τ2>0\tau_1 > \tau_2 > 00 in τ1>τ2>0\tau_1 > \tau_2 > 01 Aspect-ratio asymmetry

These patterns show that “asymmetric scaling formula” is not a single canonical equation. It is a family resemblance across formulas in which the scaling map itself encodes a directional, structural, or regime-dependent inequivalence.

2. Learning-theoretic asymmetric scaling

In knowledge distillation, the relevant asymmetry is class selective. The KD term is decomposed into Correct Guidance, Smooth Regularization, and Class Discriminability, with the last term governed by variance among wrong-class probabilities. The paper defines

τ1>τ2>0\tau_1 > \tau_2 > 02

and proves the factorization

τ1>τ2>0\tau_1 > \tau_2 > 03

The stated motivation is that complex teachers tend to be over-confident and traditional temperature scaling limits the efficacy of class discriminability, so ATS uses a larger τ1>τ2>0\tau_1 > \tau_2 > 04 on the correct class and a smaller τ1>τ2>0\tau_1 > \tau_2 > 05 on wrong classes to enlarge the variance of wrong-class probabilities (Li et al., 2022).

A distinct learning-theoretic use of asymmetric scaling appears in sparse-feature neural scaling laws. There, the data are sparse coordinates with

τ1>τ2>0\tau_1 > \tau_2 > 06

and the expected number of coordinates active at least once across τ1>τ2>0\tau_1 > \tau_2 > 07 samples is

τ1>τ2>0\tau_1 > \tau_2 > 08

This creates an asymmetric bottleneck: in the data-limited regime, rare coordinates are never observed, while in the parameter-limited regime the bottleneck is the number of learnable directions. The resulting exponents are

τ1>τ2>0\tau_1 > \tau_2 > 09

with Bayes(N,D)N(α1+α2+1)+D(α1+α2+1)/(α1+1)\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-(\alpha_1+\alpha_2+1)}+ D^{-(\alpha_1+\alpha_2+1)/(\alpha_1+1)}0 for Bayes(N,D)N(α1+α2+1)+D(α1+α2+1)/(α1+1)\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-(\alpha_1+\alpha_2+1)}+ D^{-(\alpha_1+\alpha_2+1)/(\alpha_1+1)}1, and the interpolation threshold is near

Bayes(N,D)N(α1+α2+1)+D(α1+α2+1)/(α1+1)\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-(\alpha_1+\alpha_2+1)}+ D^{-(\alpha_1+\alpha_2+1)/(\alpha_1+1)}2

The same paper derives a compute-optimal frontier under

Bayes(N,D)N(α1+α2+1)+D(α1+α2+1)/(α1+1)\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-(\alpha_1+\alpha_2+1)}+ D^{-(\alpha_1+\alpha_2+1)/(\alpha_1+1)}3

finding

Bayes(N,D)N(α1+α2+1)+D(α1+α2+1)/(α1+1)\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-(\alpha_1+\alpha_2+1)}+ D^{-(\alpha_1+\alpha_2+1)/(\alpha_1+1)}4

so the compute-optimal frontier favors more data than parameters (Sous et al., 22 May 2026).

These two cases use asymmetry differently. ATS imposes asymmetry directly in the softmax map. Sparse-feature scaling laws derive asymmetry from a latent observation bottleneck. In both cases, the asymmetry is mechanistic rather than merely empirical.

3. Geometry-, channel-, and field-dependent formulas in plasma and scattering

In tokamak turbulence, the scaling law depends on whether the flux surface shaping is non-mirror symmetric or mirror symmetric. For non-mirror symmetric up-down asymmetric shaping, the turbulent toroidal momentum flux obeys

Bayes(N,D)N(α1+α2+1)+D(α1+α2+1)/(α1+1)\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-(\alpha_1+\alpha_2+1)}+ D^{-(\alpha_1+\alpha_2+1)/(\alpha_1+1)}5

for sufficiently weak shaping, specifically when

Bayes(N,D)N(α1+α2+1)+D(α1+α2+1)/(α1+1)\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-(\alpha_1+\alpha_2+1)}+ D^{-(\alpha_1+\alpha_2+1)/(\alpha_1+1)}6

In the explicit two-mode example with Bayes(N,D)N(α1+α2+1)+D(α1+α2+1)/(α1+1)\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-(\alpha_1+\alpha_2+1)}+ D^{-(\alpha_1+\alpha_2+1)/(\alpha_1+1)}7, this becomes

Bayes(N,D)N(α1+α2+1)+D(α1+α2+1)/(α1+1)\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-(\alpha_1+\alpha_2+1)}+ D^{-(\alpha_1+\alpha_2+1)/(\alpha_1+1)}8

By contrast, for mirror-symmetric shaping,

Bayes(N,D)N(α1+α2+1)+D(α1+α2+1)/(α1+1)\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-(\alpha_1+\alpha_2+1)}+ D^{-(\alpha_1+\alpha_2+1)/(\alpha_1+1)}9

with ai(r)a_i(r)0 and ai(r)a_i(r)1 independent of ai(r)a_i(r)2. The asymmetry therefore lies not only in up-down parity, but in whether different shaping harmonics can beat together and survive averaging over the fast coordinate (Ball et al., 2016).

Asymmetric magnetic reconnection with in-plane flow shear uses another class of formulas. The X-line convection speed is

ai(r)a_i(r)3

the reconnection rate is

ai(r)a_i(r)4

and steady reconnection is suppressed above

ai(r)a_i(r)5

Here asymmetry is encoded in the combinations ai(r)a_i(r)6 and ai(r)a_i(r)7, which weight the contributions of the two upstream regions (Doss et al., 2016).

In asymmetric nuclei, the 2p2h MEC response is scaled from ai(r)a_i(r)8C by channel-specific multiplicative laws. For electron scattering, the paper proposes

ai(r)a_i(r)9

bi(r)b_i(r)0

bi(r)b_i(r)1

For charged-current neutrino scattering, the formulas are

bi(r)b_i(r)2

bi(r)b_i(r)3

The stated motivation is phase-space dominance with distinct proton and neutron Fermi seas, and the paper reports that using bi(r)b_i(r)4Ca as a proxy for bi(r)b_i(r)5Ar leads to a systematic error of approximately bi(r)b_i(r)6 (Martinez-Consentino et al., 10 Sep 2025).

4. Finite-size, anisotropic, and probabilistic asymmetric scaling

Finite-size scaling in percolating stick systems uses a generalized scaling function

bi(r)b_i(r)7

The complementarity relation

bi(r)b_i(r)8

induces parity constraints on the prefactors of the moment expansions. For the mean threshold density, bi(r)b_i(r)9 is odd in nL,r=nc+L1/νi=0ai(r)Lθi,ΔL,r2=L2/νi=0bi(r)Lθi.\langle n \rangle_{L,r} = n_c + L^{-1/\nu}\sum_{i=0}^\infty a_i(r) L^{-\theta_i}, \qquad \Delta^2_{L,r} = L^{-2/\nu}\sum_{i=0}^\infty b_i(r) L^{-\theta_i}.0 and nL,r=nc+L1/νi=0ai(r)Lθi,ΔL,r2=L2/νi=0bi(r)Lθi.\langle n \rangle_{L,r} = n_c + L^{-1/\nu}\sum_{i=0}^\infty a_i(r) L^{-\theta_i}, \qquad \Delta^2_{L,r} = L^{-2/\nu}\sum_{i=0}^\infty b_i(r) L^{-\theta_i}.1 is even; for the variance, nL,r=nc+L1/νi=0ai(r)Lθi,ΔL,r2=L2/νi=0bi(r)Lθi.\langle n \rangle_{L,r} = n_c + L^{-1/\nu}\sum_{i=0}^\infty a_i(r) L^{-\theta_i}, \qquad \Delta^2_{L,r} = L^{-2/\nu}\sum_{i=0}^\infty b_i(r) L^{-\theta_i}.2 is even and nL,r=nc+L1/νi=0ai(r)Lθi,ΔL,r2=L2/νi=0bi(r)Lθi.\langle n \rangle_{L,r} = n_c + L^{-1/\nu}\sum_{i=0}^\infty a_i(r) L^{-\theta_i}, \qquad \Delta^2_{L,r} = L^{-2/\nu}\sum_{i=0}^\infty b_i(r) L^{-\theta_i}.3 is odd. The paper also identifies a characteristic aspect ratio nL,r=nc+L1/νi=0ai(r)Lθi,ΔL,r2=L2/νi=0bi(r)Lθi.\langle n \rangle_{L,r} = n_c + L^{-1/\nu}\sum_{i=0}^\infty a_i(r) L^{-\theta_i}, \qquad \Delta^2_{L,r} = L^{-2/\nu}\sum_{i=0}^\infty b_i(r) L^{-\theta_i}.4 at which the threshold probability becomes scale invariant (Žeželj et al., 2011).

Anisotropic random fields on nL,r=nc+L1/νi=0ai(r)Lθi,ΔL,r2=L2/νi=0bi(r)Lθi.\langle n \rangle_{L,r} = n_c + L^{-1/\nu}\sum_{i=0}^\infty a_i(r) L^{-\theta_i}, \qquad \Delta^2_{L,r} = L^{-2/\nu}\sum_{i=0}^\infty b_i(r) L^{-\theta_i}.5 provide another meaning of asymmetry. Partial sums are taken over rectangles of side lengths nL,r=nc+L1/νi=0ai(r)Lθi,ΔL,r2=L2/νi=0bi(r)Lθi.\langle n \rangle_{L,r} = n_c + L^{-1/\nu}\sum_{i=0}^\infty a_i(r) L^{-\theta_i}, \qquad \Delta^2_{L,r} = L^{-2/\nu}\sum_{i=0}^\infty b_i(r) L^{-\theta_i}.6 and nL,r=nc+L1/νi=0ai(r)Lθi,ΔL,r2=L2/νi=0bi(r)Lθi.\langle n \rangle_{L,r} = n_c + L^{-1/\nu}\sum_{i=0}^\infty a_i(r) L^{-\theta_i}, \qquad \Delta^2_{L,r} = L^{-2/\nu}\sum_{i=0}^\infty b_i(r) L^{-\theta_i}.7, and the scaling transition is described by a critical exponent nL,r=nc+L1/νi=0ai(r)Lθi,ΔL,r2=L2/νi=0bi(r)Lθi.\langle n \rangle_{L,r} = n_c + L^{-1/\nu}\sum_{i=0}^\infty a_i(r) L^{-\theta_i}, \qquad \Delta^2_{L,r} = L^{-2/\nu}\sum_{i=0}^\infty b_i(r) L^{-\theta_i}.8. For congruous scaling, the critical exponent is generally nL,r=nc+L1/νi=0ai(r)Lθi,ΔL,r2=L2/νi=0bi(r)Lθi.\langle n \rangle_{L,r} = n_c + L^{-1/\nu}\sum_{i=0}^\infty a_i(r) L^{-\theta_i}, \qquad \Delta^2_{L,r} = L^{-2/\nu}\sum_{i=0}^\infty b_i(r) L^{-\theta_i}.9 or its reciprocal. For incongruous or oblique dependence axis, the paper proves the universal result

pc(τ1,τ2)\mathbf{p}_c(\tau_1,\tau_2)0

This sharply separates axis-aligned and obliquely oriented dependence structures (Pilipauskaitė et al., 2020).

Weakly asymmetric bridges classify asymptotic behavior by the size of the asymmetry

pc(τ1,τ2)\mathbf{p}_c(\tau_1,\tau_2)1

The paper identifies three thresholds: pc(τ1,τ2)\mathbf{p}_c(\tau_1,\tau_2)2 for the hydrodynamic transition between the heat equation and nonlinear Hamilton–Jacobi/Burgers behavior, pc(τ1,τ2)\mathbf{p}_c(\tau_1,\tau_2)3 for the comparison between mean shape and fluctuations in equilibrium, and pc(τ1,τ2)\mathbf{p}_c(\tau_1,\tau_2)4 for the KPZ window (Labbé, 2016). Closely related weakly asymmetric interfaces use the critical scale

pc(τ1,τ2)\mathbf{p}_c(\tau_1,\tau_2)5

which yields invariant measures tilted by continuum area functionals and dynamical limits given by stochastic heat equations or reflected stochastic heat equations with additive drift pc(τ1,τ2)\mathbf{p}_c(\tau_1,\tau_2)6 (Etheridge et al., 2014).

Asymmetric trap models exhibit a different scaling mechanism. For the finite complete graph, the rescaled trap-depth process is

pc(τ1,τ2)\mathbf{p}_c(\tau_1,\tau_2)7

while the small-time scaling for the K-process is

pc(τ1,τ2)\mathbf{p}_c(\tau_1,\tau_2)8

valid only in the regime pc(τ1,τ2)\mathbf{p}_c(\tau_1,\tau_2)9. Here the asymmetry parameter Bayes(N,D)NαN+DαD\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-\alpha_N}+D^{-\alpha_D}0 modifies the effective stable index of the limit process (Bezerra et al., 2012).

Large-deviation theory supplies perhaps the clearest statistical meaning of asymmetric scaling. For the empirical average with stretched exponential tails and Bayes(N,D)NαN+DαD\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-\alpha_N}+D^{-\alpha_D}1,

Bayes(N,D)NαN+DαD\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-\alpha_N}+D^{-\alpha_D}2

Below the typical value, the cost is collective and extensive; above it, the cost is a one-big-jump mechanism. The same logic extends to non-integer empirical moments (Monthus, 2019).

5. Anisotropic gravity and parity-dependent renormalization

In gravitational thermodynamics, anisotropic scaling means

Bayes(N,D)NαN+DαD\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-\alpha_N}+D^{-\alpha_D}3

For charged Lifshitz black holes, the entropy is organized by a generalized Cardy formula involving the black-hole energy Bayes(N,D)NαN+DαD\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-\alpha_N}+D^{-\alpha_D}4, the ground-state energy Bayes(N,D)NαN+DαD\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-\alpha_N}+D^{-\alpha_D}5, the electric and magnetic charge sectors, and a theory-dependent coefficient Bayes(N,D)NαN+DαD\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-\alpha_N}+D^{-\alpha_D}6. The same Bayes(N,D)NαN+DαD\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-\alpha_N}+D^{-\alpha_D}7 appears in the Smarr relation

Bayes(N,D)NαN+DαD\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-\alpha_N}+D^{-\alpha_D}8

which in the three-dimensional Lifshitz case becomes

Bayes(N,D)NαN+DαD\ell_{\mathrm{Bayes}}(N,D)\asymp N^{-\alpha_N}+D^{-\alpha_D}9

For hyperscaling violation, the effective dimensionality is

Πsβeγm\Pi_s \sim -\beta e^{-\gamma m}0

the entropy scales as

Πsβeγm\Pi_s \sim -\beta e^{-\gamma m}1

and the same structural formula survives with Πsβeγm\Pi_s \sim -\beta e^{-\gamma m}2 replacing the naive spatial dimension (Bravo-Gaete et al., 2015).

A very different parity-sensitive asymmetric scaling law appears in strongly asymmetric unimodal maps. Near the critical point,

Πsβeγm\Pi_s \sim -\beta e^{-\gamma m}3

Πsβeγm\Pi_s \sim -\beta e^{-\gamma m}4

The renormalization intervals Πsβeγm\Pi_s \sim -\beta e^{-\gamma m}5 satisfy

Πsβeγm\Pi_s \sim -\beta e^{-\gamma m}6

and the scaling alternates with parity. For large even Πsβeγm\Pi_s \sim -\beta e^{-\gamma m}7,

Πsβeγm\Pi_s \sim -\beta e^{-\gamma m}8

whereas for large odd Πsβeγm\Pi_s \sim -\beta e^{-\gamma m}9,

RK(X)RK(12C)×(factor)×C(X)R^K(X)\approx R^K({}^{12}\mathrm{C})\times(\text{factor})\times\mathcal C(X)0

The lengths of the renormalization intervals decay super-exponentially, with

RK(X)RK(12C)×(factor)×C(X)R^K(X)\approx R^K({}^{12}\mathrm{C})\times(\text{factor})\times\mathcal C(X)1

The paper emphasizes that this scaling is non-universal, because RK(X)RK(12C)×(factor)×C(X)R^K(X)\approx R^K({}^{12}\mathrm{C})\times(\text{factor})\times\mathcal C(X)2 depends on the map (Kozlovski et al., 2019).

6. Interpretation, scope, and non-universality

The cited literature does not support treating asymmetric scaling as a single phenomenon. In some works, asymmetry is imposed at the level of the scaling map itself, as in ATS. In others, it emerges from hidden coverage constraints, as in sparse features, from symmetry-breaking geometry, as in tokamaks, from proton-neutron imbalance and channel counting, as in 2p2h MEC scaling, or from different fluctuation mechanisms above and below a typical value, as in large deviations (Li et al., 2022, Sous et al., 22 May 2026, Ball et al., 2016, Martinez-Consentino et al., 10 Sep 2025, Monthus, 2019).

The literature also distinguishes sharply between universal and non-universal asymmetry. Oblique dependence axes in linear random fields force the universal critical exponent RK(X)RK(12C)×(factor)×C(X)R^K(X)\approx R^K({}^{12}\mathrm{C})\times(\text{factor})\times\mathcal C(X)3 (Pilipauskaitė et al., 2020). By contrast, the charge-sector coefficient RK(X)RK(12C)×(factor)×C(X)R^K(X)\approx R^K({}^{12}\mathrm{C})\times(\text{factor})\times\mathcal C(X)4 in charged Lifshitz black holes depends on the electromagnetic theory, the reduced coefficients RK(X)RK(12C)×(factor)×C(X)R^K(X)\approx R^K({}^{12}\mathrm{C})\times(\text{factor})\times\mathcal C(X)5 in asymmetric nuclear scaling are model dependent, and the super-exponential rate RK(X)RK(12C)×(factor)×C(X)R^K(X)\approx R^K({}^{12}\mathrm{C})\times(\text{factor})\times\mathcal C(X)6 in asymmetric unimodal maps depends on the map (Bravo-Gaete et al., 2015, Martinez-Consentino et al., 10 Sep 2025, Kozlovski et al., 2019).

A further point is that asymmetry need not coincide with the mere breaking of an obvious discrete symmetry. The tokamak result is explicit: mirror-symmetric flux surface shaping can be up-down asymmetric and yet produce momentum flux that is exponentially small in large shaping mode number, while non-mirror symmetric shaping yields a power law (Ball et al., 2016). This suggests that the decisive issue is often which terms survive averaging, normalization, or coarse graining, not simply whether a configuration looks asymmetric in a geometric sense.

Across fields, asymmetric scaling formulas therefore function as compact summaries of inequivalent mechanisms. They mark cases in which a single symmetric exponent or prefactor is insufficient, and in which the scaling description must remember direction, channel, parity, resource regime, or fluctuation side.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Asymmetric Scaling Formula.