---
title: Double Smoothly Broken Power Law
url: https://www.emergentmind.com/topics/double-smoothly-broken-power-law
type: topic
---

# Double Smoothly Broken Power Law

Searching arXiv for the cited works to ground the article in the current literature.
A double smoothly broken power law is a family of parameterizations in which a quantity follows one approximate power-law regime, then passes through two continuous transitions into later regimes, yielding three asymptotic segments on log-log axes without hard piecewise breaks. In the literature, the same modeling role appears in several closely related forms: an explicit multi-break product formula with two smooth breakpoints; a sum of two smooth single-break spectra; and rank-size constructions in which two power-law-like factors or two Lavalette-like terms generate a smooth bend, an intermediate shoulder, or an inflection point on a log-log plot [2210.14891][1404.3605][2211.13111].

## 1. Formal structure and parameterization

An explicit multi-break smoothly broken power law is given by the Broken Neural Scaling Law family,
\[
y = a + b x^{-c_0}\prod_{i=1}^{n}\left(1+\left(\frac{x}{d_i}\right)^{f_i}\right)^{-c_i/f_i}.
\]
In this parameterization, \(n\) is the number of smooth breaks or transitions, \(a\) is the limiting asymptote as \(x\to\infty\), \(b\) is an offset or normalization, \(c_0\) is the slope of the first segment on a log-log plot, \(c_i\) is the change in slope between segment \(i\) and \(i+1\), \(d_i\) is the breakpoint location on the \(x\)-axis, and \(f_i\) is the smoothness or sharpness of the \(i\)-th transition. Smaller nonnegative \(f_i\) gives a sharper break, while larger \(f_i\) gives a smoother or wider transition [2210.14891].

For a double smoothly broken power law in the strict sense, one sets \(n=2\):
\[
y = a + b x^{-c_0} \left(1+\left(\frac{x}{d_1}\right)^{f_1}\right)^{-c_1/f_1} \left(1+\left(\frac{x}{d_2}\right)^{f_2}\right)^{-c_2/f_2}.
\]
On log-log axes, this corresponds to a curve that is nearly linear in three regions, with smooth transitions around \(x=d_1\) and \(x=d_2\). The appendix decomposition in the same work makes the slope accumulation explicit: successive breaks add the slope changes \(c_i\), so later segments inherit all earlier transitions [2210.14891].

A related one-break building block appears in stochastic gravitational-wave searches as the smooth broken power law
\[
\Omega_\mathrm{BPL}(f; \bm{\theta}) = \Omega_* \left( \frac{f}{f_*} \right)^{n_1} \left[\frac{1+(f/f_*)^\Delta}{2}\right]^{(n_2 - n_1) / \Delta},
\]
with \(\bm{\theta}\equiv(\Omega_*, f_*, n_1, n_2, \Delta)\). Here \(\Omega_*\) is the peak amplitude at \(f_*\), \(f_*\) is the peak frequency, \(n_1\) and \(n_2\) are the low- and high-frequency slopes, and \(1/\Delta\) is the transition width. This single-break form becomes a component from which more complex spectra can be built, either by summation or by embedding it in a larger doubly broken ansatz [2211.13111].

## 2. Rank-size and generalized Lavalette constructions

A precursor to double smoothly broken behavior appears in rank-size modeling through the Lavalette family. The basic 2-parameter Lavalette law is
\[
y(r)=\kappa\;\Big[\frac{N\,r}{N-r+1}\Big]^{-\chi},
\]
or equivalently
\[
y(r)=\hat{\kappa}\;r^{-\chi}(N-r+1)^{+\chi}.
\]
This is a power-law decay multiplied by a power-law cut-off. On a log-log plot it looks like a descending scaling law at low rank, but it bends down near the upper end because of the \((N-r+1)^{+\chi}\) factor. The parameter \(\chi\) controls the common steepness, and the finite-size factor \(N\) sets where the cut-off happens; in practice, \(N\) is often just a convenient scale factor rather than a physically meaningful fitted quantity. This already plays the role of a smoothly broken power law because the transition between regimes is continuous rather than piecewise [1404.3605].

The generalized 3-parameter Lavalette form,
\[
y(r)=\Lambda\,[r]^{-\phi}[N-r+1]^{-\psi} \equiv \hat{\Lambda}\;u^{-\phi}(1-u)^{+\psi},\qquad u=\frac{r}{N+1},
\]
separates the two tails. Here \(\phi\) governs the low-rank behavior and \(\psi\) governs the high-rank behavior, so the low and high tails can have distinct slopes. If \(\psi=0\), it reduces to a Zipf or pure power-law-like behavior; if \(\phi>0\) and \(\psi<0\), one gets a Feller-Pareto-type form; if \(\phi=-1\) and \(\psi=+1\), one recovers the Verhulst or logistic structure. Additional 3- and 4-parameter variants introduce a free high-rank range parameter \(N_1\) and low-rank Mandelbrot shifts \(m\), explicitly to handle flattening at the low-rank end and or high-rank end [1404.3605].

The closest analogue in that paper to a double smoothly broken power law is the proposal to generate an inflection point on a log-log plot by summing two decreasing power-law-like terms and then imposing a high-\(x\) cut-off:
\[
y(x)=\big[A(x+m_5)^{-m_1}+B(x+m_6)^{-m_2}\big]\,(N+m_4-x^{m_7})^{m_3},
\]
or, with an exponential cut-off,
\[
y(x)=\big[A(x+m_5)^{-m_1}+B(x+m_6)^{-m_2}\big]\,e^{-m_3(\,x+\cdots\,)}.
\]
The stated intent is that the sum of two decreasing power-law terms gives an intermediate curvature change, while the multiplicative cut-off handles the final high-\(x\) suppression. The author explicitly says that such a law can produce the needed inflection point on a log-log plot. Hyper-generalized forms with nonlinear powers of \(r\) and \(N-r\) are also written down, but the paper repeatedly states that extra parameters should be considered only if simpler models fail and judged empirically rather than by aesthetic preference [1404.3605].

## 3. Physical realizations in energetic particles and gravitational waves

In numerical modeling of energetic particles accelerated at coronal shocks, event-integrated energy spectra over the whole simulation domain are reported to be well described by a double power law for all ion species. The study solves the Parker transport equation in a streamer-like coronal magnetic field, includes protons and the heavier ions He, O, Mg, and Fe, and fits the integrated spectra with the Band function, i.e. a smoothly joined broken power law. In the Kolmogorov case, the fitted low-energy indices are around \(\gamma_1\approx 1.29\text{--}1.30\), the high-energy indices are around \(\gamma_2\approx 2.35\text{--}2.42\), and the break energy follows
\[
E_B \sim \left(\frac{Q}{A}\right)^\alpha,
\]
with \(\alpha \approx 0.51\) for Kolmogorov turbulence and spanning roughly \(0.16\) to \(1.2\) across different turbulence spectral indices. The low-energy slope is consistent with the DSA expectation for a shock compression ratio \(X=3\), while the second power-law segment is attributed to spatially inhomogeneous acceleration and transport [2201.06712].

The proposed mechanism in that context is superposition of particle populations from different source regions along the shock front. In the streamer and nonstreamer regions individually, spectra are best fit by a power law with an exponential rollover, but in the transition region the spectrum becomes double-power-law-like because it contains a mix of locally accelerated particles and higher-energy particles that have diffused or escaped from the more efficient streamer region. The paper explicitly demonstrates that combining two rollover spectra with different characteristic energies can produce a composite spectrum that resembles a double power law, and that mixing a small fraction of high-energy streamer particles with a locally accelerated spectrum can reproduce the transition-region spectrum. A plausible implication is that a double smoothly broken power law can emerge from spatial superposition rather than from a single local acceleration law [2201.06712].

In stochastic gravitational-wave background searches, the distinction between related constructions is made explicit. A double-peak spectrum is modeled as the sum of two separate broken power laws,
\[
\Omega_\mathrm{DP}(f;\{\bm{\theta}_1,\bm{\theta}_2\})= \Omega_\mathrm{BPL}(f;\bm{\theta}_1)+\Omega_\mathrm{BPL}(f;\bm{\theta}_2),
\]
whereas the doubly broken spectrum is one spectrum with three power-law segments and two break frequencies \(f_l\) and \(f_h\). Its asymptotic behavior is written as
\[
\Omega_\mathrm{DB}(f)\propto
\begin{cases}
(f/f_l)^{n_l}, & f\ll f_l,\\
(f/f_l)^{n_m}, & f_l\ll f\ll f_h,\\
(f_h/f_l)^{n_m}(f/f_h)^{n_h}, & f\gg f_h.
\end{cases}
\]
Using Advanced LIGO-Virgo O1-O3 data, the analysis finds strong negative evidence for the double-peak case and weak negative evidence for the doubly broken case, placing \(95\%\) C.L. upper limits \(\Omega_\mathrm{BPL,1}<5.8\times10^{-8}\), \(\Omega_\mathrm{BPL,2}<4.4\times10^{-8}\), and \(\Omega_\mathrm{DB}<1.2\times10^{-7}\) relative to unresolved compact binary coalescence backgrounds. In that literature, the doubly broken template is closely related in spirit to a double smoothly broken power law, but it is not parameterized as two explicit smooth breaks of the same type as the single-BPL formula [2211.13111].

## 4. Multi-break scaling laws in machine learning

Broken Neural Scaling Laws provide the most explicit general-purpose multi-break formulation among the cited works. They are introduced to model how the evaluation metric of interest varies with amount of compute used for training or inference, number of model parameters, training dataset size, model input size, number of training steps, or upstream performance, across various architectures and tasks in zero-shot, prompted, and finetuned settings. The same functional family is stated to model behaviors that simpler laws cannot express, including delayed sharp inflection points, multiple regime changes, and nonmonotonic transitions such as double descent [2210.14891].

The paper is explicit about the hierarchy of special cases. When \(n=0\),
\[
y = a + b x^{-c_0},
\]
which is a plain power law. When \(n=1\), the model is a single smoothly broken power law. In general, \(n\) can be any nonnegative integer, giving \(n\) breaks and therefore \(n+1\) approximately linear segments on a log-log plot. A double smoothly broken power law is therefore not a special construction added afterward; it is the \(n=2\) member of the family [2210.14891].

The same work emphasizes expressivity beyond monotone scaling. It states that M1, M2, and M3 are monotonic and cannot model double-descent-like behavior, while M4 can model an inflection point but still cannot represent nonmonotonicity. BNSL can do both. Specific examples include delayed sharp inflection points in 4-digit arithmetic addition, temporary worsening followed by improvement in transformer double descent, and late transitions from near-random to strong performance on some downstream tasks [2210.14891].

Empirically, BNSL is reported to be best on \(69.44\%\) of tasks on the scaling laws benchmark for downstream vision tasks and on \(75\%\) of tasks on the language benchmark. The paper also states that extrapolations can extend beyond the largest training point by more than an order of magnitude and, in some GPT-4-related plots, by over \(100{,}000\times\) the largest fitted scale. The associated conceptual point is not only that multi-break smoothly broken power laws are expressive, but also that they expose intrinsic limits to predictability when a sharp break lies beyond the observed range [2210.14891].

## 5. Related but distinct constructions

Several adjacent model classes resemble double smoothly broken power laws but are not equivalent to a two-break smooth interpolation. The distinctions matter because the same qualitative language—double power law, doubly broken, double-sided power law, broken peak—can refer to different mathematical objects [1910.05364][2407.12914][2211.13111].

| Construction | Break structure | Representative source |
|---|---|---|
| BNSL with \(n=2\) | One formula, two smooth breaks, three segments | [2210.14891] |
| Double-peak spectrum | Sum of two smooth single-BPL spectra | [2211.13111] |
| Doubly broken SGWB spectrum | One spectrum with three asymptotic power-law regimes and two break frequencies | [2211.13111] |
| Lavalette-based inflection model | Sum of two power-law-like terms times a cut-off | [1404.3605] |
| Beta Rank Function | Smooth double-sided power law in log-space, single rounded peak | [1910.05364] |
| Broken power-law peak for SIGWs | One smooth peak with IR and UV power-law tails | [2407.12914] |

The Beta Rank Function,
\[
x(u)=A\frac{(1-u)^b}{u^a},
\]
is a continuous-rank analogue of the discrete generalized beta distribution. Its induced density is approximately a unimodal skewed and asymmetric two-sided power law or double Pareto or log-Laplacian distribution, but after the log transform it makes a smooth turn at the peak and does not diverge, lacking the sharp angle observed in the double Pareto or Laplace distribution. The left and right log-space tails are approximately exponential, with the left slope controlled by \(b\) and the right slope by \(a\). When \(a=b\), the BRF becomes the Lavalette distribution and is closely related to a lognormal-like shape. This is therefore a smooth double-sided power law, not a double smoothly broken power law with two distinct break locations [1910.05364].

A similar terminological caution applies to scalar-induced gravitational waves from a broken power-law curvature spectrum. The spectrum
\[
\mathcal{P}_{\rm PL}(k) = A\,\frac{\alpha+\beta}{\beta (k/k_\ast)^{-\alpha}+\alpha (k/k_\ast)^\beta}
\]
has a peak at \(k=k_\ast\), an infrared branch \(\mathcal{P}_{\mathcal R}\propto k^{\alpha}\), and an ultraviolet branch \(\mathcal{P}_{\mathcal R}\propto k^{-\beta}\). The paper explicitly states that this is not a double-break spectrum in the strict sense of two independent break scales; it is a single smooth broken power law with one peak and two asymptotic branches. Its utility lies in near-peak approximation and in the induced SIGW spectrum, whose IR tail, central peak, and UV tail carry distinct signatures of \(\alpha\), \(\beta\), and peak width [2407.12914].

## 6. Estimation, interpretability, and structural constraints

Across the cited literature, increasing the number of parameters improves flexibility but weakens interpretability. In the Lavalette study, the added parameters are said to be “understood” in a pragmatic fit sense, but their theoretical meaning is often unclear; the paper cautions that such parameters should not be added indiscriminately. It also warns against reporting parameter values too precisely, since these are nonlinear fits and the regression coefficient \(R^2\) can be misleadingly high even when the visual fit is unsatisfactory in one tail. The Levenberg–Marquardt algorithm is mentioned for nonlinear fitting, and the recommendation is to use low-rank and high-rank parameters to target visible deviations where the simple power law fails [1404.3605].

In neural scaling, the limitation is epistemic rather than merely statistical. The BNSL analysis states three implications: very sharp breaks impose a lower bound on how far one can extrapolate from pre-break data; breaks beyond the fitted range may be impossible to predict reliably; and points before a sharp early break can sometimes be unhelpful or even insufficient, but in other settings they are crucial for identifying the right post-break behavior. This makes a double smoothly broken power law not only a descriptive fit family but also a statement about the difficulty of inferring unseen regimes from small-scale observations [2210.14891].

A further limitation arises in network theory. For degree sequences distributed as a double power law with a size-dependent crossover, graphicality is governed by the Erdős–Gallai theorem and related scaling arguments. The cited work derives the graphicality of infinite sequences for all possible values of the degree exponents \(\gamma_1\) and \(\gamma_2\), finds five qualitatively distinct ways graphicality can be violated, and emphasizes that finite-size convergence toward the asymptotic phase diagram is logarithmically slow. It also states that smoothly broken or double power-law network models are subject to the same structural feasibility constraints as sharp double power laws because the asymptotic issue is governed by the scaling of the tail exponents and the crossover scale. This suggests that functional smoothness alone does not guarantee realizability: a plausible fit can still correspond to no simple graph [2512.24976].

Taken together, these results define the double smoothly broken power law less as a single universal equation than as a modeling class for multi-regime scaling. Its essential features are continuous transitions between asymptotic slopes, the possibility of distinct low-, intermediate-, and high-range behavior, and the capacity—when formulated by superposition or multi-break interpolation—to represent shoulders, delayed inflection points, nonmonotonic transitions, and species- or scale-dependent cut-offs. Its main difficulties are equally consistent across fields: parameter proliferation, ambiguity of interpretation, sensitivity of extrapolation to unseen breaks, and, in some applications, structural constraints external to the fit itself.

Source: https://www.emergentmind.com/topics/double-smoothly-broken-power-law