---
title: 'Super-Linear: Scaling, Convergence & Applications'
url: https://www.emergentmind.com/topics/super-linear
type: topic
---

# Super-Linear: Scaling, Convergence & Applications

Super-linear denotes behavior that exceeds linear order with respect to a reference variable, but the technical meaning is domain-dependent. In the cited literature it appears as faster-than-linear scaling laws such as $X \propto P^\gamma$ with $\gamma>1$, nonlinear input–output response laws such as $I_{em} \propto I_{exc}^s$ with $s>1$, coefficient-growth conditions in ODEs, SDEs, BSDEs, and BSVIEs, convergence regimes satisfying $\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 0$, and structural bounds such as code lengths or circuit sizes that exceed linear order. The term also names a specific forecasting architecture, "Super-Linear," built as a lightweight pretrained mixture of linear experts for time series forecasting [0809.4994][2403.13436][2009.00129][2509.15105].

## 1. Core meanings and formal criteria

Across the literature, “super-linear” is not a single definition but a family of related order relations. In scaling-law settings, it means an exponent strictly larger than $1$, as in $X \propto P^\gamma$ with $\gamma>1$ [1212.4914]. In optical response, it denotes a power-law slope $s>1$ in the excitation–emission relation $I_{em} \propto (I_{exc})^s$ [2403.13436]. In numerical analysis and optimization, it means convergence faster than linear, formally $\|x_{k+1}-x^*\|/\|x_k-x^*\| \to 0$ as $k\to\infty$ [2009.00129]. In nonlinear dynamics, a function $f$ may be called superlinear when $f(x)/x$ is ultimately increasing and $\lim_{x\to\infty} f(x)/x = \infty$ [1702.06427].

| Context | Canonical form | Meaning |
|---|---|---|
| Scaling laws | $X \propto P^\gamma$, $\gamma>1$ | Output grows faster than proportionally |
| Optical response | $I_{em} \propto I_{exc}^s$, $s>1$ | Emission localizes more sharply than linear response |
| Convergence theory | $\|x_{k+1}-x^*\|/\|x_k-x^*\| \to 0$ | Faster-than-linear iteration |
| Nonlinear dynamics | $f(x)/x \to \infty$ | State-dependent forcing exceeds linear growth |

A recurring misconception is to equate super-linear with quadratic. The cited numerical papers explicitly distinguish the two: quadratic convergence is a special case, whereas super-linear convergence only requires asymptotically faster-than-linear decay of error [2009.00129][2403.11115]. A second misconception is to equate super-linearity with monotone growth in time. The ODE literature shows that superlinear systems may exhibit finite-time blow-up, while other superlinear dissipative systems decay algebraically rather than exponentially [1702.06427][2211.00209].

## 2. Super-linear scaling in collective systems

In urban science, superlinear scaling refers to sociological quantities such as economic productivity, creative output, patents, inventors, and GDP increasing faster than city population. Reported empirical exponents are typically between $1$ and $1.5$, with a mean around $1.2$ [0809.4994]. The network model proposed for this phenomenon places individuals on the leaves of a hierarchical tree, defines social distance $d$ as the height of the lowest common ancestor, assigns connection probability proportional to $b^{-\alpha d}$, counts $b^d$ individuals at distance $d$, and assumes a productivity benefit per tie proportional to $b^{\beta d}$. This yields
$$
P(N)=N\sum_{d=1}^{\log_b N} b^{-\alpha d}b^d b^{\beta d}
= N\sum_{d=1}^{\log_b N} b^{(\beta-\alpha+1)d},
$$
and for large $N$ with $\alpha<1+\beta$,
$$
P(N)\sim N^{\beta-\alpha+2}.
$$
The mechanism is the proliferation of socially distant links, interpreted as productive “weak ties,” so that larger cities create more opportunities for novelty and creative collaboration [0809.4994].

A geometrically generative account appears in growing random geometric graph models, where new nodes survive only if they are placed within distance $r$ of existing nodes and then connect to all existing nodes within radius $r$. In that setting the total number of edges obeys a super-linear power law $X \propto P^\gamma$, and the geometric dimension $d$ is the primary parameter controlling $\gamma$ [1212.4914]. Simulations reported $\gamma \approx 1.78$ for $d=1$, $\gamma \approx 1.57$ for $d=2$, and $\gamma \approx 1.22$ for $d=3$, while the same framework also reproduced fractal growth, asymptotically size-invariant clustering coefficient, and sub-linear area–population and diversity–population relations [1212.4914].

In tumor ecology, super-linear growth is written as
$$
\frac{dC_T}{dt} \propto C_T^\beta,
$$
with $\beta>1$ defining the super-linear regime [2104.00079]. The cited model attributes this regime not to competition or fitter subclones alone, but to tumor–microenvironment interaction through angiogenesis. In the oxygen dynamics,
$$
\frac{\partial n(r,t)}{\partial t}
= D_n \nabla^2 n(r,t)-a_n n(r,t)+b_{n0}+b_{n1} C_T-c_n C_t(r,t)n(r,t),
$$
the term $b_{n1}C_T$ increases oxygen supply in proportion to tumor mass, producing positive feedback compatible with an Allee effect [2104.00079]. The same paper states that recent empirical work found average tumor growth exponents $\langle \beta \rangle = 1.25$ across various human cancers [2104.00079].

Robot learning supplies a further operational meaning. The CASHER pipeline reports super-linear scaling with human effort by crowdsourcing digital twins of real scenes, collecting behavioral data in simulation, and gradually replacing human demonstrations with model-generated demonstrations as a generalist policy improves [2412.01770]. The authors state that required human demonstrations per environment decrease as the number of environments grows, while zero-shot and few-shot scaling laws are demonstrated on three real-world tasks [2412.01770]. A plausible commonality across cities, networks, tumors, and CASHER is that super-linear scaling emerges when larger system size increases the density or efficacy of productive interactions faster than it increases the relevant cost base.

## 3. Super-linear optical response and super-resolution

In fluorescence microscopy, super-linearity is used in a response-law sense. Super-linear image scanning microscopy extends conventional ISM by exploiting nonlinear upconversion emission from lanthanide-doped UCNPs, with
$$
I_{em} \propto (I_{exc})^s.
$$
When $s>1$, emission is more tightly localized at the excitation center, narrowing the emission PSF and pushing resolution beyond the twofold ISM limit [2403.13436]. In the reported implementation, NaYF\(_4\) doped with $20\%$ Yb\(^{3+}\) and $8\%$ Tm\(^{3+}\) was excited at $976$ nm by a single low-power continuous-wave near-infrared laser. The five-photon $455$ nm emission reached $s \approx 4.5$ at about $0.1$ mW, producing measured FWHM $\sim 124$ nm, or $\lambda/8$ for $\lambda=976$ nm; Fourier ring correlation gave about $116$ nm [2403.13436]. The same work also reported a multifocal structured-excitation variant with roughly $20\times20$ foci, a field of view of $20\,\mu\mathrm{m} \times 20\,\mu\mathrm{m}$, and frame rates up to $1$ Hz [2403.13436].

A distinct optical mechanism appears in bistable scattering from nano-silicon Mie resonators. There, photo-thermo-optical feedback produces optical bistability in a silicon resonator with volume size $10^{-3}\,\mu\mathrm{m}^3$ and $Q$-factor $<10$, and the bistable transition yields a large effective super-linear scattering–excitation law with measured slope $p \approx 10$ [2307.01490]. For a Gaussian excitation beam, the cited analysis gives
$$
I_{\text{em}}(x)\propto \left[I_{\text{exc}}(x)\right]^p,
$$
so the emission PSF narrows by a factor $1/\sqrt{p}$ [2307.01490]. Experimentally, diffraction-limited laser scanning microscopy with FWHM $\approx 500$ nm was sharpened to about $140$ nm, a resolution enhancement of more than $3$ times [2307.01490].

These two optical lines use different nonlinearities. UCNP-based SL-ISM relies on multiphoton upconversion with experimentally observed $s \approx 4.5$, whereas nano-silicon bistable scattering relies on thermally driven resonance shifts and hysteresis, reaching $p \sim 10$ near the transition [2403.13436][2307.01490]. The shared consequence is PSF compression through a super-linear emission law rather than through purely linear optical transfer.

## 4. Differential, stochastic, and backward equations

For forced ODEs,
$$
x'(t)=f(x(t))+h(t), \qquad x(0)=\psi>0,
$$
superlinearity is imposed by requiring $f$ to be continuous, positive, increasing on $(0,\infty)$, with $f(x)/x$ ultimately increasing and $\lim_{x\to\infty} f(x)/x=\infty$ [1702.06427]. Defining
$$
F(x)=\int_1^x \frac{du}{f(u)},
$$
finite-time blow-up occurs if $F(\infty)<\infty$; when $F(\infty)=\infty$, the forcing–nonlinearity competition can be classified sharply. If $\limsup_{t\to\infty} F(H(t))/t \in [0,1]$, then $F(x(t))/t \to 1$. If the limit superior equals $K\in(1,\infty)$, then $\limsup_{t\to\infty} F(x(t))/t = K$. Under an additional negligibility condition, $x(t)/H(t)\to 1$ [1702.06427]. Thus “superlinear” in the state variable does not determine the asymptotic regime by itself; the forcing scale matters.

A different use appears in genuinely nonlinear dissipative systems
$$
y'=-H(y)Ay+G(t,y),
$$
where $H$ is positively homogeneous of degree $a>0$ and positive away from the origin [2211.00209]. Here the principal effect of superlinearity is not explosion but non-exponential decay. For sufficiently small initial data, nontrivial decaying solutions satisfy
$$
y(t)=\xi\, t^{-1/a}+O(t^{-1/a-\epsilon}),
$$
with $\xi\neq 0$ an eigenvector of $A$ satisfying $aH(\xi)\lambda=1$ for the corresponding eigenvalue $\lambda$ [2211.00209]. The contrast with linear theory is explicit: decay is algebraic, not exponential [2211.00209].

The SDE and SFDE literature uses “super-linear” chiefly as a coefficient-growth condition. For multidimensional SDEs with non-Lipschitz coefficients, the local logarithmic hypothesis
$$
\|\sigma(x)-\sigma(y)\| \leq C \sqrt{\log N}\,|x-y| + C\frac{\log N}{N^p}, \qquad
|b(x)-b(y)| \leq C \log N\, |x-y| + C\frac{\log N}{N^p}
$$
on each ball $B(N)$ yields pathwise uniqueness, non-contact, a stochastic flow of continuous maps, and a Freidlin–Wentzell-type large deviations principle [1502.04915]. For super-linear SFDEs, an explicit truncated Euler–Maruyama scheme with linear interpolation achieves boundedness, strong convergence in $L^p$, convergence rate $1/2$, and preservation of exponential stability, without requiring global Lipschitz continuity of the diffusion coefficient [2208.10214]. For SDEs with superlinearly growing drift and diffusion coefficients, explicit Milstein schemes based on tamed coefficients converge in $\mathcal L^p$ with the optimal strong rate $1.0$ under mild assumptions [1601.02695].

Backward equations sharpen the threshold structure. Multi-dimensional BSVIEs with generators that are diagonally strictly quadratic in $Z$ and sub-quadratically coupled off-diagonally admit unique adapted solutions for bounded free term; when the free term is unbounded but has exponential moments of arbitrary order, unique solvability persists only in the diagonal at-most-quadratic case [2211.04078]. The same paper presents negative results for super-quadratic growth in $Z$: in general, even bounded free term does not guarantee bounded solutions [2211.04078]. Scalar BSDEs with generator growth
$$
|y||\ln |y||^{(\lambda+1/2)\wedge 1}+|z||\ln |z||^\lambda
$$
exhibit four different integrability thresholds for the terminal condition, according to $\lambda=0$, $\lambda\in(0,1/2)$, $\lambda=1/2$, and $\lambda>1/2$; comparison and uniqueness follow when one generator is convex or concave in $(y,z)$, or when it satisfies a one-sided Osgood condition in $y$ and uniform continuity in $z$ [2107.12694]. In this branch of the literature, “super-linear” therefore marks a solvability frontier rather than a uniform dynamical effect.

## 5. Convergence, speedup, and lower bounds

In iterative computation, super-linear often refers to convergence rate. The refined $p$-adic QR algorithm defines super-linear convergence by the standard criterion $\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 0$ and proves a stronger block-deflation estimate under eigenvalue separation: if a size-sorted Hessenberg matrix has suitably separated eigenvalues in the small block, then after the corresponding QR cycle
$$
|\epsilon'| \leq |\epsilon|^2.
$$
The resulting convergence is essentially quadratic in many cases, while the algorithm falls back to linear behavior when the favorable separation structure fails [2009.00129]. When the characteristic polynomial modulo $p$ is square-free and splits completely, the paper states that all eigenvalues can be obtained up to error $O(p^N)$ in at most
$$
\frac{1}{2} n^3 \log_2 N + o(n^3 \log_2 N)
$$
arithmetic operations at $N$ $p$-adic digits of precision [2009.00129].

Optimization papers use the same convergence terminology but different mechanisms. “Superlinear Optimization Algorithms” proposes trajectory-inspired updates for minimizing nonlinear objectives, with several variants remaining applicable when the Hessian is singular [2403.11115]. One family avoids calculating the inverse of the Hessian matrix or an identical-dimension matrix; another requires only the diagonal elements of the Hessian; all are reported to be superlinear convergent when appropriate parameters are selected [2403.11115]. In mixed linear regression, alternating minimization is shown to contract estimation error super-linearly under proper initialization. The main recurrence is of the form
$$
\mathrm{dist}_{t+1} \lesssim C\, \mathrm{dist}_t^{3/2},
$$
which yields $\mathcal O(\log\log(1/\epsilon))$ iterations to reach $\epsilon$-accuracy, with a quadratic regime in a narrower neighborhood of the optimum [2004.10914].

Program transformation and computational complexity supply two further meanings. Repeated recursion unfolding repeatedly unfolds a recursive rule with itself, so that each unfolding doubles the number of recursive steps covered; with optimal rule application, the runtime recurrence changes from
$$
r(n)=c(n)+r(n-1)
$$
to
$$
r'(n)=c(n)+r'(n/2),
$$
and, in the best case, the method lowers time complexity class within a chosen bound on recursion depth [2009.05314]. The paper explicitly lists examples such as quadratic becoming linear and linear becoming constant time [2009.05314]. By contrast, threshold-circuit complexity uses “super-linear” comparatively: the paper on depth-two and depth-three threshold circuits proves the first super-linear gate lower bounds and the first super-quadratic wire lower bounds for explicit functions. For Andreev’s function, any depth-two linear threshold circuit agreeing on a $(1/2+\epsilon)$-fraction of inputs requires at least $\Omega(\epsilon^3 n^{3/2}/\log^3 n)$ gates or $\Omega(\epsilon^3 n^{5/2}/\log^{7/2} n)$ wires, while PARITY has tight average-case complexity $\Theta(\sqrt n)$ gates and $\Theta(n^{3/2})$ wires in this setting [1511.07860]. This usage is about lower-bound magnitude, not iterative improvement.

## 6. Named architectures and broader distinctions

“Super-Linear” is also the title of a time-series forecasting model: a lightweight pretrained mixture-of-experts architecture built from frequency-specialized linear experts and a spectral router [2509.15105]. Given input sequence $X_{1:L}$ and forecast $Y_{1:H}$, the model writes the prediction as
$$
\hat{Y} = \sum_{i=1}^{N} \mathrm{diag}(G_i(X)) F_i(X),
$$
where $F_i$ are linear experts and the gating network is driven by normalized spectral features,
$$
g(X) = \frac{I(X-\bar{X})}{\|I(X-\bar{X})\|_1} W_g + b_g,
$$
followed by sparse Top-$k$ softmax selection [2509.15105]. Training proceeds in two stages: independent pretraining of experts on aggressively resampled data to match target frequency regimes, then freezing those experts while jointly training the router and complementary experts [2509.15105].

The reported empirical profile is explicitly tied to efficiency. Super-Linear uses about $2.5$M parameters, trains and runs on a single GPU, and is described as much smaller than Timer-XL, at $3\%$ of its size [2509.15105]. On the LTSF zero-shot benchmark it reports average MSE reductions of $26.2\%$ relative to Chronos, $14.2\%$ relative to TimesFM, $8.5\%$ relative to Moirai, and $7.1\%$ relative to Time-MoE Large, while on GIFT-Eval it reports a $16\%$ MASE reduction over the lightweight TTM model [2509.15105]. The “super” in the model name is therefore nominal, but the paper explicitly ties that architecture to robustness across sampling rates, sparse interpretability through spectral routing, and strong zero-shot/full-shot performance [2509.15105].

Coding theory uses the term in yet another strictly order-theoretic sense. For optimal locally repairable codes with all-symbol $(r,\delta)$-locality, the paper derives alphabet-dependent upper bounds on length and constructs order-optimal codes whose length is super-linear in the alphabet size [1812.11942]. With
$$
t=\left\lfloor \frac{d-1}{\delta}\right\rfloor,
$$
the constructions based on union-intersection-bounded families, packings, and Steiner systems achieve
$$
n=\Theta(q^t),
$$
so that code length can exceed linear order in $q$ [1812.11942]. This use is neither dynamical nor algorithmic; it is a structural asymptotic classification.

Taken together, these literatures show that “super-linear” is a relational descriptor rather than a single phenomenon. It may denote exponents greater than one, response slopes greater than one, convergence faster than linear, code lengths above linear order, or lower bounds above linear order. The unifying feature is asymptotic comparison with a linear baseline; the mechanisms—weak ties in cities, geometric densification, angiogenic feedback, nonlinear emission, coefficient growth, spectral routing, or circuit lower-bound constructions—are domain-specific [0809.4994][1212.4914][2307.01490][2509.15105].

Source: https://www.emergentmind.com/topics/super-linear