Super-Linear: Scaling, Convergence & Applications
- Super-linear is a descriptor for phenomena where outputs grow faster than linearly with respect to a variable, seen in power-law scaling and accelerated convergence rates.
- It appears across disciplines—from urban productivity and optical imaging to numerical algorithms—with key exponents greater than one driving distinct system behaviors.
- Named architectures like the Super-Linear forecasting model leverage mixtures-of-experts and spectral routing for robust, efficient time-series prediction with minimal parameters.
Super-linear denotes behavior that exceeds linear order with respect to a reference variable, but the technical meaning is domain-dependent. In the cited literature it appears as faster-than-linear scaling laws such as with , nonlinear input–output response laws such as with , coefficient-growth conditions in ODEs, SDEs, BSDEs, and BSVIEs, convergence regimes satisfying , and structural bounds such as code lengths or circuit sizes that exceed linear order. The term also names a specific forecasting architecture, "Super-Linear," built as a lightweight pretrained mixture of linear experts for time series forecasting (0809.4994, Wang et al., 2024, Kulkarni et al., 2020, Nochumsohn et al., 18 Sep 2025).
1. Core meanings and formal criteria
Across the literature, “super-linear” is not a single definition but a family of related order relations. In scaling-law settings, it means an exponent strictly larger than $1$, as in with (Zhang, 2012). In optical response, it denotes a power-law slope in the excitation–emission relation (Wang et al., 2024). In numerical analysis and optimization, it means convergence faster than linear, formally 0 as 1 (Kulkarni et al., 2020). In nonlinear dynamics, a function 2 may be called superlinear when 3 is ultimately increasing and 4 (Appleby et al., 2017).
| Context | Canonical form | Meaning |
|---|---|---|
| Scaling laws | 5, 6 | Output grows faster than proportionally |
| Optical response | 7, 8 | Emission localizes more sharply than linear response |
| Convergence theory | 9 | Faster-than-linear iteration |
| Nonlinear dynamics | 0 | State-dependent forcing exceeds linear growth |
A recurring misconception is to equate super-linear with quadratic. The cited numerical papers explicitly distinguish the two: quadratic convergence is a special case, whereas super-linear convergence only requires asymptotically faster-than-linear decay of error (Kulkarni et al., 2020, Wang et al., 2024). A second misconception is to equate super-linearity with monotone growth in time. The ODE literature shows that superlinear systems may exhibit finite-time blow-up, while other superlinear dissipative systems decay algebraically rather than exponentially (Appleby et al., 2017, Hoang, 2022).
2. Super-linear scaling in collective systems
In urban science, superlinear scaling refers to sociological quantities such as economic productivity, creative output, patents, inventors, and GDP increasing faster than city population. Reported empirical exponents are typically between 1 and 2, with a mean around 3 (0809.4994). The network model proposed for this phenomenon places individuals on the leaves of a hierarchical tree, defines social distance 4 as the height of the lowest common ancestor, assigns connection probability proportional to 5, counts 6 individuals at distance 7, and assumes a productivity benefit per tie proportional to 8. This yields
9
and for large 0 with 1,
2
The mechanism is the proliferation of socially distant links, interpreted as productive “weak ties,” so that larger cities create more opportunities for novelty and creative collaboration (0809.4994).
A geometrically generative account appears in growing random geometric graph models, where new nodes survive only if they are placed within distance 3 of existing nodes and then connect to all existing nodes within radius 4. In that setting the total number of edges obeys a super-linear power law 5, and the geometric dimension 6 is the primary parameter controlling 7 (Zhang, 2012). Simulations reported 8 for 9, 0 for 1, and 2 for 3, while the same framework also reproduced fractal growth, asymptotically size-invariant clustering coefficient, and sub-linear area–population and diversity–population relations (Zhang, 2012).
In tumor ecology, super-linear growth is written as
4
with 5 defining the super-linear regime (Azimzade et al., 2021). The cited model attributes this regime not to competition or fitter subclones alone, but to tumor–microenvironment interaction through angiogenesis. In the oxygen dynamics,
6
the term 7 increases oxygen supply in proportion to tumor mass, producing positive feedback compatible with an Allee effect (Azimzade et al., 2021). The same paper states that recent empirical work found average tumor growth exponents 8 across various human cancers (Azimzade et al., 2021).
Robot learning supplies a further operational meaning. The CASHER pipeline reports super-linear scaling with human effort by crowdsourcing digital twins of real scenes, collecting behavioral data in simulation, and gradually replacing human demonstrations with model-generated demonstrations as a generalist policy improves (Torne et al., 2024). The authors state that required human demonstrations per environment decrease as the number of environments grows, while zero-shot and few-shot scaling laws are demonstrated on three real-world tasks (Torne et al., 2024). A plausible commonality across cities, networks, tumors, and CASHER is that super-linear scaling emerges when larger system size increases the density or efficacy of productive interactions faster than it increases the relevant cost base.
3. Super-linear optical response and super-resolution
In fluorescence microscopy, super-linearity is used in a response-law sense. Super-linear image scanning microscopy extends conventional ISM by exploiting nonlinear upconversion emission from lanthanide-doped UCNPs, with
9
When $1$0, emission is more tightly localized at the excitation center, narrowing the emission PSF and pushing resolution beyond the twofold ISM limit (Wang et al., 2024). In the reported implementation, NaYF$1$1 doped with $1$2 Yb$1$3 and $1$4 Tm$1$5 was excited at $1$6 nm by a single low-power continuous-wave near-infrared laser. The five-photon $1$7 nm emission reached $1$8 at about $1$9 mW, producing measured FWHM 0 nm, or 1 for 2 nm; Fourier ring correlation gave about 3 nm (Wang et al., 2024). The same work also reported a multifocal structured-excitation variant with roughly 4 foci, a field of view of 5, and frame rates up to 6 Hz (Wang et al., 2024).
A distinct optical mechanism appears in bistable scattering from nano-silicon Mie resonators. There, photo-thermo-optical feedback produces optical bistability in a silicon resonator with volume size 7 and 8-factor 9, and the bistable transition yields a large effective super-linear scattering–excitation law with measured slope 0 (Tseng et al., 2023). For a Gaussian excitation beam, the cited analysis gives
1
so the emission PSF narrows by a factor 2 (Tseng et al., 2023). Experimentally, diffraction-limited laser scanning microscopy with FWHM 3 nm was sharpened to about 4 nm, a resolution enhancement of more than 5 times (Tseng et al., 2023).
These two optical lines use different nonlinearities. UCNP-based SL-ISM relies on multiphoton upconversion with experimentally observed 6, whereas nano-silicon bistable scattering relies on thermally driven resonance shifts and hysteresis, reaching 7 near the transition (Wang et al., 2024, Tseng et al., 2023). The shared consequence is PSF compression through a super-linear emission law rather than through purely linear optical transfer.
4. Differential, stochastic, and backward equations
For forced ODEs,
8
superlinearity is imposed by requiring 9 to be continuous, positive, increasing on 0, with 1 ultimately increasing and 2 (Appleby et al., 2017). Defining
3
finite-time blow-up occurs if 4; when 5, the forcing–nonlinearity competition can be classified sharply. If 6, then 7. If the limit superior equals 8, then 9. Under an additional negligibility condition, 0 (Appleby et al., 2017). Thus “superlinear” in the state variable does not determine the asymptotic regime by itself; the forcing scale matters.
A different use appears in genuinely nonlinear dissipative systems
1
where 2 is positively homogeneous of degree 3 and positive away from the origin (Hoang, 2022). Here the principal effect of superlinearity is not explosion but non-exponential decay. For sufficiently small initial data, nontrivial decaying solutions satisfy
4
with 5 an eigenvector of 6 satisfying 7 for the corresponding eigenvalue 8 (Hoang, 2022). The contrast with linear theory is explicit: decay is algebraic, not exponential (Hoang, 2022).
The SDE and SFDE literature uses “super-linear” chiefly as a coefficient-growth condition. For multidimensional SDEs with non-Lipschitz coefficients, the local logarithmic hypothesis
9
on each ball 00 yields pathwise uniqueness, non-contact, a stochastic flow of continuous maps, and a Freidlin–Wentzell-type large deviations principle (Bahlali et al., 2015). For super-linear SFDEs, an explicit truncated Euler–Maruyama scheme with linear interpolation achieves boundedness, strong convergence in 01, convergence rate 02, and preservation of exponential stability, without requiring global Lipschitz continuity of the diffusion coefficient (Li et al., 2022). For SDEs with superlinearly growing drift and diffusion coefficients, explicit Milstein schemes based on tamed coefficients converge in 03 with the optimal strong rate 04 under mild assumptions (Kumar et al., 2016).
Backward equations sharpen the threshold structure. Multi-dimensional BSVIEs with generators that are diagonally strictly quadratic in 05 and sub-quadratically coupled off-diagonally admit unique adapted solutions for bounded free term; when the free term is unbounded but has exponential moments of arbitrary order, unique solvability persists only in the diagonal at-most-quadratic case (Fan et al., 2022). The same paper presents negative results for super-quadratic growth in 06: in general, even bounded free term does not guarantee bounded solutions (Fan et al., 2022). Scalar BSDEs with generator growth
07
exhibit four different integrability thresholds for the terminal condition, according to 08, 09, 10, and 11; comparison and uniqueness follow when one generator is convex or concave in 12, or when it satisfies a one-sided Osgood condition in 13 and uniform continuity in 14 (Fan et al., 2021). In this branch of the literature, “super-linear” therefore marks a solvability frontier rather than a uniform dynamical effect.
5. Convergence, speedup, and lower bounds
In iterative computation, super-linear often refers to convergence rate. The refined 15-adic QR algorithm defines super-linear convergence by the standard criterion 16 and proves a stronger block-deflation estimate under eigenvalue separation: if a size-sorted Hessenberg matrix has suitably separated eigenvalues in the small block, then after the corresponding QR cycle
17
The resulting convergence is essentially quadratic in many cases, while the algorithm falls back to linear behavior when the favorable separation structure fails (Kulkarni et al., 2020). When the characteristic polynomial modulo 18 is square-free and splits completely, the paper states that all eigenvalues can be obtained up to error 19 in at most
20
arithmetic operations at 21 22-adic digits of precision (Kulkarni et al., 2020).
Optimization papers use the same convergence terminology but different mechanisms. “Superlinear Optimization Algorithms” proposes trajectory-inspired updates for minimizing nonlinear objectives, with several variants remaining applicable when the Hessian is singular (Wang et al., 2024). One family avoids calculating the inverse of the Hessian matrix or an identical-dimension matrix; another requires only the diagonal elements of the Hessian; all are reported to be superlinear convergent when appropriate parameters are selected (Wang et al., 2024). In mixed linear regression, alternating minimization is shown to contract estimation error super-linearly under proper initialization. The main recurrence is of the form
23
which yields 24 iterations to reach 25-accuracy, with a quadratic regime in a narrower neighborhood of the optimum (Ghosh et al., 2020).
Program transformation and computational complexity supply two further meanings. Repeated recursion unfolding repeatedly unfolds a recursive rule with itself, so that each unfolding doubles the number of recursive steps covered; with optimal rule application, the runtime recurrence changes from
26
to
27
and, in the best case, the method lowers time complexity class within a chosen bound on recursion depth (Fruehwirth, 2020). The paper explicitly lists examples such as quadratic becoming linear and linear becoming constant time (Fruehwirth, 2020). By contrast, threshold-circuit complexity uses “super-linear” comparatively: the paper on depth-two and depth-three threshold circuits proves the first super-linear gate lower bounds and the first super-quadratic wire lower bounds for explicit functions. For Andreev’s function, any depth-two linear threshold circuit agreeing on a 28-fraction of inputs requires at least 29 gates or 30 wires, while PARITY has tight average-case complexity 31 gates and 32 wires in this setting (Kane et al., 2015). This usage is about lower-bound magnitude, not iterative improvement.
6. Named architectures and broader distinctions
“Super-Linear” is also the title of a time-series forecasting model: a lightweight pretrained mixture-of-experts architecture built from frequency-specialized linear experts and a spectral router (Nochumsohn et al., 18 Sep 2025). Given input sequence 33 and forecast 34, the model writes the prediction as
35
where 36 are linear experts and the gating network is driven by normalized spectral features,
37
followed by sparse Top-38 softmax selection (Nochumsohn et al., 18 Sep 2025). Training proceeds in two stages: independent pretraining of experts on aggressively resampled data to match target frequency regimes, then freezing those experts while jointly training the router and complementary experts (Nochumsohn et al., 18 Sep 2025).
The reported empirical profile is explicitly tied to efficiency. Super-Linear uses about 39M parameters, trains and runs on a single GPU, and is described as much smaller than Timer-XL, at 40 of its size (Nochumsohn et al., 18 Sep 2025). On the LTSF zero-shot benchmark it reports average MSE reductions of 41 relative to Chronos, 42 relative to TimesFM, 43 relative to Moirai, and 44 relative to Time-MoE Large, while on GIFT-Eval it reports a 45 MASE reduction over the lightweight TTM model (Nochumsohn et al., 18 Sep 2025). The “super” in the model name is therefore nominal, but the paper explicitly ties that architecture to robustness across sampling rates, sparse interpretability through spectral routing, and strong zero-shot/full-shot performance (Nochumsohn et al., 18 Sep 2025).
Coding theory uses the term in yet another strictly order-theoretic sense. For optimal locally repairable codes with all-symbol 46-locality, the paper derives alphabet-dependent upper bounds on length and constructs order-optimal codes whose length is super-linear in the alphabet size (Cai et al., 2018). With
47
the constructions based on union-intersection-bounded families, packings, and Steiner systems achieve
48
so that code length can exceed linear order in 49 (Cai et al., 2018). This use is neither dynamical nor algorithmic; it is a structural asymptotic classification.
Taken together, these literatures show that “super-linear” is a relational descriptor rather than a single phenomenon. It may denote exponents greater than one, response slopes greater than one, convergence faster than linear, code lengths above linear order, or lower bounds above linear order. The unifying feature is asymptotic comparison with a linear baseline; the mechanisms—weak ties in cities, geometric densification, angiogenic feedback, nonlinear emission, coefficient growth, spectral routing, or circuit lower-bound constructions—are domain-specific (0809.4994, Zhang, 2012, Tseng et al., 2023, Nochumsohn et al., 18 Sep 2025).