Papers
Topics
Authors
Recent
Search
2000 character limit reached

Super-Linear: Scaling, Convergence & Applications

Updated 12 July 2026
  • Super-linear is a descriptor for phenomena where outputs grow faster than linearly with respect to a variable, seen in power-law scaling and accelerated convergence rates.
  • It appears across disciplines—from urban productivity and optical imaging to numerical algorithms—with key exponents greater than one driving distinct system behaviors.
  • Named architectures like the Super-Linear forecasting model leverage mixtures-of-experts and spectral routing for robust, efficient time-series prediction with minimal parameters.

Super-linear denotes behavior that exceeds linear order with respect to a reference variable, but the technical meaning is domain-dependent. In the cited literature it appears as faster-than-linear scaling laws such as XPγX \propto P^\gamma with γ>1\gamma>1, nonlinear input–output response laws such as IemIexcsI_{em} \propto I_{exc}^s with s>1s>1, coefficient-growth conditions in ODEs, SDEs, BSDEs, and BSVIEs, convergence regimes satisfying xk+1x/xkx0\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 0, and structural bounds such as code lengths or circuit sizes that exceed linear order. The term also names a specific forecasting architecture, "Super-Linear," built as a lightweight pretrained mixture of linear experts for time series forecasting (0809.4994, Wang et al., 2024, Kulkarni et al., 2020, Nochumsohn et al., 18 Sep 2025).

1. Core meanings and formal criteria

Across the literature, “super-linear” is not a single definition but a family of related order relations. In scaling-law settings, it means an exponent strictly larger than $1$, as in XPγX \propto P^\gamma with γ>1\gamma>1 (Zhang, 2012). In optical response, it denotes a power-law slope s>1s>1 in the excitation–emission relation Iem(Iexc)sI_{em} \propto (I_{exc})^s (Wang et al., 2024). In numerical analysis and optimization, it means convergence faster than linear, formally γ>1\gamma>10 as γ>1\gamma>11 (Kulkarni et al., 2020). In nonlinear dynamics, a function γ>1\gamma>12 may be called superlinear when γ>1\gamma>13 is ultimately increasing and γ>1\gamma>14 (Appleby et al., 2017).

Context Canonical form Meaning
Scaling laws γ>1\gamma>15, γ>1\gamma>16 Output grows faster than proportionally
Optical response γ>1\gamma>17, γ>1\gamma>18 Emission localizes more sharply than linear response
Convergence theory γ>1\gamma>19 Faster-than-linear iteration
Nonlinear dynamics IemIexcsI_{em} \propto I_{exc}^s0 State-dependent forcing exceeds linear growth

A recurring misconception is to equate super-linear with quadratic. The cited numerical papers explicitly distinguish the two: quadratic convergence is a special case, whereas super-linear convergence only requires asymptotically faster-than-linear decay of error (Kulkarni et al., 2020, Wang et al., 2024). A second misconception is to equate super-linearity with monotone growth in time. The ODE literature shows that superlinear systems may exhibit finite-time blow-up, while other superlinear dissipative systems decay algebraically rather than exponentially (Appleby et al., 2017, Hoang, 2022).

2. Super-linear scaling in collective systems

In urban science, superlinear scaling refers to sociological quantities such as economic productivity, creative output, patents, inventors, and GDP increasing faster than city population. Reported empirical exponents are typically between IemIexcsI_{em} \propto I_{exc}^s1 and IemIexcsI_{em} \propto I_{exc}^s2, with a mean around IemIexcsI_{em} \propto I_{exc}^s3 (0809.4994). The network model proposed for this phenomenon places individuals on the leaves of a hierarchical tree, defines social distance IemIexcsI_{em} \propto I_{exc}^s4 as the height of the lowest common ancestor, assigns connection probability proportional to IemIexcsI_{em} \propto I_{exc}^s5, counts IemIexcsI_{em} \propto I_{exc}^s6 individuals at distance IemIexcsI_{em} \propto I_{exc}^s7, and assumes a productivity benefit per tie proportional to IemIexcsI_{em} \propto I_{exc}^s8. This yields

IemIexcsI_{em} \propto I_{exc}^s9

and for large s>1s>10 with s>1s>11,

s>1s>12

The mechanism is the proliferation of socially distant links, interpreted as productive “weak ties,” so that larger cities create more opportunities for novelty and creative collaboration (0809.4994).

A geometrically generative account appears in growing random geometric graph models, where new nodes survive only if they are placed within distance s>1s>13 of existing nodes and then connect to all existing nodes within radius s>1s>14. In that setting the total number of edges obeys a super-linear power law s>1s>15, and the geometric dimension s>1s>16 is the primary parameter controlling s>1s>17 (Zhang, 2012). Simulations reported s>1s>18 for s>1s>19, xk+1x/xkx0\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 00 for xk+1x/xkx0\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 01, and xk+1x/xkx0\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 02 for xk+1x/xkx0\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 03, while the same framework also reproduced fractal growth, asymptotically size-invariant clustering coefficient, and sub-linear area–population and diversity–population relations (Zhang, 2012).

In tumor ecology, super-linear growth is written as

xk+1x/xkx0\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 04

with xk+1x/xkx0\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 05 defining the super-linear regime (Azimzade et al., 2021). The cited model attributes this regime not to competition or fitter subclones alone, but to tumor–microenvironment interaction through angiogenesis. In the oxygen dynamics,

xk+1x/xkx0\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 06

the term xk+1x/xkx0\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 07 increases oxygen supply in proportion to tumor mass, producing positive feedback compatible with an Allee effect (Azimzade et al., 2021). The same paper states that recent empirical work found average tumor growth exponents xk+1x/xkx0\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 08 across various human cancers (Azimzade et al., 2021).

Robot learning supplies a further operational meaning. The CASHER pipeline reports super-linear scaling with human effort by crowdsourcing digital twins of real scenes, collecting behavioral data in simulation, and gradually replacing human demonstrations with model-generated demonstrations as a generalist policy improves (Torne et al., 2024). The authors state that required human demonstrations per environment decrease as the number of environments grows, while zero-shot and few-shot scaling laws are demonstrated on three real-world tasks (Torne et al., 2024). A plausible commonality across cities, networks, tumors, and CASHER is that super-linear scaling emerges when larger system size increases the density or efficacy of productive interactions faster than it increases the relevant cost base.

3. Super-linear optical response and super-resolution

In fluorescence microscopy, super-linearity is used in a response-law sense. Super-linear image scanning microscopy extends conventional ISM by exploiting nonlinear upconversion emission from lanthanide-doped UCNPs, with

xk+1x/xkx0\|x_{k+1}-x^*\|/\|x_k-x^*\|\to 09

When $1$0, emission is more tightly localized at the excitation center, narrowing the emission PSF and pushing resolution beyond the twofold ISM limit (Wang et al., 2024). In the reported implementation, NaYF$1$1 doped with $1$2 Yb$1$3 and $1$4 Tm$1$5 was excited at $1$6 nm by a single low-power continuous-wave near-infrared laser. The five-photon $1$7 nm emission reached $1$8 at about $1$9 mW, producing measured FWHM XPγX \propto P^\gamma0 nm, or XPγX \propto P^\gamma1 for XPγX \propto P^\gamma2 nm; Fourier ring correlation gave about XPγX \propto P^\gamma3 nm (Wang et al., 2024). The same work also reported a multifocal structured-excitation variant with roughly XPγX \propto P^\gamma4 foci, a field of view of XPγX \propto P^\gamma5, and frame rates up to XPγX \propto P^\gamma6 Hz (Wang et al., 2024).

A distinct optical mechanism appears in bistable scattering from nano-silicon Mie resonators. There, photo-thermo-optical feedback produces optical bistability in a silicon resonator with volume size XPγX \propto P^\gamma7 and XPγX \propto P^\gamma8-factor XPγX \propto P^\gamma9, and the bistable transition yields a large effective super-linear scattering–excitation law with measured slope γ>1\gamma>10 (Tseng et al., 2023). For a Gaussian excitation beam, the cited analysis gives

γ>1\gamma>11

so the emission PSF narrows by a factor γ>1\gamma>12 (Tseng et al., 2023). Experimentally, diffraction-limited laser scanning microscopy with FWHM γ>1\gamma>13 nm was sharpened to about γ>1\gamma>14 nm, a resolution enhancement of more than γ>1\gamma>15 times (Tseng et al., 2023).

These two optical lines use different nonlinearities. UCNP-based SL-ISM relies on multiphoton upconversion with experimentally observed γ>1\gamma>16, whereas nano-silicon bistable scattering relies on thermally driven resonance shifts and hysteresis, reaching γ>1\gamma>17 near the transition (Wang et al., 2024, Tseng et al., 2023). The shared consequence is PSF compression through a super-linear emission law rather than through purely linear optical transfer.

4. Differential, stochastic, and backward equations

For forced ODEs,

γ>1\gamma>18

superlinearity is imposed by requiring γ>1\gamma>19 to be continuous, positive, increasing on s>1s>10, with s>1s>11 ultimately increasing and s>1s>12 (Appleby et al., 2017). Defining

s>1s>13

finite-time blow-up occurs if s>1s>14; when s>1s>15, the forcing–nonlinearity competition can be classified sharply. If s>1s>16, then s>1s>17. If the limit superior equals s>1s>18, then s>1s>19. Under an additional negligibility condition, Iem(Iexc)sI_{em} \propto (I_{exc})^s0 (Appleby et al., 2017). Thus “superlinear” in the state variable does not determine the asymptotic regime by itself; the forcing scale matters.

A different use appears in genuinely nonlinear dissipative systems

Iem(Iexc)sI_{em} \propto (I_{exc})^s1

where Iem(Iexc)sI_{em} \propto (I_{exc})^s2 is positively homogeneous of degree Iem(Iexc)sI_{em} \propto (I_{exc})^s3 and positive away from the origin (Hoang, 2022). Here the principal effect of superlinearity is not explosion but non-exponential decay. For sufficiently small initial data, nontrivial decaying solutions satisfy

Iem(Iexc)sI_{em} \propto (I_{exc})^s4

with Iem(Iexc)sI_{em} \propto (I_{exc})^s5 an eigenvector of Iem(Iexc)sI_{em} \propto (I_{exc})^s6 satisfying Iem(Iexc)sI_{em} \propto (I_{exc})^s7 for the corresponding eigenvalue Iem(Iexc)sI_{em} \propto (I_{exc})^s8 (Hoang, 2022). The contrast with linear theory is explicit: decay is algebraic, not exponential (Hoang, 2022).

The SDE and SFDE literature uses “super-linear” chiefly as a coefficient-growth condition. For multidimensional SDEs with non-Lipschitz coefficients, the local logarithmic hypothesis

Iem(Iexc)sI_{em} \propto (I_{exc})^s9

on each ball γ>1\gamma>100 yields pathwise uniqueness, non-contact, a stochastic flow of continuous maps, and a Freidlin–Wentzell-type large deviations principle (Bahlali et al., 2015). For super-linear SFDEs, an explicit truncated Euler–Maruyama scheme with linear interpolation achieves boundedness, strong convergence in γ>1\gamma>101, convergence rate γ>1\gamma>102, and preservation of exponential stability, without requiring global Lipschitz continuity of the diffusion coefficient (Li et al., 2022). For SDEs with superlinearly growing drift and diffusion coefficients, explicit Milstein schemes based on tamed coefficients converge in γ>1\gamma>103 with the optimal strong rate γ>1\gamma>104 under mild assumptions (Kumar et al., 2016).

Backward equations sharpen the threshold structure. Multi-dimensional BSVIEs with generators that are diagonally strictly quadratic in γ>1\gamma>105 and sub-quadratically coupled off-diagonally admit unique adapted solutions for bounded free term; when the free term is unbounded but has exponential moments of arbitrary order, unique solvability persists only in the diagonal at-most-quadratic case (Fan et al., 2022). The same paper presents negative results for super-quadratic growth in γ>1\gamma>106: in general, even bounded free term does not guarantee bounded solutions (Fan et al., 2022). Scalar BSDEs with generator growth

γ>1\gamma>107

exhibit four different integrability thresholds for the terminal condition, according to γ>1\gamma>108, γ>1\gamma>109, γ>1\gamma>110, and γ>1\gamma>111; comparison and uniqueness follow when one generator is convex or concave in γ>1\gamma>112, or when it satisfies a one-sided Osgood condition in γ>1\gamma>113 and uniform continuity in γ>1\gamma>114 (Fan et al., 2021). In this branch of the literature, “super-linear” therefore marks a solvability frontier rather than a uniform dynamical effect.

5. Convergence, speedup, and lower bounds

In iterative computation, super-linear often refers to convergence rate. The refined γ>1\gamma>115-adic QR algorithm defines super-linear convergence by the standard criterion γ>1\gamma>116 and proves a stronger block-deflation estimate under eigenvalue separation: if a size-sorted Hessenberg matrix has suitably separated eigenvalues in the small block, then after the corresponding QR cycle

γ>1\gamma>117

The resulting convergence is essentially quadratic in many cases, while the algorithm falls back to linear behavior when the favorable separation structure fails (Kulkarni et al., 2020). When the characteristic polynomial modulo γ>1\gamma>118 is square-free and splits completely, the paper states that all eigenvalues can be obtained up to error γ>1\gamma>119 in at most

γ>1\gamma>120

arithmetic operations at γ>1\gamma>121 γ>1\gamma>122-adic digits of precision (Kulkarni et al., 2020).

Optimization papers use the same convergence terminology but different mechanisms. “Superlinear Optimization Algorithms” proposes trajectory-inspired updates for minimizing nonlinear objectives, with several variants remaining applicable when the Hessian is singular (Wang et al., 2024). One family avoids calculating the inverse of the Hessian matrix or an identical-dimension matrix; another requires only the diagonal elements of the Hessian; all are reported to be superlinear convergent when appropriate parameters are selected (Wang et al., 2024). In mixed linear regression, alternating minimization is shown to contract estimation error super-linearly under proper initialization. The main recurrence is of the form

γ>1\gamma>123

which yields γ>1\gamma>124 iterations to reach γ>1\gamma>125-accuracy, with a quadratic regime in a narrower neighborhood of the optimum (Ghosh et al., 2020).

Program transformation and computational complexity supply two further meanings. Repeated recursion unfolding repeatedly unfolds a recursive rule with itself, so that each unfolding doubles the number of recursive steps covered; with optimal rule application, the runtime recurrence changes from

γ>1\gamma>126

to

γ>1\gamma>127

and, in the best case, the method lowers time complexity class within a chosen bound on recursion depth (Fruehwirth, 2020). The paper explicitly lists examples such as quadratic becoming linear and linear becoming constant time (Fruehwirth, 2020). By contrast, threshold-circuit complexity uses “super-linear” comparatively: the paper on depth-two and depth-three threshold circuits proves the first super-linear gate lower bounds and the first super-quadratic wire lower bounds for explicit functions. For Andreev’s function, any depth-two linear threshold circuit agreeing on a γ>1\gamma>128-fraction of inputs requires at least γ>1\gamma>129 gates or γ>1\gamma>130 wires, while PARITY has tight average-case complexity γ>1\gamma>131 gates and γ>1\gamma>132 wires in this setting (Kane et al., 2015). This usage is about lower-bound magnitude, not iterative improvement.

6. Named architectures and broader distinctions

“Super-Linear” is also the title of a time-series forecasting model: a lightweight pretrained mixture-of-experts architecture built from frequency-specialized linear experts and a spectral router (Nochumsohn et al., 18 Sep 2025). Given input sequence γ>1\gamma>133 and forecast γ>1\gamma>134, the model writes the prediction as

γ>1\gamma>135

where γ>1\gamma>136 are linear experts and the gating network is driven by normalized spectral features,

γ>1\gamma>137

followed by sparse Top-γ>1\gamma>138 softmax selection (Nochumsohn et al., 18 Sep 2025). Training proceeds in two stages: independent pretraining of experts on aggressively resampled data to match target frequency regimes, then freezing those experts while jointly training the router and complementary experts (Nochumsohn et al., 18 Sep 2025).

The reported empirical profile is explicitly tied to efficiency. Super-Linear uses about γ>1\gamma>139M parameters, trains and runs on a single GPU, and is described as much smaller than Timer-XL, at γ>1\gamma>140 of its size (Nochumsohn et al., 18 Sep 2025). On the LTSF zero-shot benchmark it reports average MSE reductions of γ>1\gamma>141 relative to Chronos, γ>1\gamma>142 relative to TimesFM, γ>1\gamma>143 relative to Moirai, and γ>1\gamma>144 relative to Time-MoE Large, while on GIFT-Eval it reports a γ>1\gamma>145 MASE reduction over the lightweight TTM model (Nochumsohn et al., 18 Sep 2025). The “super” in the model name is therefore nominal, but the paper explicitly ties that architecture to robustness across sampling rates, sparse interpretability through spectral routing, and strong zero-shot/full-shot performance (Nochumsohn et al., 18 Sep 2025).

Coding theory uses the term in yet another strictly order-theoretic sense. For optimal locally repairable codes with all-symbol γ>1\gamma>146-locality, the paper derives alphabet-dependent upper bounds on length and constructs order-optimal codes whose length is super-linear in the alphabet size (Cai et al., 2018). With

γ>1\gamma>147

the constructions based on union-intersection-bounded families, packings, and Steiner systems achieve

γ>1\gamma>148

so that code length can exceed linear order in γ>1\gamma>149 (Cai et al., 2018). This use is neither dynamical nor algorithmic; it is a structural asymptotic classification.

Taken together, these literatures show that “super-linear” is a relational descriptor rather than a single phenomenon. It may denote exponents greater than one, response slopes greater than one, convergence faster than linear, code lengths above linear order, or lower bounds above linear order. The unifying feature is asymptotic comparison with a linear baseline; the mechanisms—weak ties in cities, geometric densification, angiogenic feedback, nonlinear emission, coefficient growth, spectral routing, or circuit lower-bound constructions—are domain-specific (0809.4994, Zhang, 2012, Tseng et al., 2023, Nochumsohn et al., 18 Sep 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (20)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Super-Linear.