Higher-Order Linearization
- Higher-Order Linearization is the extension of first-order tangent approximations by incorporating higher derivatives and lifted features to preserve nonlinear structure.
- It forms a hierarchical framework that transforms complex nonlinear dynamics into a series of linear or multilinear models across fields like PDEs, dynamical systems, and machine learning.
- Its applications span improved numerical schemes, enhanced kernel descriptors in action recognition, and rigorous integrability tests in complex systems such as gravity.
Higher-order linearization denotes a family of constructions that extend ordinary first-order tangent or Jacobian approximations by retaining higher derivatives, higher variational equations, higher-order Taylor terms, or explicit lifted features. Across nonlinear PDEs, dynamical systems, geometric structures, optimization, and machine learning, the resulting object may be a recursive hierarchy of linear equations, a jet-space linear system, a lifted affine model on monomials, or an explicit tensor descriptor whose inner product approximates a nonlinear kernel (Armstrong et al., 2019, Simon, 2013, Akiba et al., 7 May 2026, Cherian et al., 2017). The unifying aim is to preserve more of the nonlinear structure than ordinary linearization while keeping some form of linear or multilinear tractability.
1. Core patterns of the concept
At first order, linearization replaces a nonlinear object by its tangent model: a Jacobian for a map, a linearized operator for a PDE, or a tangent dynamics along a trajectory. Higher-order linearization goes further by incorporating second and higher derivatives, or by enlarging the state or feature space so that nonlinear interactions become linear in lifted coordinates. In the PDE literature this typically appears as a hierarchy of linear equations whose source terms depend polynomially on lower-order jets; in lifted-system methods it appears as an affine or linear evolution on monomials or tensorized features; in representation learning it may appear as an explicit higher-order descriptor obtained after kernel linearization (Armstrong et al., 2019, Uhlmann et al., 11 Jun 2026, Akiba et al., 7 May 2026, Cherian et al., 2017).
Taken together, these works suggest three recurrent forms. The first is hierarchical differentiation, in which repeated differentiation of a nonlinear equation produces a triangular family of linearized equations, each forced by lower-order terms. The second is lift-and-linearize, in which nonlinearities are represented explicitly in a larger space, such as symmetric jets, monomials, or tensorized feature maps. The third is reduction of higher-derivative structure, where a theory that is formally higher order is transformed, often by gauge fixing or exact reformulation, into second-order master equations for physical modes (Simon, 2013, Akiba et al., 7 May 2026, Zhong et al., 2016).
A further recurring theme is that higher-order linearization is often paired with a compatibility question. In some settings it delivers better approximations or richer descriptors; in others it tests whether first-order solutions are actually integrable into genuine nonlinear ones. This distinction is especially sharp in gravity, where second-order consistency conditions can exclude spurious first-order perturbations (Altas, 2018).
2. Jet hierarchies, variational equations, and geometric linearization
For autonomous holomorphic dynamical systems , higher derivatives of the flow with respect to initial data satisfy a nonlinear recursive hierarchy. The paper "Linearised Higher Variational Equations" (Simon, 2013) shows that these nonlinear higher variational equations can be repackaged into a linear block-triangular system on jets. Using the symmetric product , the full jet object satisfies a linear system
and its finite truncations provide explicit system matrices, principal fundamental matrices, and monodromy matrices. In this setting, higher-order linearization means linearizing not merely tangent vectors but the full -jet of the flow along a particular solution. This formulation is designed for constructive use in Ziglin–Morales–Ramis theory, because differential Galois and monodromy methods require genuinely linear systems (Simon, 2013).
A related but more geometric formulation appears in "Equivalence of variational problems of higher order" (Doubrov et al., 2010). There, higher-order Lagrangians, variational ODEs of order $2n$, and rank $2$ distributions associated with underdetermined ODEs are shown, for , to define essentially the same equivalence problem. The crucial construction is linearization along distinguished curves: solutions of the Euler–Lagrange equation, or abnormal extremals of the associated distribution. This produces self-dual projective curves in 0, and their generalized Wilczynski invariants become the fundamental invariants of the original nonlinear problem. Maximal symmetry occurs exactly when all such linearized curves are rational normal curves (Doubrov et al., 2010).
An analogous dual-bundle viewpoint underlies "Linearization of the higher analogue of Courant algebroids" (2002.03656). For a vector bundle 1, sections of 2 are identified with linear 3-vector fields on 4, and sections of 5 with linear 6-forms on 7. This yields two higher linearizations of 8: the pseudo-linearization 9, and the Weinstein-linearization 0. The second is particularly significant because integrable graph subbundles of the resulting omni 1-Lie algebroid characterize 2-Lie algebroids, local 3-Lie algebras, and Nambu–Jacobi structures (2002.03656).
3. Nonlinear PDEs, homogenization, and inverse problems
In nonlinear stochastic homogenization, higher-order linearization organizes the passage from microscopic nonlinearity to effective macroscopic regularity. The paper "Higher-order linearization and regularity in nonlinear homogenization" (Armstrong et al., 2019) studies uniformly convex divergence-form equations
4
and constructs a hierarchy of linearized correctors 5 around a slope 6. The source term 7 at order 8 depends polynomially on lower-order increments, and the 9-th linearized corrector solves a linear equation with coefficient field 0. The main structural result is that higher-order linearization and homogenization commute: derivatives of the homogenized Lagrangian 1 can be expressed through homogenized fluxes of higher linearized correctors, and this yields large-scale 2-type regularity, a Liouville theorem for the linearized hierarchy, and a heterogeneous Taylor expansion with optimally controlled remainder (Armstrong et al., 2019). Here higher-order linearization is not merely an approximation device; it is the mechanism tying together homogenization, effective regularity, and the derivative structure of 3.
In inverse problems for the nonlinear Monge–Ampère equation,
4
higher-order linearization serves a different purpose: it converts nonlinear boundary measurements into successive linear inverse problems. In "Inverse Problems for the Monge--Ampère Equation: Linearization and Nonlinear Recovery" (Uhlmann et al., 11 Jun 2026), linearizing around a strictly convex background solution 5 produces the uniformly elliptic operator
6
where 7, 8, and 9. Multi-parameter differentiation of the nonlinear solution map yields higher variations 0, and Faà di Bruno partitions generate recursive linear PDEs whose source terms involve higher derivatives of both 1 and 2. The resulting higher-order integral identities recover higher derivatives 3 along the background jet, and under suitable uniqueness and density assumptions determine the full Taylor expansion of the nonlinearity there (Uhlmann et al., 11 Jun 2026). In this setting, higher-order linearization is an identifiability tool rather than a surrogate model.
4. Higher-derivative field theories and second-order obstructions
In higher-curvature gravity, higher-order linearization often means handling field equations that are formally fourth order in the metric. "Linearization of a warped 4 theory in the higher-order frame" (Zhong et al., 2016) studies five-dimensional warped backgrounds directly in the original 5 formulation, rather than moving to the Einstein frame. The crucial observation is that the problematic higher-derivative term in the quadratic action is 6. After scalar-vector-tensor decomposition, the curvature gauge
7
eliminates the explicit higher-derivative scalar-curvature fluctuation, and the tensor, vector, and scalar normal modes all reduce to second-order equations. The tensor and scalar sectors take Schrödinger-type form, and the quadratic actions of the normal modes are equivalent to those found in the Einstein frame (Zhong et al., 2016).
The companion paper "Linearization of a warped 8 theory in the higher-order frame II: the equation of motion approach" (Zhong et al., 2017) derives the same reduction directly from the perturbed field equations. In curvature gauge, the scalar sector again collapses to a second-order master equation
9
while the tensor equation acquires the characteristic higher-order-frame coefficient 0. The equation-of-motion approach agrees with the quadratic-action method in the tensor and scalar sectors, but additionally yields a vector constraint absent from the action-based derivation (Zhong et al., 2017).
A distinct use of higher-order linearization appears in "Linearization Instability in Gravity Theories" (Altas, 2018). There the issue is not better approximation, but whether a first-order perturbation is tangent to an exact family of solutions. If 1 satisfies the linearized equations, then integrability requires the second-order consistency equation
2
for some second-order correction 3. Contracting with a background Killing field and integrating over a hypersurface produces the Taub and ADT charges. In chiral gravity and critical gravity, the relevant ADT charges vanish while the quadratic Taub charge need not, so some linearized solutions are spurious and the background is linearization unstable (Altas, 2018). Here higher-order linearization functions as an integrability test on perturbation theory itself.
5. Numerical schemes and lifted solvers
Several numerical traditions use higher-order linearization to obtain more accurate or more stable surrogates than ordinary Jacobian methods. In "Equivalent linearization finds nonzero frequency corrections beyond first order" (Chattopadhyay et al., 2016), the nonlinear oscillator is rewritten as a linear oscillator with an effective frequency plus residual forcing. When the first-order resonant term vanishes, the method does not stop; instead it computes the first-order waveform correction, reinserts that correction into the nonlinear term, and extracts the next nonzero frequency correction. The paper carries this out for conservative anharmonic oscillators and the van der Pol oscillator, showing that equivalent linearization can recover second-order corrections even when the first-order shift is zero (Chattopadhyay et al., 2016).
For dissipative PDEs, "Arbitrarily High-order Linear Schemes for Gradient Flow Models" (Gong et al., 2019) combines energy quadratization with extrapolative linearization and algebraically stable Runge–Kutta discretization. The nonlinear free energy is rewritten using an auxiliary variable 4, so the reformulated system has quadratic energy. On each time slab 5, the nonlinear coefficients are frozen by extrapolating 6 from previously computed data, yielding a linear variable-coefficient system. Because the Runge–Kutta method is algebraically stable, the fully discrete scheme is linear, unconditionally energy stable, uniquely solvable, and may reach arbitrarily high order (Gong et al., 2019). This is higher-order linearization as a time-stepping paradigm.
Finite-order Carleman methods provide another lifted form. "Duplicate-Aware Shift-and-Lift Carleman Linearization:Structure, Complexity, and Comparative Evaluation" (Akiba et al., 7 May 2026) studies polynomial ODEs 7 by lifting monomials 8 up to degree 9, producing a truncated affine system
$2n$0
The paper’s contribution is not a new convergence theorem but a practical higher-order construction: a shift-and-lift architecture, symmetry-reduced monomial bases, packed exponent-key indexing, sparse triplet coalescing, and a moving-center expansion that rebuilds the lifted model around the current state. The comparative conclusion is deliberately balanced: higher-order shift-and-lift can outperform Jacobian linearization when the dominant nonlinearities align with a moderate-order lifted basis and the lifted dynamics remain bounded and well conditioned, but the gains are regime-dependent rather than universal (Akiba et al., 7 May 2026).
6. Data-driven and machine-learning formulations
In data-driven dynamics, higher-order linearization often means constructing a nonlinear change of variables that makes reduced observable dynamics linear to higher order. "Data-Driven Linearization of Dynamical Systems" (Haller et al., 2024) reinterprets DMD as a local leading-order model for dynamics restricted to an attracting slow spectral submanifold $2n$1. The proposed data-driven linearization (DDL) then seeks a near-identity transformation
$2n$2
such that the reduced dynamics in $2n$3-coordinates are linear. Under hyperbolicity, observability, and nonresonance assumptions, DDL gives a systematic higher-order linearization of the dominant slow dynamics, whereas DMD is exactly the leading-order truncation $2n$4 (Haller et al., 2024).
A closely related trajectory-correction viewpoint appears in optimization. "Higher-Order Corrections to Optimisers based on Newton's Method" (Brooks, 2023) introduces a path $2n$5 defined by
$2n$6
whose tangent at $2n$7 is the Newton, Gauss–Newton, or Levenberg–Marquardt direction. Differentiating this path produces higher derivatives $2n$8; the second derivative recovers geodesic acceleration, and third- and fourth-order corrections are obtained by repeated differentiation via Faà di Bruno. The resulting updates extend Jacobian-based first-order linearization into higher-order local approximations tailored to narrow curved valleys in nonlinear least squares (Brooks, 2023).
In feature learning, "Higher-order Pooling of CNN Features via Kernel Linearization for Action Recognition" (Cherian et al., 2017) uses the term in yet another precise sense. A kernel on framewise CNN score vectors and normalized time indices is linearized by explicit feature maps $2n$9 and $2$0, then raised to order $2$1. The identity
$2$2
turns the sequence kernel into the inner product of explicit higher-order pooled tensors, yielding the Higher-order Kernel (HOK) descriptor. For $2$3 this recovers mean pooling; for $2$4 it captures covariance-like statistics; for $2$5 it captures third-order interactions among score and temporal features. The descriptor is super-symmetric, power normalized by HOSVD, and used with a linear classifier for video-level action recognition (Cherian et al., 2017).
The same movement beyond first-order linearization appears in wide-network theory. "Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks" (Bai et al., 2019) studies two-layer networks in a local Taylor regime beyond NTK. Standard lazy training couples the network to the first-order term $2$6; the paper introduces randomization schemes that suppress lower-order contributions and couple training to the quadratic term
$2$7
and more generally to $2$8-th order Taylor terms. This yields a mathematically controlled regime that is still local, still Taylor-governed, but no longer NTK-dominated (Bai et al., 2019).
7. Limits, contrasts, and recurrent caveats
The literature consistently rejects the idea that higher order is automatically superior. In HOK descriptors, third-order pooling improves over average pooling but is not uniformly better than second-order pooling, and the paper attributes weaker third-order performance on MPII to highly unbalanced sequence lengths and class counts (Cherian et al., 2017). In shift-and-lift Carleman linearization, higher truncation order improves coarse-grid behavior in some regimes but can lose its advantage when closure error or conditioning dominates (Akiba et al., 7 May 2026). DDL is explicitly local, depends on the existence of an attracting slow SSM, and relies on nonresonance and data-dominance assumptions (Haller et al., 2024). Wide-network coupling to quadratic and higher Taylor terms likewise depends on randomization and smooth activation assumptions (Bai et al., 2019).
An important contrast is that some problems do not admit a useful local first-order linearization at all. "Exact solutions of Kondo problems in higher-order fermions" (Song et al., 2022) argues that the standard Affleck–Ludwig conformal-field-theoretic treatment of Kondo impurities relies on linearizing low-energy excitations near a Fermi surface, but this fails for higher-order fermion systems with Fermi points and higher-order dispersion. The paper shows that naive linearization generates artificial flat bands, and replaces low-energy linearization by an exact full-energy mapping and bosonization of the entire spectrum. This contrast is instructive: sometimes the correct response to higher-order structure is a more elaborate higher-order linearization, but sometimes it is the abandonment of local linearization in favor of an exact reformulation (Song et al., 2022).
This suggests that higher-order linearization is best understood not as a single method but as a technical family of strategies for retaining nonlinear information beyond the tangent level. Depending on the field, it may mean jet-space linear systems, hierarchies of linearized PDEs, lifted affine models, higher-order optimization-path corrections, or explicit tensorized descriptors. Its success depends on whether the chosen higher-order object remains faithful, stable, and computationally tractable for the nonlinear structure under study.