Papers
Topics
Authors
Recent
Search
2000 character limit reached

Cubic-Regularized Riemannian Newton Method

Updated 12 July 2026
  • The paper presents a second-order cubic-regularized Newton algorithm that ensures O(1/ε^(3/2)) complexity for reaching second-order ε-stationarity on Riemannian manifolds.
  • The method builds a cubic model in the tangent space and uses a retraction to project steps back onto the manifold, preserving intrinsic geometric feasibility.
  • Local acceleration regimes enable superlinear and quadratic convergence near nondegenerate minimizers, highlighting its practical efficiency and theoretical rigor.

to=arxiv_search 񹚗json {"2query2 OR \2"A Cubic Regularized Newton's Method over Riemannian Manifolds\"","max_results":5,"sort_by":"relevance"} to=search_arxiv ашәаးjson {"2query2 The second-order cubic-regularized Riemannian Newton algorithm is a Newton-type method for minimizing a smooth function over a Riemannian manifold by solving, at each iterate, a cubic-regularized second-order model in the tangent space and then mapping the step back to the manifold by a retraction. In its canonical form, developed as the cubic regularized Riemannian Newton (CRRN) method, it addresses problems of the form

PRESERVED_PLACEHOLDER_2query2^

where PRESERVED_PLACEHOLDER_2(Zhang et al., 2018) OR \2^ is an embedded Riemannian submanifold of a Euclidean space EE. Its defining feature is that it transfers Euclidean cubic-regularized Newton methodology to the manifold setting while preserving geometric feasibility and retaining the worst-case second-order complexity O(1/ϵ3/2)\mathcal O(1/\epsilon^{3/2}) for reaching a second-order ϵ\epsilon-stationary point under stated smoothness and retraction assumptions (&&&2query2&&&).

The algorithm is designed for smooth unconstrained optimization on a manifold, but the unconstrained character is intrinsic rather than ambient: the variable is restricted to MM, and the ambient Euclidean problem is generally nonconvex even when the manifold itself is smooth (&&&2query2&&&). The fundamental geometric objects are the tangent space TxMT_xM, the Riemannian gradient gradf(x)\operatorname{grad} f(x), the Riemannian Hessian Hessf(x)\operatorname{Hess} f(x), and a retraction Retr(x,ξ)Retr(x,\xi) that maps a tangent vector PRESERVED_PLACEHOLDER_2(Zhang et al., 2018) OR \2query2^ back to the manifold.

A central analytic device is the pullback of PRESERVED_PLACEHOLDER_2(Zhang et al., 2018) OR \2(Zhang et al., 2018) OR \2^ at a base point PRESERVED_PLACEHOLDER_2(Zhang et al., 2018) OR \22,

PRESERVED_PLACEHOLDER_2(Zhang et al., 2018) OR \23

To differentiate this object in ambient coordinates, the CRRN analysis uses an extended retraction defined by

PRESERVED_PLACEHOLDER_2(Zhang et al., 2018) OR \24

for ambient vectors PRESERVED_PLACEHOLDER_2(Zhang et al., 2018) OR \25, where PRESERVED_PLACEHOLDER_2(Zhang et al., 2018) OR \26 is the orthogonal projection onto the tangent space (&&&2query2&&&). This yields the identity

PRESERVED_PLACEHOLDER_2(Zhang et al., 2018) OR \27

and a precise relation between the pullback Hessian and the Riemannian Hessian: PRESERVED_PLACEHOLDER_2(Zhang et al., 2018) OR \28

If the retraction is second-order, the additional term vanishes, so the pullback Hessian and the Riemannian Hessian coincide on tangent directions (&&&2query2&&&). This correspondence is structurally important because cubic regularization is formulated in the tangent space, whereas stationarity is defined on the manifold itself. When only a regular retraction is available, the analysis still controls the mismatch through

PRESERVED_PLACEHOLDER_2(Zhang et al., 2018) OR \29

This formulation clarifies a common misconception: the method is not simply Euclidean Newton followed by projection. Feasibility is maintained by a retraction at every iteration, and both the model and the convergence guarantees are expressed in intrinsic Riemannian terms rather than by post hoc projection arguments (&&&2query2&&&).

2. Cubic model and tangent-space subproblem

At iteration EE2query2, the method forms a second-order Taylor model of the pullback together with a cubic regularizer. In tangent coordinates, the model is

EE2(Zhang et al., 2018) OR \2^

Equivalently, with EE2 the orthogonal projection onto EE3,

EE4

Because EE5, the linear term is already tangent (&&&2query2&&&).

The subproblem is an unconstrained cubic-regularized minimization over the tangent space. In this sense, the CRRN step is the exact Riemannian analogue of Euclidean cubic Newton: the manifold enters through the tangent geometry and the retraction, while the local step computation remains a cubic-regularized second-order model minimization (&&&2query2&&&). The iterate update is correspondingly simple: EE6

The global minimizer EE7 of the cubic model satisfies Euclidean-style optimality conditions on the tangent space: EE8 where EE9 and O(1/ϵ3/2)\mathcal O(1/\epsilon^{3/2})2query2^ (&&&2query2&&&). The analysis also uses the bound

O(1/ϵ3/2)\mathcal O(1/\epsilon^{3/2})2(Zhang et al., 2018) OR \2^

which keeps the step inside the region where local pullback Lipschitz estimates are valid.

The cubic term has the same role as in Euclidean cubic regularization: it stabilizes the second-order model, controls step length without an explicit trust-region radius, and enables complexity results at approximate second-order points. This places CRRN in direct contrast with Riemannian trust-region methods, whose subproblems are radius-constrained rather than cubically penalized (&&&2query2&&&).

3. Second-order stationarity and global complexity

The target guarantee is a second-order O(1/ϵ3/2)\mathcal O(1/\epsilon^{3/2})2-stationary point, defined by

O(1/ϵ3/2)\mathcal O(1/\epsilon^{3/2})3

This is the natural Riemannian counterpart of approximate first-order stationarity combined with approximate positive semidefiniteness of the Hessian (&&&2query2&&&).

The main global theorem for CRRN states that the method reaches such a point within O(1/ϵ3/2)\mathcal O(1/\epsilon^{3/2})4 iterations, provided several assumptions hold. The key requirements are that O(1/ϵ3/2)\mathcal O(1/\epsilon^{3/2})5 is compact, the retraction is second-order in the global complexity proof, and the pullback Hessian is locally Lipschitz continuous uniformly on bounded tangent sets: O(1/ϵ3/2)\mathcal O(1/\epsilon^{3/2})6 Compactness is used to guarantee uniform constants and a global lower bound on O(1/ϵ3/2)\mathcal O(1/\epsilon^{3/2})7, while the second-order retraction assumption ensures the cleanest transfer of pullback information to manifold optimality conditions (&&&2query2&&&).

Under a sufficiently large O(1/ϵ3/2)\mathcal O(1/\epsilon^{3/2})8, the method achieves a cubic decrease: O(1/ϵ3/2)\mathcal O(1/\epsilon^{3/2})9 with an explicit positive ϵ\epsilon2query2^ depending on ϵ\epsilon2(Zhang et al., 2018) OR \2^ (&&&2query2&&&). Summation of these decreases yields a bound on ϵ\epsilon2, and choosing the best iterate among ϵ\epsilon3 steps gives the stated complexity. The returned point ϵ\epsilon4 satisfies

ϵ\epsilon5

A further ingredient in the proof is a Lipschitz bound on the smallest Riemannian Hessian eigenvalue,

ϵ\epsilon6

for nearby points on a compact manifold (&&&2query2&&&). This controls the transfer of approximate second-order information between successive tangent spaces.

The complexity statement is notable because it matches the best-known Euclidean worst-case second-order complexity, and the paper explicitly contrasts it with Riemannian trust-region methods, which typically guarantee second-order ϵ\epsilon7-stationarity in ϵ\epsilon8 iterations (&&&2query2&&&). A plausible implication is that cubic regularization is not merely a globalization device on manifolds; it is the mechanism that preserves the canonical Euclidean second-order complexity order in the intrinsic setting.

4. Local superlinear and quadratic convergence

Beyond worst-case guarantees, the CRRN analysis identifies two local acceleration regimes near a local minimizer ϵ\epsilon9 (&&&2query2&&&). The first is based on local gradient domination: MM2query2^ in a neighborhood of MM2(Zhang et al., 2018) OR \2. In that setting, the method improves on the global rate through a sharper descent recursion. The theorem introduces an auxiliary sequence MM2 and shows:

  • if MM3, then MM4 decreases linearly;
  • if MM5, then

MM6

and in particular MM7 gives MM8;

  • if MM9, then the convergence is superlinear of order TxMT_xM2query2^ once the iterates are sufficiently close.

The heuristic behind these regimes is explicit in the paper: the cubic model produces

TxMT_xM2(Zhang et al., 2018) OR \2^

while the next gradient is controlled by

TxMT_xM2

Combined with gradient domination, this yields a stronger recursion than the global analysis (&&&2query2&&&).

The second regime concerns a nondegenerate local minimum. If there exists TxMT_xM3 such that

TxMT_xM4

then the method attains local quadratic convergence. Specifically, once the iterates are sufficiently close and the initial tangent step is sufficiently small,

TxMT_xM5

which implies

TxMT_xM6

The proof uses the positive definiteness of the operator

TxMT_xM7

together with Lipschitz continuity of the Riemannian Hessian to preserve positivity in a neighborhood (&&&2query2&&&).

These local results underscore another important distinction. The cubic regularizer does not preclude Newton-like fast local behavior; rather, when curvature is favorable, the method transitions from globally safeguarded second-order descent to locally quadratic convergence.

5. Retractions, compactness, and manifold-specific structure

Several assumptions in CRRN are geometric rather than purely analytic. The extended retraction is needed to differentiate pullbacks in ambient Euclidean coordinates, and second-order retractions are used to align pullback Hessians with Riemannian Hessians at the origin (&&&2query2&&&). The regularity bounds

TxMT_xM8

are repeatedly invoked in the descent and neighborhood arguments.

Compactness of TxMT_xM9 is also central in the original global theory. It ensures uniform constants, a global lower bound on the objective, and local Lipschitz properties of the pullback Hessian on bounded tangent sets (&&&2query2&&&). This compactness requirement is one of the most visible limitations of the original formulation, and later work on adaptive, inexact, and quasi-Newton variants can be read as efforts to relax or bypass aspects of that framework while preserving second-order guarantees (&&&22(Zhang et al., 2018) OR \2&&&).

The Stiefel specialization in the original paper makes the abstract assumptions concrete. For the Stiefel manifold gradf(x)\operatorname{grad} f(x)2query2,

gradf(x)\operatorname{grad} f(x)2(Zhang et al., 2018) OR \2^

and the polar retraction

gradf(x)\operatorname{grad} f(x)2

is a second-order retraction with gradf(x)\operatorname{grad} f(x)3 (&&&2query2&&&). In this case the paper provides explicit constants for the pullback Hessian Lipschitz bound and gradient comparison, including

gradf(x)\operatorname{grad} f(x)4

This explicit specialization shows that the theory is not purely formal: the regularity assumptions can be verified in an important matrix-manifold setting.

A common misunderstanding is that such constants are merely bookkeeping artifacts. In the CRRN framework they determine the admissible regularization magnitude, the validity region of pullback Taylor estimates, and the quantitative form of the descent inequality (&&&2query2&&&).

Subsequent work has broadened the cubic-regularized Riemannian Newton paradigm in several directions while preserving the core idea of solving a cubic-regularized second-order model on tangent spaces.

Method Distinguishing feature Stated complexity
CRRN (&&&2query2&&&) Exact cubic-regularized Newton on the full tangent space gradf(x)\operatorname{grad} f(x)5 to a second-order gradf(x)\operatorname{grad} f(x)6-stationary point
RDRSOM (Tang et al., 2023) Cubic model solved on a low-dimensional subspace gradf(x)\operatorname{grad} f(x)7 gradf(x)\operatorname{grad} f(x)8 for gradf(x)\operatorname{grad} f(x)9-stationarity under Assumption 2(Zhang et al., 2018) OR \2^
Adaptive inexact ARC on manifolds (&&&22(Zhang et al., 2018) OR \2&&&) Inexact gradient and Hessian with acceptance ratio Hessf(x)\operatorname{Hess} f(x)2query2^ and adaptive Hessf(x)\operatorname{Hess} f(x)2(Zhang et al., 2018) OR \2^ Hessf(x)\operatorname{Hess} f(x)2
Adaptive cubic quasi-Newton (Louzeiro et al., 2024) Inexact Hessf(x)\operatorname{Hess} f(x)3, Hessf(x)\operatorname{Hess} f(x)4, possibly derivative-free finite differences Hessf(x)\operatorname{Hess} f(x)5 first-order and Hessf(x)\operatorname{Hess} f(x)6 second-order
Inexact Sub-RN-CR (Deng et al., 2023) Subsampled gradient and Hessian plus cubic regularization Hessf(x)\operatorname{Hess} f(x)7, improved to Hessf(x)\operatorname{Hess} f(x)8 under stronger subproblem conditions

The dimension-reduced method RDRSOM restricts the cubic model to a subspace Hessf(x)\operatorname{Hess} f(x)9, with a notable two-dimensional choice

Retr(x,ξ)Retr(x,\xi)2query2^

thereby reducing per-iteration cost while retaining an Retr(x,ξ)Retr(x,\xi)2(Zhang et al., 2018) OR \2^ iteration bound under its model-accuracy assumptions (Tang et al., 2023). Its practical Barzilai–Borwein-like extension replaces exact Hessian-vector products by finite differences and drops to Retr(x,ξ)Retr(x,\xi)2 complexity.

Adaptive inexact cubic regularization on manifolds allows Retr(x,ξ)Retr(x,\xi)3 and Retr(x,ξ)Retr(x,\xi)4, uses the model

Retr(x,ξ)Retr(x,\xi)5

and accepts or rejects steps through

Retr(x,ξ)Retr(x,\xi)6

Its main theorem gives termination after at most

Retr(x,ξ)Retr(x,\xi)7

iterations under pullback accuracy, inexactness, and subproblem-quality assumptions (&&&22(Zhang et al., 2018) OR \2&&&).

The adaptive cubic quasi-Newton method on manifolds replaces exact second-order information by approximations satisfying step-dependent accuracy bounds and can be implemented with forward finite differences, including a fully function-evaluation-based Hessian approximation (Louzeiro et al., 2024). The paper explicitly contrasts this adaptive framework with earlier Riemannian cubic methods that required compactness, constants such as Retr(x,ξ)Retr(x,\xi)8, and a fixed Retr(x,ξ)Retr(x,\xi)9 satisfying a complicated lower bound.

Subsampled cubic-regularized Riemannian Newton-type methods address large finite-sum problems by replacing exact gradient and Hessian evaluations with sampled estimates and by solving the cubic subproblem approximately with Lanczos or nonlinear conjugate gradient methods (Deng et al., 2023). These methods aim to preserve the saddle-escape and second-order properties of cubic regularization while reducing the cost of curvature computation.

Later application-driven work has also made the paradigm more concrete. RDRSOM was applied to sensor network localization on a feasible set described as an intersection of spheres with different centers and radii, which the authors characterize as a Riemannian manifold not previously considered in the literature (Tang et al., 2023). A more recent study formulates PRESERVED_PLACEHOLDER_2(Zhang et al., 2018) OR \2query2query2-means clustering as smooth unconstrained optimization over a submanifold and states that a second-order cubic-regularized Riemannian Newton algorithm can be implemented so that each Newton subproblem is solved in linear time in the number of samples after product-manifold factorization (Xu et al., 25 Sep 2025).

These developments clarify the present status of the subject. The exact CRRN method is the foundational second-order cubic-regularized Riemannian Newton algorithm (&&&2query2&&&). Many later methods are best understood not as replacements, but as structural variants: dimension-reduced, adaptive, inexact, quasi-Newton, subsampled, or application-specialized extensions of the same tangent-space cubic-regularization principle (Tang et al., 2023).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Second-Order Cubic-Regularized Riemannian Newton Algorithm.