Papers
Topics
Authors
Recent
Search
2000 character limit reached

Term Sparsity-Based Hierarchy

Updated 29 November 2025
  • Term sparsity-based hierarchy is a framework that exploits the inherent sparsity in polynomial terms to decompose and reduce the dimensionality of convex relaxations.
  • It constructs term-adjacency graphs and employs chordal completions to block-diagonalize semidefinite programs, significantly enhancing computational efficiency.
  • Extensions include symmetry adaptation, noncommutative formulations, and applications to complex polynomial systems, dynamical systems, and regression models.

The term sparsity-based hierarchy is a class of optimization methodologies and algorithmic frameworks that exploit the sparsity of polynomial terms to achieve dramatic decomposition and dimensionality reduction in polynomial optimization and related convex relaxations. Term sparsity refers to the property that, in high-dimensional polynomial systems, only a small subset of all possible monomials is present in each polynomial. By precisely tracking the interactions between these nonzero monomials, rather than variables, one builds graph-theoretic representations which may be iteratively refined and chordally completed; semidefinite programming (SDP) relaxations are then decomposed over maximal cliques in these graphs, yielding block-diagonalized constraints with size dictated by the local term structure and symmetry, rather than overall variable count or relaxation order. Recent advances further extend such term sparsity methodologies to symmetry-adapted bases, complex polynomial, noncommutative polynomial, and hierarchical representations, producing convergent hierarchies of relaxations, block-decomposable moment and localizing matrices, and strong guarantees of tightness under mild regularity conditions (Klep et al., 22 Nov 2025, Wang et al., 2020, Wang et al., 2019, Magron et al., 2021, Wang et al., 2021, Wang et al., 2020, Wang et al., 2018, Balagansky et al., 30 May 2025).

1. Fundamental Framework: Polynomial Optimization and Moment-SOS Hierarchies

Polynomial optimization problems (POPs) are typically formulated as

min⁡x∈Rnf(x)s.t.gj(x)≥0,  j=1,...,m,\min_{x \in \mathbb{R}^n} f(x) \quad \text{s.t.} \quad g_j(x) \ge 0,\; j=1,...,m,

where f,gj∈R[x1,...,xn]f,g_j \in \mathbb{R}[x_1,...,x_n] and the feasible set K={x:gj(x)≥0}K = \{x: g_j(x) \ge 0\} is typically assumed compact (Archimedean). The moment-SOS (sum-of-squares) relaxation hierarchy constructs moment matrices Mt(y)M_t(y) of order tt indexed by multi-indices α,β\alpha,\beta, which encode the candidate moments yαy_\alpha associated to a measure on KK; localizing matrices Mt−dj(gjy)M_{t-d_j}(g_j y) are similarly built for the constraints.

Dense SOS relaxations enforce PSD constraints on Mt(y)M_t(y) and each f,gj∈R[x1,...,xn]f,g_j \in \mathbb{R}[x_1,...,x_n]0, with matrix sizes scaling as f,gj∈R[x1,...,xn]f,g_j \in \mathbb{R}[x_1,...,x_n]1 for f,gj∈R[x1,...,xn]f,g_j \in \mathbb{R}[x_1,...,x_n]2 variables and order f,gj∈R[x1,...,xn]f,g_j \in \mathbb{R}[x_1,...,x_n]3. This rapidly becomes intractable for many moderate-to-large scale applications.

2. Term Sparsity: Graph-Based Decomposition and Block-Diagonalization

Term sparsity is detected by constructing the term-support set

f,gj∈R[x1,...,xn]f,g_j \in \mathbb{R}[x_1,...,x_n]4

where f,gj∈R[x1,...,xn]f,g_j \in \mathbb{R}[x_1,...,x_n]5 denotes the set of exponent vectors f,gj∈R[x1,...,xn]f,g_j \in \mathbb{R}[x_1,...,x_n]6 appearing in polynomial f,gj∈R[x1,...,xn]f,g_j \in \mathbb{R}[x_1,...,x_n]7. We then build a term-adjacency graph f,gj∈R[x1,...,xn]f,g_j \in \mathbb{R}[x_1,...,x_n]8 whose vertices are basis monomials f,gj∈R[x1,...,xn]f,g_j \in \mathbb{R}[x_1,...,x_n]9, and an undirected edge K={x:gj(x)≥0}K = \{x: g_j(x) \ge 0\}0 is present if K={x:gj(x)≥0}K = \{x: g_j(x) \ge 0\}1 appears in K={x:gj(x)≥0}K = \{x: g_j(x) \ge 0\}2, some K={x:gj(x)≥0}K = \{x: g_j(x) \ge 0\}3, or their product: K={x:gj(x)≥0}K = \{x: g_j(x) \ge 0\}4

The sparsity pattern of K={x:gj(x)≥0}K = \{x: g_j(x) \ge 0\}5 encodes interactions between term pairs. Chordal completion and clique decomposition of K={x:gj(x)≥0}K = \{x: g_j(x) \ge 0\}6 allow the moment and localizing matrices to be split into much smaller blocks indexed by maximal cliques. Each block corresponds to a set of monomials which interact according to the true algebraic structure of the POP data.

For example, in (Wang et al., 2019), the TSSOS (Term Sparsity SOS) hierarchy uses a support-extension process followed by block-closure to partition the monomial basis into connected components, each giving rise to a PSD constraint for that component; as the process iterates, the block structure refines and eventually matches that of the dense SOS.

3. Hierarchical Iterative Refinement and Block Closure

The hierarchy is constructed via an iterative process:

  • Support extension: At each sparse order K={x:gj(x)≥0}K = \{x: g_j(x) \ge 0\}7, update the support set by including all K={x:gj(x)≥0}K = \{x: g_j(x) \ge 0\}8 monomials present in current blocks. For each block, declare an edge between indices K={x:gj(x)≥0}K = \{x: g_j(x) \ge 0\}9 if Mt(y)M_t(y)0 has nonempty intersection with current support.
  • Block (clique) closure: Refine each term-sparsity graph by completing connected components to cliques (or by chordal extension).
  • Sparse block SDP construction: For each block, assemble the restricted moment matrix and impose PSD conditions only on these blocks.

The resulting relaxation has the structure: Mt(y)M_t(y)1 where the blocks are much smaller than the full matrix (Klep et al., 22 Nov 2025, Wang et al., 2019, Wang et al., 2020).

This process can be extended to a two-level hierarchy: first exploit variable-based correlative sparsity to partition Mt(y)M_t(y)2-dimensional problems into variable cliques, and second, apply term sparsity within each clique to further reduce block sizes, yielding the CS-TSSOS methodology (Wang et al., 2020).

4. Symmetry-Adapted Hierarchy and Extensions

When the POP exhibits finite group invariance (e.g., under permutation, cyclic, dihedral, or symmetric groups), further decomposition can be achieved by passing to a symmetry-adapted basis. The isotypic decomposition of Mt(y)M_t(y)3 splits the polynomial space into Mt(y)M_t(y)4-invariant subspaces Mt(y)M_t(y)5. The associated basis organizes the moment matrix into blocks indexed by irreducible representations and their multiplicities.

Term sparsity can be exploited directly within each symmetry-adapted block: one builds block-specific support sets Mt(y)M_t(y)6, constructs block-specific term-sparsity graphs and cliques, and imposes PSD conditions only on principal submatrices defined by these cliques. The resulting relaxation reads

Mt(y)M_t(y)7

where each block is restricted to its reduced support. This yields dramatic, often problem-size-independent, block-size reduction for highly symmetric instances (e.g., maximal block size Mt(y)M_t(y)8 for 1D Ising quartic with dihedral symmetry even as Mt(y)M_t(y)9 grows) (Klep et al., 22 Nov 2025).

5. Convergence, Theoretical Guarantees, and Computational Trade-Offs

For the hierarchy where maximal chordal extension is used at each block closure:

  • Finite-step recovery: For fixed tt0, the hierarchy converges after finitely many sparse order steps tt1, i.e., tt2 matches the full symmetric SOS bound tt3 at some finite tt4.
  • Asymptotic convergence: As tt5 and tt6 maximal per tt7, the sparse symmetric hierarchy satisfies tt8, recovering the global optimum.
  • Monotonicity: For each tt9, optimal values increase with the sparse order α,β\alpha,\beta0 and are bounded above by the dense SOS bound. Increasing relaxation order α,β\alpha,\beta1 leads to monotone increase towards the global optimum (Klep et al., 22 Nov 2025, Wang et al., 2020, Wang et al., 2019).

Computationally, block sizes decrease from α,β\alpha,\beta2 to α,β\alpha,\beta3, often yielding orders of magnitude speedup. Applications to symmetric quartic POPs, torus grid quartics, and large contact-rich motion planning demonstrate the scalability and efficacy of the approach (Klep et al., 22 Nov 2025, Kang et al., 5 Feb 2025).

6. Applications and Extensions: Complex, Noncommutative, Dynamical, and Regression Models

Term sparsity-based hierarchies extend beyond real polynomial optimization:

  • Complex POPs: The complex moment-HSOS hierarchy adapts term sparsity via graphs on Hermitian monomials and supports, with convergence and monotonicity guarantees in the sparse order. Implementations achieve block sizes (e.g., α,β\alpha,\beta4 in complex vs. α,β\alpha,\beta5 in real for random quartics) and run-time reduction of α,β\alpha,\beta6 or more (Wang et al., 2021).
  • Noncommutative POPs: NCTSSOS exploits term sparsity patterns among noncommutative words and products (e.g., α,β\alpha,\beta7), yielding small blocks (typically size α,β\alpha,\beta8–α,β\alpha,\beta9) and finite-step convergence to NPA-style relaxations (Wang et al., 2020).
  • Dynamical Systems: Term sparsity yields block-decomposed relaxations for SDP-based region of attraction or invariant set computation; sign symmetries and causality can be layered on top for further savings and convergence (Wang et al., 2021).
  • Hierarchical Regression Priors: Bayesian regression with term-level hierarchical sparsity priors (latent scales yαy_\alpha0) encode heredity (strong/weak) and enable adaptivity, borrowing strength, and model-consistent shrinkage (Griffin et al., 2013).
  • Neural Interpretable Sparse Models: Recent work (HierarchicalTopK) trains a single sparse autoencoder across a spectrum of sparsity budgets, leveraging a prefix-averaged reconstruction loss to maintain disentangled monosemantic features at all sparsity regimes; practical Pareto-optimality in tradeoff between reconstruction error and number of active terms is empirically observed (Balagansky et al., 30 May 2025).

7. Algorithmic Implementation, Software Ecosystem, and Practical Considerations

Efficient implementation involves:

  • Basis construction (standard or Newton polytope basis up to degree yαy_\alpha1).
  • Extraction of initial support, iterative sparse-extension and block closure (or chordal extension), extraction of maximal cliques.
  • Block-assembly of moment and localizing matrices for each clique.
  • Solve block-wise SDP using commercial solvers (MOSEK).
  • Empirical observation: only few iterations (yαy_\alpha2) suffice for equality with dense SOS in most cases.

Major libraries include the Julia package TSSOS for real, complex, and noncommutative polynomial optimization, with options for minimal/maximal chordal completion, block-closure, and model selection via flat extension (Magron et al., 2021, Wang et al., 2019, Wang et al., 2020).

A practical implication is the ability to scale global polynomial optimization (and dynamical system verification) to instances with thousands of variables and constraints previously inaccessible to dense methods, at minimal loss of relaxation tightness.


References

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Term Sparsity-Based Hierarchy.