Papers
Topics
Authors
Recent
Search
2000 character limit reached

Support-Basis Decomposition Overview

Updated 14 July 2026
  • Support-Basis Decomposition is a method that defines global objects via components with explicit support and basis roles, enabling clear geometric and combinatorial interpretations.
  • It organizes positive spanning sets through simplicial and conical decompositions, precisely characterizing minimal positive bases and maximal negatively independent subsets.
  • Applications range from efficient sparse–dense attention approximations and lossless multi-head attention reformulations to locally supported basis decompositions in neural field models.

Support-basis decomposition denotes a family of decomposition schemes in which a global object is represented through components that have explicit support and basis-like roles. In the foundational geometric formulation, it concerns the decomposition of a finite set XRdX\subset \mathbb R^d that positively spans Rd\mathbb R^d into minimal positive bases and maximal negatively independent subsets (Schoch, 2020). In later literature, closely related terminology is used for efficient attention algorithms, sparse additive decompositions after orthogonal basis transforms, decomposition of probability marginals via admissible support candidates, and neural architectures built from locally supported basis contributions (Aliakbarpour et al., 2 Oct 2025, Zhao, 2 Oct 2025, Ba et al., 2024, Matuschke, 2023, Moseley et al., 2021, Kumar et al., 24 May 2026).

1. Positively spanning sets and simplicial decomposition

For a finite set XRdX\subset \mathbb R^d, the linear span and positive span are

(X)={iaixi:xiX, aiR},(X)={iaixi:xiX, ai0}.\ell(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\in\mathbb R\Bigr\},\qquad \P(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\ge 0\Bigr\}.

A positive spanning set for a subspace LRdL\subset \mathbb R^d is a set XX with (X)=L\P(X)=L. A positive basis is a positive spanning set that is inclusion-minimal. A simplex is a subset SXS\subset X such that 0(S)0\in \P(S) and (T)(S)\P(T)\subsetneq \P(S) for every proper subset Rd\mathbb R^d0; equivalently, it is a minimal positive basis. Schoch denotes the family of simplices by Rd\mathbb R^d1 and studies when a positively spanning set can be reconstructed from such components (Schoch, 2020).

The central structural condition is the factorisation condition,

Rd\mathbb R^d2

where Rd\mathbb R^d3. The main characterization states that Rd\mathbb R^d4 is a positive basis of Rd\mathbb R^d5 if and only if it is the union of at most Rd\mathbb R^d6 simplices and satisfies the factorisation condition. In that case,

Rd\mathbb R^d7

for some linear basis Rd\mathbb R^d8, with each Rd\mathbb R^d9 for a unique XRdX\subset \mathbb R^d0, and the sets XRdX\subset \mathbb R^d1 are pairwise disjoint. Each simplex is then of the form XRdX\subset \mathbb R^d2 (Schoch, 2020).

This yields the standard cardinality bounds for positive bases in XRdX\subset \mathbb R^d3: XRdX\subset \mathbb R^d4 Equality XRdX\subset \mathbb R^d5 occurs if and only if XRdX\subset \mathbb R^d6 is a cross, namely a union of XRdX\subset \mathbb R^d7 XRdX\subset \mathbb R^d8-simplices derived from a linear basis, exemplified by XRdX\subset \mathbb R^d9. In this setting, support-basis decomposition is not an approximation scheme but an exact classification of positive bases by their simplex structure.

2. Subbases, Boolean lattices, and conical decomposition

The same geometric framework has a second layer built from negatively independent subsets. A set (X)={iaixi:xiX, aiR},(X)={iaixi:xiX, ai0}.\ell(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\in\mathbb R\Bigr\},\qquad \P(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\ge 0\Bigr\}.0 is negatively independent if (X)={iaixi:xiX, aiR},(X)={iaixi:xiX, ai0}.\ell(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\in\mathbb R\Bigr\},\qquad \P(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\ge 0\Bigr\}.1 for every (X)={iaixi:xiX, aiR},(X)={iaixi:xiX, ai0}.\ell(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\in\mathbb R\Bigr\},\qquad \P(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\ge 0\Bigr\}.2; equivalently, (X)={iaixi:xiX, aiR},(X)={iaixi:xiX, ai0}.\ell(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\in\mathbb R\Bigr\},\qquad \P(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\ge 0\Bigr\}.3 is pointed. Maximal negatively independent subsets are collected in

(X)={iaixi:xiX, aiR},(X)={iaixi:xiX, ai0}.\ell(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\in\mathbb R\Bigr\},\qquad \P(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\ge 0\Bigr\}.4

These subsets are the maximal pointed-cone frames carried by (X)={iaixi:xiX, aiR},(X)={iaixi:xiX, ai0}.\ell(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\in\mathbb R\Bigr\},\qquad \P(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\ge 0\Bigr\}.5 (Schoch, 2020).

For any positively spanning set (X)={iaixi:xiX, aiR},(X)={iaixi:xiX, ai0}.\ell(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\in\mathbb R\Bigr\},\qquad \P(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\ge 0\Bigr\}.6, the subbases

(X)={iaixi:xiX, aiR},(X)={iaixi:xiX, ai0}.\ell(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\in\mathbb R\Bigr\},\qquad \P(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\ge 0\Bigr\}.7

form a Boolean lattice under inclusion. Moreover, the map (X)={iaixi:xiX, aiR},(X)={iaixi:xiX, ai0}.\ell(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\in\mathbb R\Bigr\},\qquad \P(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\ge 0\Bigr\}.8 embeds (X)={iaixi:xiX, aiR},(X)={iaixi:xiX, ai0}.\ell(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\in\mathbb R\Bigr\},\qquad \P(X)=\Bigl\{\sum_i a_i x_i:x_i\in X,\ a_i\ge 0\Bigr\}.9 into the Boolean algebra LRdL\subset \mathbb R^d0, and if LRdL\subset \mathbb R^d1 is a positive basis then this embedding is onto, so LRdL\subset \mathbb R^d2. This identifies the combinatorics of positively spanning subcollections with the combinatorics of simplex families. A plausible implication is that the factorisation condition does not merely constrain intersections of spans; it organizes the full subbasis structure into a distributive combinatorial object.

Schoch’s second main theorem is conical rather than simplicial: any finite LRdL\subset \mathbb R^d3 can be written as the union of at most LRdL\subset \mathbb R^d4 maximal sets spanning pointed cones. If LRdL\subset \mathbb R^d5 is a positive basis, these sets are tantamount to frames of the cones. The bound is sharp if and only if LRdL\subset \mathbb R^d6 is a cross. More generally, there can be at most LRdL\subset \mathbb R^d7 maximal subsets of LRdL\subset \mathbb R^d8 spanning pointed cones when intersections of two of them do not span a set of full dimension (Schoch, 2020).

Low-dimensional examples make the extremal cases explicit. In LRdL\subset \mathbb R^d9, a triangle with XX0 is a single simplex, whereas the planar cross XX1 attains XX2 and XX3. In XX4, a tetrahedron with XX5 is again a single simplex, while the standard cross XX6 attains XX7. These examples separate the simplex count, which is at most XX8, from the pointed-cone count, which is at most XX9.

3. Sparse–dense support decomposition in softmax attention

A distinct usage appears in efficient attention, where support-basis decomposition is introduced to approximate softmax attention beyond the bounded-entry assumption. The framework assumes that the entries of (X)=L\P(X)=L0 are independent centered sub-Gaussian random variables with variance proxies (X)=L\P(X)=L1, and uses this tail behavior to split each matrix into large and small entries at a threshold (X)=L\P(X)=L2 (Aliakbarpour et al., 2 Oct 2025).

With

(X)=L\P(X)=L3

the method defines the “support basis” of (X)=L\P(X)=L4 by extracting the entries activated by (X)=L\P(X)=L5: (X)=L\P(X)=L6 The matrices (X)=L\P(X)=L7 and (X)=L\P(X)=L8 are disjoint and sum to (X)=L\P(X)=L9. Under the stated tail bounds, SXS\subset X0 and SXS\subset X1 with high probability. The sparse part is handled exactly, while the dense bounded part is approximated by a degree-SXS\subset X2 Chebyshev polynomial SXS\subset X3 on SXS\subset X4, giving a low-rank factorization SXS\subset X5. Polynomial-kernel sketching with Count-Sketch or OSNAP reduces the cost of applying the approximation (Aliakbarpour et al., 2 Oct 2025).

The single-threshold algorithm builds SXS\subset X6 explicitly on its sparse support, constructs SXS\subset X7 so that SXS\subset X8, forms

SXS\subset X9

computes normalization terms from the row sums, and returns

0(S)0\in \P(S)0

which approximates 0(S)0\in \P(S)1. The stated runtime is 0(S)0\in \P(S)2 for the sparse construction plus 0(S)0\in \P(S)3 for the Chebyshev factorization and 0(S)0\in \P(S)4 for sketching, yielding overall sub-quadratic time 0(S)0\in \P(S)5 for some 0(S)0\in \P(S)6. The accuracy theorem states that if 0(S)0\in \P(S)7 have at most 0(S)0\in \P(S)8 entries exceeding 0(S)0\in \P(S)9, then for any (T)(S)\P(T)\subsetneq \P(S)0 the algorithm returns (T)(S)\P(T)\subsetneq \P(S)1 in (T)(S)\P(T)\subsetneq \P(S)2 time such that

(T)(S)\P(T)\subsetneq \P(S)3

The multi-threshold extension removes all distributional assumptions by partitioning entry magnitudes into exponentially growing buckets (T)(S)\P(T)\subsetneq \P(S)4, producing (T)(S)\P(T)\subsetneq \P(S)5 disjoint blocks (T)(S)\P(T)\subsetneq \P(S)6. The support-basis and disjointness guarantee

(T)(S)\P(T)\subsetneq \P(S)7

The paper further states that combining these blockwise polynomial attentions gives a better global fit in (T)(S)\P(T)\subsetneq \P(S)8 norms, for (T)(S)\P(T)\subsetneq \P(S)9, than any single polynomial. Empirically, on Rd\mathbb R^d00 Gaussian inputs, the method becomes faster than exact attention once Rd\mathbb R^d01, and on LLaDA-8B-Instruct its degree-6 approximation is reported at approximately Rd\mathbb R^d02 average accuracy versus approximately Rd\mathbb R^d03 for exact attention, while the Alman–Song baseline fails at Rd\mathbb R^d04 (Aliakbarpour et al., 2 Oct 2025).

4. Basis decomposition as an exact reformulation of multi-head attention

A second attention-related usage, presented as Basis Decomposition (BD), is exact rather than approximate. Its starting point is the matrix identity that if Rd\mathbb R^d05 has rank Rd\mathbb R^d06, then after selecting Rd\mathbb R^d07 linearly independent rows to form a basis matrix Rd\mathbb R^d08, the remaining rows are reconstructed by unique coefficients collected in Rd\mathbb R^d09, and

Rd\mathbb R^d10

in “row-first” form, or Rd\mathbb R^d11 in “row-last” form (Zhao, 2 Oct 2025).

In multi-head attention, each head has matrices Rd\mathbb R^d12, and the products Rd\mathbb R^d13 and Rd\mathbb R^d14 have rank at most Rd\mathbb R^d15. Applying BD with Rd\mathbb R^d16 yields per-head basis matrices and support coefficients: Rd\mathbb R^d17 These are stacked across heads into shared Rd\mathbb R^d18 and Rd\mathbb R^d19. The corresponding transformed projections satisfy

Rd\mathbb R^d20

so the scaled dot products, softmax weights, and reconstructed outputs are preserved exactly. The method is therefore described as the first lossless algorithmic reformulation of attention (Zhao, 2 Oct 2025).

The algorithm has an offline preparation stage and an inference stage. Preparation computes two BD decompositions for each Rd\mathbb R^d21 and Rd\mathbb R^d22, chooses the orientation with smaller average residual, and stores the resulting basis and coefficient matrices. Inference computes Rd\mathbb R^d23, reconstructs Rd\mathbb R^d24 and Rd\mathbb R^d25 through the coefficient matrices, applies standard scaled-dot-product attention headwise, concatenates the outputs, and multiplies by Rd\mathbb R^d26 to produce the final output. The stated savings factor in key/value projection FLOPs and parameters is

Rd\mathbb R^d27

For Rd\mathbb R^d28 and Rd\mathbb R^d29, the reduction is Rd\mathbb R^d30. On DeepSeek-V2-Lite (16B, FP16), the reported figures are 4 seconds of offline preparation, Rd\mathbb R^d31 faster key/value projections in FP16, Rd\mathbb R^d32 in BF16, Rd\mathbb R^d33 smaller KV weights, and end-to-end perplexity changes of Rd\mathbb R^d34 in FP32 and Rd\mathbb R^d35 in FP16 (Zhao, 2 Oct 2025).

A common misconception is to treat the two attention papers as variants of the same algorithm. They are not. In the support-basis attention framework, the decomposition is sparse–dense and the result is a provably accurate approximation. In BD Attention, the decomposition is a rank-exact factorization of projection products and the result preserves every pairwise dot product and final projected output exactly.

5. Basis transforms and sparse additive function decompositions

In high-dimensional approximation, support-basis decomposition refers to finding an orthogonal transform under which an ANOVA or anchored decomposition becomes sparse. For a Rd\mathbb R^d36 function Rd\mathbb R^d37 on a box Rd\mathbb R^d38, both ANOVA and anchored representations take the form

Rd\mathbb R^d39

with the distinction lying in whether coordinates are integrated out or fixed at an anchor point. Sparsity is characterized through the function graph, whose vertices are Rd\mathbb R^d40 and edges are Rd\mathbb R^d41. Minimal additive decompositions use only cliques of this graph (Ba et al., 2024).

The key idea is that Rd\mathbb R^d42 may not be sparse in the ambient coordinates, but Rd\mathbb R^d43 may be sparse for some orthogonal Rd\mathbb R^d44. The directional derivatives transform as

Rd\mathbb R^d45

so choosing Rd\mathbb R^d46 to annihilate most first and second directional derivatives yields a sparse graph Rd\mathbb R^d47. The proposed algorithm has three steps. First, gradient samples are assembled into

Rd\mathbb R^d48

and an SVD is used to minimize the number of active coordinates. Second, Hessian samples Rd\mathbb R^d49 are approximately joint-block-diagonalized by a commutant-based procedure. Third, residual off-block edges are sparsified by minimizing

Rd\mathbb R^d50

over Rd\mathbb R^d51 by Riemannian gradient descent or the Landing method (Ba et al., 2024).

The resulting transformed representation is typically

Rd\mathbb R^d52

with Rd\mathbb R^d53 small and Rd\mathbb R^d54 containing only the surviving mixed-derivative pairs. Because the ANOVA and anchored terms are minimal, the decomposition into at-most-two-variable summands is unique. The reported numerics include a synthetic Rd\mathbb R^d55 example in which step 1 reduces the active dimension to Rd\mathbb R^d56, step 2 identifies two connected components of sizes Rd\mathbb R^d57 and Rd\mathbb R^d58, and step 3 drives off-block entries to Rd\mathbb R^d59 within Rd\mathbb R^d60 Riemannian-gradient iterations, as well as a noisy Rd\mathbb R^d61 case recovering Rd\mathbb R^d62 with more than Rd\mathbb R^d63 success over Rd\mathbb R^d64 trials (Ba et al., 2024).

6. Admissible-support decomposition of probability marginals

In combinatorial optimization, the decomposition problem begins with a finite ground set Rd\mathbb R^d65, a family of subsets Rd\mathbb R^d66, a requirement function Rd\mathbb R^d67, and a marginal vector Rd\mathbb R^d68. The goal is to find a distribution Rd\mathbb R^d69 on Rd\mathbb R^d70 such that

Rd\mathbb R^d71

or certify infeasibility. The associated feasible marginal set is

Rd\mathbb R^d72

and every Rd\mathbb R^d73 satisfies the covering inequalities

Rd\mathbb R^d74

(Matuschke, 2023).

The decomposition algorithm is driven by admissible support candidates (ASCs). For residual marginals Rd\mathbb R^d75 and residual requirements Rd\mathbb R^d76, a set Rd\mathbb R^d77 is an ASC if it satisfies three conditions: Rd\mathbb R^d78, Rd\mathbb R^d79 for every tight requirement set Rd\mathbb R^d80, and Rd\mathbb R^d81 for every non-dominated Rd\mathbb R^d82 with Rd\mathbb R^d83. Domination is defined by a partial order Rd\mathbb R^d84 if either Rd\mathbb R^d85 or

Rd\mathbb R^d86

Algorithm 1 repeatedly selects an ASC Rd\mathbb R^d87, chooses a peeling weight Rd\mathbb R^d88, updates the distribution mass Rd\mathbb R^d89, and reduces residual marginals and requirements. The invariants guarantee nonnegativity of residual marginals, preservation of the covering inequalities, and strict progress. Termination occurs after at most Rd\mathbb R^d90 iterations, and the overall runtime is Rd\mathbb R^d91 (Matuschke, 2023).

When a suitable ASC oracle exists, the framework yields an exact polyhedral characterization: Rd\mathbb R^d92 The paper constructs ASCs for supermodular requirements, abstract networks with weak conservation, and Hoffman–Schwartz-type lattice polyhedra. It also characterizes balanced hypergraphs as precisely those systems Rd\mathbb R^d93 for which every Rd\mathbb R^d94 admits a perfect decomposition achieving

Rd\mathbb R^d95

Here the “support” in support-basis decomposition is set-theoretic rather than geometric, and the basis role is played by the selected ASC subsets supporting the distribution.

7. Local support, partition-of-unity structure, and neural field decompositions

In scientific machine learning, the decomposition is spatial and local. Finite Basis PINNs (FBPINNs) represent the approximate solution of a differential equation on Rd\mathbb R^d96 as

Rd\mathbb R^d97

where Rd\mathbb R^d98 is partitioned into overlapping subdomains Rd\mathbb R^d99, XRdX\subset \mathbb R^d00 is a smooth window with compact or near-compact support, XRdX\subset \mathbb R^d01 rescales coordinates to a standard range such as XRdX\subset \mathbb R^d02, and XRdX\subset \mathbb R^d03 is a constraining operator that enforces boundary conditions exactly (Moseley et al., 2021). A strict partition of unity may be imposed by XRdX\subset \mathbb R^d04, but the framework can also sum XRdX\subset \mathbb R^d05 directly. The loss is the strong-form residual

XRdX\subset \mathbb R^d06

with no boundary-loss term and no interface-matching penalty. The stated motivation is to mitigate spectral bias and reduce optimization stiffness by replacing one large network with many smaller local networks.

Courant adopts a related but attention-based local-support decomposition. It places latent anchors XRdX\subset \mathbb R^d07 in the computational domain and computes decoder cross-attention weights

XRdX\subset \mathbb R^d08

so that XRdX\subset \mathbb R^d09 and XRdX\subset \mathbb R^d10 for every head XRdX\subset \mathbb R^d11 and point XRdX\subset \mathbb R^d12. Because the decoder is affine in the latent values, the predicted field decomposes into local contributions XRdX\subset \mathbb R^d13, and the model can be written in basis-expansion form

XRdX\subset \mathbb R^d14

with XRdX\subset \mathbb R^d15 induced by the attention weights and XRdX\subset \mathbb R^d16 determined by the latent state (Kumar et al., 24 May 2026).

The two frameworks differ architecturally—FBPINNs use overlapping subdomain networks and window functions, whereas Courant uses geometry-anchored latent queries, shared random Fourier features, and an affine decoder—but both impose locality structurally rather than through an explicit sparsity regularizer. Courant is trained only with an XRdX\subset \mathbb R^d17 prediction loss in physical space, yet its latents are reported to specialize geometrically: in steady 2D cylinder flow, the top-8 largest XRdX\subset \mathbb R^d18 fields tile the boundary layer and wake; in transient vortex shedding, some anchors remain pinned to the cylinder wall while others track moving vortices, and the latent trajectory exhibits a PSD peak at the physical shedding frequency XRdX\subset \mathbb R^d19 (Kumar et al., 24 May 2026). FBPINNs similarly use local normalization and local support to handle large and multi-scale domains without interface penalties (Moseley et al., 2021).

Across these usages, the term does not denote a single standardized formalism. In positive-span theory it classifies finite vector configurations via simplices and pointed cones; in attention it names either an approximate sparse–dense split or an exact rank factorization; in additive decomposition and marginal decomposition it organizes sparsity through basis transforms or admissible support candidates; and in neural PDE and surrogate models it refers to locally supported basis contributions assembled into a global field. A common pattern is the replacement of a monolithic global object by finitely many components whose supports are explicit and whose recombination is algebraically controlled.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Support-Basis Decomposition.