Papers
Topics
Authors
Recent
Search
2000 character limit reached

Nearly Instance Optimal Sparse Matrix Approximation from Matrix-Vector Products

Published 10 Jun 2026 in cs.DS and math.NA | (2606.12179v1)

Abstract: A large body of work studies the problem of learning an approximation to an implicit matrix A∈R<sup>m×</sup>nA\in \mathbb{R}<sup>{m\times</sup> n} that is only accessible implicitly via matrix-vector product queries (matvec queries) of the form x→Ax{x} \rightarrow {A}{x} or x→A<sup>Tx{x} \rightarrow {A}<sup>T{x}. Of particular interest are methods that learn a near-optimal approximation with a fixed sparsity pattern. For example, we might want to learn a near-optimal diagonal, banded, or arrow-head approximation to an implicit matrix AA. Naturally, the number of matvec queries required to solve this problem depends on the sparsity pattern, which can be encoded as a binary matrix S∈0,1<sup>m×</sup>n{S}\in {0,1}<sup>{m\times</sup> n}. The query complexity of previous algorithms scales with quantities like the total number of ones in S{S}, its maximum column/row sparsity, or the chromatic number of a its "conflict graph". These quantities are incomparable: for a given S{S}, parameterizing by one might yield lower query complexity than another. In this work, we unify and tighten these prior results by providing a nearly sharp characterization of the matvec query complexity of sparse matrix approximation. Generalizing a definition from graph algorithms, let the degeneracy, degen(S){degen}({S}), denote the smallest number kk so that, if we iteratively delete all rows and columns of S{S} with ≤k\leq k ones, we are left with an empty matrix. We show that a near-optimal approximation to AA with sparsity pattern SS can be learned with O~(degen(S))\tilde{O}({degen}({S})) matrix-vector product queries, and Ω(degen(S))Ω({degen}({S})) queries are necessary, for any sparsity pattern S{S}. Moreover, unlike prior work based on graph coloring, all of our methods run in polynomial time.

Authors (2)

Summary

  • The paper introduces degeneracy as the key parameter that sharply characterizes the minimal number of matvec queries needed for a (1+ε)-optimal sparse approximation.
  • It develops a polynomial-time, non-adaptive algorithm using masked Gaussian projections to iteratively recover and approximate matrices with a fixed sparsity pattern.
  • The results unify and improve existing bounds by providing nearly tight upper and lower query complexities for common structured patterns in scientific computing and machine learning.

Nearly Instance Optimal Sparse Matrix Approximation from Matrix-Vector Products

Problem Setting and Context

The paper formalizes and addresses the problem of learning a sparse approximation to an implicit matrix A∈Rm×nA \in \mathbb{R}^{m \times n}, accessible solely through matrix-vector product (matvec) or adjoint-matvec queries, and constrained to a fixed, known sparsity pattern S∈{0,1}m×nS \in \{0,1\}^{m \times n}. The setting models scenarios in scientific computing and machine learning where explicit matrix access is prohibitive, and only matvecs are tractable, e.g., Hessians computed via autodiff or implicit operators in SciML. The canonical goal is to produce a matrix B~\tilde{B} with sparsity pattern SS such that

∥A−B~∥F≤(1+ϵ)min⁡B′:B′=S∘B′∥A−B′∥F.\|A - \tilde{B}\|_F \leq (1+\epsilon) \min_{B': B' = S \circ B'} \|A - B'\|_F.

Prior literature offers query complexity characterizations in terms of row/column sparsity, total sparsity, or coloring-based metrics (chromatic number of associated conflict graphs), each sometimes optimal and sometimes loose, depending on the pattern geometry. However, these metrics are not universally tight, and coloring-based methods are computationally intractable (NP-hard in general).

Degeneracy as the Query Complexity Parameter

The paper's principal contribution is the identification and theoretical justification of a single combinatorial parameter, the degeneracy degen(S)\mathrm{degen}(S) of the sparsity pattern, that sharply characterizes the minimum query complexity (up to logarithmic and ϵ\epsilon factors) for achieving near-optimal sparse approximation. The degeneracy is the minimal kk such that, repeatedly removing any row or column of SS with ≤k\leq k ones reduces S∈{0,1}m×nS \in \{0,1\}^{m \times n}0 to the zero matrix.

This structural parameter unifies and strictly improves previous bounds, yielding nearly instance-optimal, polynomial-time algorithms, and eschewing the NP-hardness of previous coloring-based methods. Figure 1

Figure 1

Figure 1

Figure 1: Visual depictions of canonical sparsity pattern families: S∈{0,1}m×nS \in \{0,1\}^{m \times n}1-banded, S∈{0,1}m×nS \in \{0,1\}^{m \times n}2-modular, and arrowhead, for S∈{0,1}m×nS \in \{0,1\}^{m \times n}3 matrices with S∈{0,1}m×nS \in \{0,1\}^{m \times n}4.

Upper and Lower Bounds: The Role of Degeneracy

The authors show two central theorems:

  • Upper Bound: There exists a polynomial-time, fully non-adaptive algorithm returning, with high probability, a S∈{0,1}m×nS \in \{0,1\}^{m \times n}5-optimal sparse approximation using S∈{0,1}m×nS \in \{0,1\}^{m \times n}6 matvecs.
  • Lower Bound: Any (possibly adaptive) algorithm must use at least S∈{0,1}m×nS \in \{0,1\}^{m \times n}7 matvecs, even for constant or large approximation factors.

The degeneracy parameter simultaneously captures the best-possible query complexities for classical patterns:

  • For S∈{0,1}m×nS \in \{0,1\}^{m \times n}8-banded or S∈{0,1}m×nS \in \{0,1\}^{m \times n}9-modular forms, B~\tilde{B}0, recovering and improving upon prior results.
  • For highly structured patterns like the arrowhead family, degeneracy yields constant query complexity where chromatic or total sparsity-based results are loose.

The bounds apply both in the recovery setting (learning B~\tilde{B}1 exactly, when B~\tilde{B}2 is itself B~\tilde{B}3-sparse), as well as in the strictly harder approximation setting (general B~\tilde{B}4). The upper bound exploits non-adaptive randomized sketching (masked Gaussian projection), requiring only polynomial pre-processing and post-processing, avoiding the intractability of coloring partition computations.

Algorithmic Framework

The methodology iteratively recovers all rows and columns with at most B~\tilde{B}5 non-zeros using randomized matvecs. Specifically:

  • For recovery, B~\tilde{B}6 matvecs suffice, using Gaussian query vectors and Moore-Penrose pseudoinversion to extract entries with the appropriate sparsity pattern over repeated rounds.
  • For approximation, matvec queries proportional to B~\tilde{B}7 are issued for each round, and with each round peeling off all rows/columns of sparsity up to B~\tilde{B}8. Geometric decay in the problem size ensures only B~\tilde{B}9 rounds are required.
  • Non-adaptivity is preserved, enabling parallelizable and practical implementation.
  • The total time is polynomial in the problem size, including pre-computation (degeneracy computation and query scheduling) as well as post-processing (reconstruction via least squares on masked projections).

The critical technical insight is that, by masking the Gaussian sketches in each round and issuing new sketches for the residual, the dependency structure is controlled and strong accuracy guarantees are maintained.

Numerical and Structural Comparisons

The paper demonstrates, both analytically and by example, that the degeneracy-based framework outperforms or matches the best prior bounds for all common pattern classes.

Pattern Max Row/Col Sparsity Chromatic Number Total Sparsity Degeneracy (This Work)
SS0-banded SS1 SS2 SS3 SS4
SS5-modular SS6 SS7 SS8 SS9
Arrowhead ∥A−B~∥F≤(1+ϵ)min⁡B′:B′=S∘B′∥A−B′∥F.\|A - \tilde{B}\|_F \leq (1+\epsilon) \min_{B': B' = S \circ B'} \|A - B'\|_F.0 ∥A−B~∥F≤(1+ϵ)min⁡B′:B′=S∘B′∥A−B′∥F.\|A - \tilde{B}\|_F \leq (1+\epsilon) \min_{B': B' = S \circ B'} \|A - B'\|_F.1 ∥A−B~∥F≤(1+ϵ)min⁡B′:B′=S∘B′∥A−B′∥F.\|A - \tilde{B}\|_F \leq (1+\epsilon) \min_{B': B' = S \circ B'} \|A - B'\|_F.2 ∥A−B~∥F≤(1+ϵ)min⁡B′:B′=S∘B′∥A−B′∥F.\|A - \tilde{B}\|_F \leq (1+\epsilon) \min_{B': B' = S \circ B'} \|A - B'\|_F.3

Degeneracy offers tighter bounds and can yield exponential improvements for structured patterns.

Implications, Extensions, and Open Problems

This work establishes degeneracy as a unified principle for measuring the intrinsic matvec difficulty of structured sparsity-constrained matrix learning in implicit-access models. Practically, this enables the design of adaptive, computationally feasible algorithms for a wide range of scientific and machine learning applications in operator learning, kernel approximation, optimization preconditioning, and analysis of high-dimensional systems.

Potential future directions include:

  • Removing the residual ∥A−B~∥F≤(1+ϵ)min⁡B′:B′=S∘B′∥A−B′∥F.\|A - \tilde{B}\|_F \leq (1+\epsilon) \min_{B': B' = S \circ B'} \|A - B'\|_F.4 factor from the upper bound.
  • Generalizing the degeneracy-based complexity measure to broader classes of structured matrix families, such as arbitrary linear subspaces or non-linear matrix manifolds.
  • Investigating whether "unstructured" (i.e., truly random) query vectors suffice, rather than the masked construction utilized here.
  • Extending the analysis to other error norms or more general sketching settings.

Conclusion

The paper provides a comprehensive characterization of the instance-wise complexity of recovering or approximating sparse matrices from matvec queries, introducing degeneracy as the governing parameter. This paradigm delivers nearly tight upper and lower bounds, efficient algorithms, and a unification of prior disparate results, setting a new foundation for sparse approximation in implicit operator settings. The degeneracy metric will likely emerge as the standard measure for query complexity in future work on structured linear algebra in the matvec model.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 5 likes about this paper.