- The paper introduces degeneracy as the key parameter that sharply characterizes the minimal number of matvec queries needed for a (1+ε)-optimal sparse approximation.
- It develops a polynomial-time, non-adaptive algorithm using masked Gaussian projections to iteratively recover and approximate matrices with a fixed sparsity pattern.
- The results unify and improve existing bounds by providing nearly tight upper and lower query complexities for common structured patterns in scientific computing and machine learning.
Nearly Instance Optimal Sparse Matrix Approximation from Matrix-Vector Products
Problem Setting and Context
The paper formalizes and addresses the problem of learning a sparse approximation to an implicit matrix A∈Rm×n, accessible solely through matrix-vector product (matvec) or adjoint-matvec queries, and constrained to a fixed, known sparsity pattern S∈{0,1}m×n. The setting models scenarios in scientific computing and machine learning where explicit matrix access is prohibitive, and only matvecs are tractable, e.g., Hessians computed via autodiff or implicit operators in SciML. The canonical goal is to produce a matrix B~ with sparsity pattern S such that
∥A−B~∥F≤(1+ϵ)B′:B′=S∘B′min∥A−B′∥F.
Prior literature offers query complexity characterizations in terms of row/column sparsity, total sparsity, or coloring-based metrics (chromatic number of associated conflict graphs), each sometimes optimal and sometimes loose, depending on the pattern geometry. However, these metrics are not universally tight, and coloring-based methods are computationally intractable (NP-hard in general).
Degeneracy as the Query Complexity Parameter
The paper's principal contribution is the identification and theoretical justification of a single combinatorial parameter, the degeneracy degen(S) of the sparsity pattern, that sharply characterizes the minimum query complexity (up to logarithmic and ϵ factors) for achieving near-optimal sparse approximation. The degeneracy is the minimal k such that, repeatedly removing any row or column of S with ≤k ones reduces S∈{0,1}m×n0 to the zero matrix.
This structural parameter unifies and strictly improves previous bounds, yielding nearly instance-optimal, polynomial-time algorithms, and eschewing the NP-hardness of previous coloring-based methods.


Figure 1: Visual depictions of canonical sparsity pattern families: S∈{0,1}m×n1-banded, S∈{0,1}m×n2-modular, and arrowhead, for S∈{0,1}m×n3 matrices with S∈{0,1}m×n4.
Upper and Lower Bounds: The Role of Degeneracy
The authors show two central theorems:
- Upper Bound: There exists a polynomial-time, fully non-adaptive algorithm returning, with high probability, a S∈{0,1}m×n5-optimal sparse approximation using S∈{0,1}m×n6 matvecs.
- Lower Bound: Any (possibly adaptive) algorithm must use at least S∈{0,1}m×n7 matvecs, even for constant or large approximation factors.
The degeneracy parameter simultaneously captures the best-possible query complexities for classical patterns:
- For S∈{0,1}m×n8-banded or S∈{0,1}m×n9-modular forms, B~0, recovering and improving upon prior results.
- For highly structured patterns like the arrowhead family, degeneracy yields constant query complexity where chromatic or total sparsity-based results are loose.
The bounds apply both in the recovery setting (learning B~1 exactly, when B~2 is itself B~3-sparse), as well as in the strictly harder approximation setting (general B~4). The upper bound exploits non-adaptive randomized sketching (masked Gaussian projection), requiring only polynomial pre-processing and post-processing, avoiding the intractability of coloring partition computations.
Algorithmic Framework
The methodology iteratively recovers all rows and columns with at most B~5 non-zeros using randomized matvecs. Specifically:
- For recovery, B~6 matvecs suffice, using Gaussian query vectors and Moore-Penrose pseudoinversion to extract entries with the appropriate sparsity pattern over repeated rounds.
- For approximation, matvec queries proportional to B~7 are issued for each round, and with each round peeling off all rows/columns of sparsity up to B~8. Geometric decay in the problem size ensures only B~9 rounds are required.
- Non-adaptivity is preserved, enabling parallelizable and practical implementation.
- The total time is polynomial in the problem size, including pre-computation (degeneracy computation and query scheduling) as well as post-processing (reconstruction via least squares on masked projections).
The critical technical insight is that, by masking the Gaussian sketches in each round and issuing new sketches for the residual, the dependency structure is controlled and strong accuracy guarantees are maintained.
Numerical and Structural Comparisons
The paper demonstrates, both analytically and by example, that the degeneracy-based framework outperforms or matches the best prior bounds for all common pattern classes.
| Pattern |
Max Row/Col Sparsity |
Chromatic Number |
Total Sparsity |
Degeneracy (This Work) |
| S0-banded |
S1 |
S2 |
S3 |
S4 |
| S5-modular |
S6 |
S7 |
S8 |
S9 |
| Arrowhead |
∥A−B~∥F≤(1+ϵ)B′:B′=S∘B′min∥A−B′∥F.0 |
∥A−B~∥F≤(1+ϵ)B′:B′=S∘B′min∥A−B′∥F.1 |
∥A−B~∥F≤(1+ϵ)B′:B′=S∘B′min∥A−B′∥F.2 |
∥A−B~∥F≤(1+ϵ)B′:B′=S∘B′min∥A−B′∥F.3 |
Degeneracy offers tighter bounds and can yield exponential improvements for structured patterns.
Implications, Extensions, and Open Problems
This work establishes degeneracy as a unified principle for measuring the intrinsic matvec difficulty of structured sparsity-constrained matrix learning in implicit-access models. Practically, this enables the design of adaptive, computationally feasible algorithms for a wide range of scientific and machine learning applications in operator learning, kernel approximation, optimization preconditioning, and analysis of high-dimensional systems.
Potential future directions include:
- Removing the residual ∥A−B~∥F≤(1+ϵ)B′:B′=S∘B′min∥A−B′∥F.4 factor from the upper bound.
- Generalizing the degeneracy-based complexity measure to broader classes of structured matrix families, such as arbitrary linear subspaces or non-linear matrix manifolds.
- Investigating whether "unstructured" (i.e., truly random) query vectors suffice, rather than the masked construction utilized here.
- Extending the analysis to other error norms or more general sketching settings.
Conclusion
The paper provides a comprehensive characterization of the instance-wise complexity of recovering or approximating sparse matrices from matvec queries, introducing degeneracy as the governing parameter. This paradigm delivers nearly tight upper and lower bounds, efficient algorithms, and a unification of prior disparate results, setting a new foundation for sparse approximation in implicit operator settings. The degeneracy metric will likely emerge as the standard measure for query complexity in future work on structured linear algebra in the matvec model.