- The paper establishes the Biorthogonal Welch Floor, a geometric lower bound that limits linear readouts to O(d^(3/2)) capacity due to unavoidable cross-talk.
- It contrasts linear interfaces with threshold recovery, demonstrating that threshold-based mechanisms bypass strict geometric constraints to achieve nearly quadratic feature loads.
- It synthesizes concepts from compressed sensing and mechanistic interpretability to propose a taxonomy of recursive capacity regimes and outlines open problems in nonlinear reset design.
Geometric Bounds and Interface Invariants in Computation in Superposition
Introduction
"Linear-Readout Floors and Threshold Recovery in Computation in Superposition" (2605.01192) investigates the computational capacity of neural networks operating in the superposition regime, focusing on the distinction between linear and threshold-based interfaces for recursive computation with sparse Boolean features. The paper rigorously analyzes and formalizes the geometric constraints imposed by linear readouts, derives tight lower bounds for cross-talk, and contrasts these with thresholded Boolean recovery mechanisms, clarifying the compatibility scales of several prominent recent constructions. The work synthesizes results from compressed sensing, feature geometry, and mechanistic interpretability to establish a taxonomy of capacity regimes available to neural architectures.
Superposition and Recent Capacity Regimes
The superposition hypothesis posits that neural networks represent or compute more conceptual features than neurons, making superposition a central concept in mechanistic interpretability. Two approaches to recursive computation in superposition have produced divergent capacity regimes:
- Hänni et al.: Their ε-linear recursive template achieves O(d3/2) computable features in width-d networks, relying on linear decodability with error-correction layers.
- Adler and Shavit: Their threshold-based Boolean recovery leverages margin-based thresholding to enable near-quadratic capacity, O(d2/logd), avoiding the accumulation of linear error.
The resulting gap spans a factor of O(d1/2) between active computable feature loads.
Biorthogonal Welch Floor: Linear Readout Bound
The primary theoretical contribution is a geometric constraint on linear readout interfaces, termed the Biorthogonal Welch Floor. For a feature code Ψ∈Rd×F and a readout G∈RF×d with unit diagonal, the worst-case off-diagonal cross-talk is bounded below:
F(F−1)1i=j∑∣(GΨ)ij∣2≥d(F−1)F−d
When F≫d, this implies unavoidable cross-talk of order Ω(d−1/2). The bound is shown to be tight for unit-norm tight frames. As a corollary, small-error O(d3/2)0-linear reset with O(d3/2)1 is impossible for superlinear feature loads, explaining why linear recursive templates (as in Hänni et al.) naturally stop at the O(d3/2)2 scale.
Threshold Recovery and Interface Distinction
Threshold recovery mechanisms, such as those used by Adler and Shavit, are not obstructed by the Welch floor. Coherence-based threshold recovery is analyzed: for a unit-norm code with coherence O(d3/2)3, exact recovery for O(d3/2)4-sparse states by thresholding is feasible if O(d3/2)5 for noise O(d3/2)6. Random feature codes with O(d3/2)7 yield coherence O(d3/2)8, so threshold recovery succeeds for O(d3/2)9 even at quadratic feature loads.
Crucially, linear readout interfaces require every inactive score to be d0, whereas threshold recovery only demands that aggregate interference across the sparse support remains below a fixed margin. This interface distinction is formally separated, demonstrating that thresholded Boolean recovery can coexist with geometric constraints that linear readout faces.
Distributional Separation of Recovery Criteria
A distributional analysis quantifies the separation:
- Linear Readout Error: Any unit-diagonal linear readout at d1 must incur average squared per-coordinate error d2 for Bernoulli d3-sparse states.
- Threshold Recovery Success: Random-support threshold recovery in the same regime succeeds with probability d4 for supports of size d5.
This demonstrates non-equivalence between linear readout and threshold recovery invariants and clarifies why threshold-based frameworks escape linear interface floor constraints.
Interpolation and Open Problems
An open-problem framework is formulated around robust nonlinear or threshold resets. Any reset mechanism attaining input tolerance scaling as d6 would achieve feature load d7, for d8. However, such resets necessarily leave the d9-linear class to avoid contradictions with the Welch floor. This interpolation formalizes the gap and identifies the construction of such resets as a central open problem in recursive computation in superposition.
Taxonomy of Capacity Regimes
The paper synthesizes capacity scales from different domains:
- Compressed Sensing: Passive storage, O(d2/logd)0, (O(d2/logd)1 from compressed sensing theory).
- Recursive Linear Computation: Hänni-compatible, O(d2/logd)2.
- Thresholded Recovery: Adler–Shavit, between O(d2/logd)3 and O(d2/logd)4.
- Passive Representation: Johnson–Lindenstrauss, O(d2/logd)5.
The hierarchy is not monotonic for all tasks; recursive computation, passive representation, and linear decodability measure distinct capacities.
Implications and Future Directions
The geometric floor derived for linear interfaces mathematically limits the recursive capacity attainable by linear decoders in superposition. This necessitates either a change in recursion invariants or the introduction of nonlinear interfaces for scaling capacity beyond O(d2/logd)6. Threshold recovery mechanisms, leveraging margin separation and sparse interference, exemplify such invariants and are compatible with quadratic-scale feature loads in randomized settings.
The distinction between passive storage and active recursive computation is formalized, and implications extend to neural architecture design: error-correction, thresholding, and nonlinear reset modules may become indispensable for exploiting higher-dimensional superpositional capacities. Future theoretical work is required to construct explicit robust threshold resets, establish universal recursive capacity bounds, and explain empirical scaling phenomena observed in sparse autoencoder studies.
Conclusion
This paper rigorously delineates the geometric obstruction imposed on linear interfaces in computation in superposition. The Biorthogonal Welch Floor shows that linear readouts at superlinear feature loads incur unavoidable cross-talk, explaining the compatibility scale of the Hänni recursive template. Threshold recovery, especially in random feature codes, escapes this obstruction, enabling near-quadratic capacity as achieved by Adler and Shavit. The distinction between linear and threshold invariants underpins both theoretical taxonomy and practical architecture design. The central open problem is the explicit construction of robust nonlinear or threshold reset mechanisms that extend recursive computable capacity beyond the linear floor, potentially unlocking new scaling laws and interpretability regimes for large neural systems.