Dimension complexity implied by adaptive statistical-query learning

Prove that every binary hypothesis class learnable under every input distribution by a randomized (m,τ)-statistical-query algorithm has dimension complexity at most C·m/τ² for a universal constant C, without assuming an additional finite-dimensional polynomial-rank certificate for the seed-averaged terminal responses.

Background

The source problem concerns whether adaptive statistical-query learning under arbitrary distributions forces a low-dimensional shared feature representation. The paper obtains a conditional theorem assuming that the span of seed-averaged terminal responses over all complete response rules has a prescribed finite polynomial rank.

That conditional rank assumption is not derived from the SQ interface. The paper identifies the unresolved task as obtaining the source-scale O(m/τ²) compression directly from adaptive SQ learning.

References

The remaining technical problem is to derive the source-scale compression from the SQ interface itself: if $F0_{\mathcal D,h}$ is the seed-averaged terminal predictor under the canonical exact-center policy, then

\dim\operatorname{span}{F0_{\mathcal D,h}:\mathcal D,h} \le C m/\tau2.

VALG: An Agentic System for ML Theory Research  (2608.13060 - Zhang et al., 13 Aug 2026) in Section 4, Subsection “Is the Power of Deep Learning over Linear Models Inherently Distribution Dependent?”, Subproblem 2: Statistical Query Learning