Papers
Topics
Authors
Recent
Search
2000 character limit reached

Agnostic learning in (almost) optimal time via Gaussian surface area

Published 6 Mar 2026 in cs.LG, cs.DS, and stat.ML | (2603.06027v1)

Abstract: The complexity of learning a concept class under Gaussian marginals in the difficult agnostic model is closely related to its L1L_1-approximability by low-degree polynomials. For any concept class with Gaussian surface area at most ΓΓ, Klivans et al. (2008) show that degree d=O(Γ<sup>2</sup>/ε<sup>4)d = O(Γ<sup>2</sup> / \varepsilon<sup>4) suffices to achieve an ε\varepsilon-approximation. This leads to the best-known bounds on the complexity of learning a variety of concept classes. In this note, we improve their analysis by showing that degree d=O~(Γ<sup>2</sup>/ε<sup>2)d = \tilde O (Γ<sup>2</sup> / \varepsilon<sup>2) is enough. In light of lower bounds due to Diakonikolas et al. (2021), this yields (near) optimal bounds on the complexity of agnostically learning polynomial threshold functions in the statistical query model. Our proof relies on a direct analogue of a construction of Feldman et al. (2020), who considered L1L_1-approximation on the Boolean hypercube.

Summary

  • The paper proves that every Boolean function with Gaussian surface area at most Γ has an L1 polynomial approximation of degree O(Γ² log(1/ε)/ε²), improving the prior O(Γ²/ε⁴) bound.
  • The method combines Ornstein–Uhlenbeck smoothing, Gaussian noise sensitivity, and truncated Hermite expansions to balance approximation error against polynomial degree.
  • The result yields agnostic learners running in n^{Ō(Γ²/ε²)} time and nearly matches known lower bounds for degree-k polynomial threshold functions, while leaving the logarithmic factor and several lower-bound gaps open.

Overview

This paper, by Pesenti, Slot, and Wiedmer (2603.06027), improves the degree bounds for L1L_1-polynomial approximation of Boolean functions under the standard Gaussian distribution, and consequently the runtime of the L1L_1-polynomial regression algorithm for agnostic learning. The central result is that any measurable function f:Rn{±1}f : \mathbb{R}^n \to \{\pm 1\} with Gaussian surface area (GSA) at most Γ\Gamma admits a polynomial of degree

d=O ⁣(log(1/ε)Γ2ε2)d = O\!\left(\frac{\log(1/\varepsilon)\cdot \Gamma^2}{\varepsilon^2}\right)

with EXN[f(X)p(X)]ε\mathbb{E}_{X \sim N}[\lvert f(X) - p(X)\rvert] \leq \varepsilon. This improves the classical bound of Klivans, O'Donnell, and Servedio (d=O(Γ2/ε4)d = O(\Gamma^2/\varepsilon^4)) by a factor of roughly Ω~(1/ε2)\tilde{\Omega}(1/\varepsilon^2), and recovers — up to a single log(1/ε)\log(1/\varepsilon) factor — the optimal O(1/ε2)O(1/\varepsilon^2) bound known specifically for halfspaces via a construction of Diakonikolas, Kane, and Nelson.

Background and context

In the agnostic learning model of Kearns, Schapire, and Sellie, an algorithm receives labeled examples from an arbitrary joint distribution L1L_10 on L1L_11 and must output a hypothesis with error at most L1L_12, where L1L_13 is the error of the best concept in the target class. Under Gaussian marginals, the standard algorithm is L1L_14-polynomial regression: compute the best degree-L1L_15 polynomial fit to the labels in L1L_16-norm (via linear programming), then threshold. This runs in time L1L_17 and yields excess error L1L_18 whenever every concept in the class admits a degree-L1L_19, f:Rn{±1}f : \mathbb{R}^n \to \{\pm 1\}0-accurate f:Rn{±1}f : \mathbb{R}^n \to \{\pm 1\}1-approximation. The choice of f:Rn{±1}f : \mathbb{R}^n \to \{\pm 1\}2 over f:Rn{±1}f : \mathbb{R}^n \to \{\pm 1\}3 is essential: f:Rn{±1}f : \mathbb{R}^n \to \{\pm 1\}4 regression only guarantees error f:Rn{±1}f : \mathbb{R}^n \to \{\pm 1\}5, tolerating limited noise.

The complexity-theoretic picture is essentially settled. Diakonikolas, Kane, Pittas, and Zarifis showed that if f:Rn{±1}f : \mathbb{R}^n \to \{\pm 1\}6 is the smallest degree needed to f:Rn{±1}f : \mathbb{R}^n \to \{\pm 1\}7-approximate a class in f:Rn{±1}f : \mathbb{R}^n \to \{\pm 1\}8, then the SQ-complexity of agnostically learning that class up to error f:Rn{±1}f : \mathbb{R}^n \to \{\pm 1\}9 is approximately Γ\Gamma0. Thus upper bounds on the approximation degree translate directly into near-optimal algorithmic guarantees.

Prior work obtained such bounds through GSA. Klivans–O'Donnell–Servedio proved that any concept with GSA at most Γ\Gamma1 has an Γ\Gamma2-approximation of degree Γ\Gamma3, which converts to Γ\Gamma4 via Cauchy–Schwarz. This bound was suboptimal even for halfspaces, where a direct construction achieves Γ\Gamma5, matching both the lower bound of Ganzburg and the SQ lower bounds. That construction, however, does not extend beyond halfspaces without incurring dimension-dependent factors or worse Γ\Gamma6-dependence. The question addressed here is whether a guarantee can be simultaneously general (all bounded-GSA concepts) and optimal in Γ\Gamma7.

Main results

The main theorem states that for any measurable Γ\Gamma8 and any Γ\Gamma9, there exists a polynomial of degree d=O ⁣(log(1/ε)Γ2ε2)d = O\!\left(\frac{\log(1/\varepsilon)\cdot \Gamma^2}{\varepsilon^2}\right)0 approximating d=O ⁣(log(1/ε)Γ2ε2)d = O\!\left(\frac{\log(1/\varepsilon)\cdot \Gamma^2}{\varepsilon^2}\right)1 within d=O ⁣(log(1/ε)Γ2ε2)d = O\!\left(\frac{\log(1/\varepsilon)\cdot \Gamma^2}{\varepsilon^2}\right)2 in d=O ⁣(log(1/ε)Γ2ε2)d = O\!\left(\frac{\log(1/\varepsilon)\cdot \Gamma^2}{\varepsilon^2}\right)3. The intermediate result is a noise-sensitivity parametrized bound: for each d=O ⁣(log(1/ε)Γ2ε2)d = O\!\left(\frac{\log(1/\varepsilon)\cdot \Gamma^2}{\varepsilon^2}\right)4 and degree d=O ⁣(log(1/ε)Γ2ε2)d = O\!\left(\frac{\log(1/\varepsilon)\cdot \Gamma^2}{\varepsilon^2}\right)5, there is a degree-d=O ⁣(log(1/ε)Γ2ε2)d = O\!\left(\frac{\log(1/\varepsilon)\cdot \Gamma^2}{\varepsilon^2}\right)6 polynomial d=O ⁣(log(1/ε)Γ2ε2)d = O\!\left(\frac{\log(1/\varepsilon)\cdot \Gamma^2}{\varepsilon^2}\right)7 with

d=O ⁣(log(1/ε)Γ2ε2)d = O\!\left(\frac{\log(1/\varepsilon)\cdot \Gamma^2}{\varepsilon^2}\right)8

where d=O ⁣(log(1/ε)Γ2ε2)d = O\!\left(\frac{\log(1/\varepsilon)\cdot \Gamma^2}{\varepsilon^2}\right)9 is the probability that EXN[f(X)p(X)]ε\mathbb{E}_{X \sim N}[\lvert f(X) - p(X)\rvert] \leq \varepsilon0 disagrees on two EXN[f(X)p(X)]ε\mathbb{E}_{X \sim N}[\lvert f(X) - p(X)\rvert] \leq \varepsilon1-correlated Gaussians. Combining this with the Ledoux-based inequality EXN[f(X)p(X)]ε\mathbb{E}_{X \sim N}[\lvert f(X) - p(X)\rvert] \leq \varepsilon2 and choosing EXN[f(X)p(X)]ε\mathbb{E}_{X \sim N}[\lvert f(X) - p(X)\rvert] \leq \varepsilon3 appropriately yields the main theorem.

Via the KKMS reduction, this immediately gives an agnostic learner running in time EXN[f(X)p(X)]ε\mathbb{E}_{X \sim N}[\lvert f(X) - p(X)\rvert] \leq \varepsilon4 for any class with GSA bounded by EXN[f(X)p(X)]ε\mathbb{E}_{X \sim N}[\lvert f(X) - p(X)\rvert] \leq \varepsilon5. Concrete consequences include:

Concept class Previous UB New UB Known LB
Halfspaces EXN[f(X)p(X)]ε\mathbb{E}_{X \sim N}[\lvert f(X) - p(X)\rvert] \leq \varepsilon6 EXN[f(X)p(X)]ε\mathbb{E}_{X \sim N}[\lvert f(X) - p(X)\rvert] \leq \varepsilon7 EXN[f(X)p(X)]ε\mathbb{E}_{X \sim N}[\lvert f(X) - p(X)\rvert] \leq \varepsilon8
Degree-EXN[f(X)p(X)]ε\mathbb{E}_{X \sim N}[\lvert f(X) - p(X)\rvert] \leq \varepsilon9 PTFs d=O(Γ2/ε4)d = O(\Gamma^2/\varepsilon^4)0 d=O(Γ2/ε4)d = O(\Gamma^2/\varepsilon^4)1 d=O(Γ2/ε4)d = O(\Gamma^2/\varepsilon^4)2
Intersections of d=O(Γ2/ε4)d = O(\Gamma^2/\varepsilon^4)3 halfspaces d=O(Γ2/ε4)d = O(\Gamma^2/\varepsilon^4)4 d=O(Γ2/ε4)d = O(\Gamma^2/\varepsilon^4)5 d=O(Γ2/ε4)d = O(\Gamma^2/\varepsilon^4)6
Convex sets d=O(Γ2/ε4)d = O(\Gamma^2/\varepsilon^4)7 d=O(Γ2/ε4)d = O(\Gamma^2/\varepsilon^4)8

The most consequential case is degree-d=O(Γ2/ε4)d = O(\Gamma^2/\varepsilon^4)9 PTFs: the new bound Ω~(1/ε2)\tilde{\Omega}(1/\varepsilon^2)0 nearly matches the SQ lower bound of Ω~(1/ε2)\tilde{\Omega}(1/\varepsilon^2)1, so Ω~(1/ε2)\tilde{\Omega}(1/\varepsilon^2)2-polynomial regression is now known to be near-optimal for this class. For intersections of halfspaces and convex sets, the improvement is a factor of Ω~(1/ε2)\tilde{\Omega}(1/\varepsilon^2)3 over prior best bounds.

Proof technique

The proof is short and modular. It proceeds in two steps: first approximate Ω~(1/ε2)\tilde{\Omega}(1/\varepsilon^2)4 by its Ornstein–Uhlenbeck smoothing Ω~(1/ε2)\tilde{\Omega}(1/\varepsilon^2)5, whose Ω~(1/ε2)\tilde{\Omega}(1/\varepsilon^2)6 error equals Ω~(1/ε2)\tilde{\Omega}(1/\varepsilon^2)7 exactly; second, approximate Ω~(1/ε2)\tilde{\Omega}(1/\varepsilon^2)8 by its truncated Hermite expansion, whose Ω~(1/ε2)\tilde{\Omega}(1/\varepsilon^2)9 error is at most log(1/ε)\log(1/\varepsilon)0 because log(1/ε)\log(1/\varepsilon)1 damps Hermite coefficients by log(1/ε)\log(1/\varepsilon)2. The triangle inequality gives the intermediate proposition; substituting the GSA-to-noise-sensitivity inequality and balancing the two terms gives the theorem.

The authors are explicit that the technical contribution is modest: the construction is a direct Gaussian analogue of a lemma of Feldman, Kothari, and Vondrák, who established the same two-step scheme on the Boolean hypercube using the Boolean noise operator and Fourier truncation. All ingredients were available in the literature; the contribution lies in assembling them for the Gaussian setting and observing the resulting optimality.

Comparison with previous constructions

Three comparisons sharpen the picture. Against Klivans–O'Donnell–Servedio, the authors identify precisely where the old argument loses: reducing the entire problem to log(1/ε)\log(1/\varepsilon)3 via Cauchy–Schwarz is inherently lossy, since even origin-centered halfspaces have log(1/ε)\log(1/\varepsilon)4 Hermite-truncation error log(1/ε)\log(1/\varepsilon)5. Notably, they prove that the plain Hermite truncation log(1/ε)\log(1/\varepsilon)6 itself achieves the improved rate for origin-centered halfspaces — namely log(1/ε)\log(1/\varepsilon)7, via Plancherel–Rotach asymptotics and Christoffel–Darboux identities — implying the suboptimality of the old bound stems from the analysis, not the choice of approximating polynomial. They note, however, that they are not aware of an example showing log(1/ε)\log(1/\varepsilon)8 fails to match the new guarantee in general, leaving open whether the smoothing step is necessary.

Against Diakonikolas–Kane–Nelson, the paper explains their construction as a smoothing with uniformly bounded derivatives, whose Hermite coefficients decay fast enough by Gaussian integration by parts. Their method achieves the optimal log(1/ε)\log(1/\varepsilon)9 for halfspaces (better than the new result by one logarithmic factor), but extending it to higher dimensions appears to incur dimension-dependent factors; for PTFs it requires a much worse dependence on O(1/ε2)O(1/\varepsilon^2)0. The new result thus occupies the previously vacant position of a general, near-optimal guarantee.

Against Feldman–Kothari–Vondrák, the correspondence is exact: the Gaussian proposition mirrors their Boolean lemma term-for-term, with the Ornstein–Uhlenbeck operator replacing the hypercube noise operator and Hermite truncation replacing Fourier truncation.

Limitations and open questions

Several caveats bear directly on the strength of the results. First, the main bound carries a O(1/ε2)O(1/\varepsilon^2)1 factor, so it matches the optimal halfspace bound only up to this factor; whether the logarithm can be removed for all bounded-GSA classes is not resolved. Second, the lower-bound side is incomplete: for intersections of O(1/ε2)O(1/\varepsilon^2)2 halfspaces, the best lower bound O(1/ε2)O(1/\varepsilon^2)3 does not match the new upper bound O(1/ε2)O(1/\varepsilon^2)4, and no lower bounds are known for convex sets. Third, the general noise-sensitivity lower bound of Diakonikolas et al. is typically loose — for halfspaces it gives O(1/ε2)O(1/\varepsilon^2)5 versus the true O(1/ε2)O(1/\varepsilon^2)6 — so tightness of the new upper bounds cannot be certified through it except in special cases. Finally, as noted above, it remains open whether the unsmoothed Hermite truncation O(1/ε2)O(1/\varepsilon^2)7 already satisfies the improved guarantee for all bounded-GSA concepts, which would simplify the construction further.

Conclusion

The paper establishes that degree O(1/ε2)O(1/\varepsilon^2)8 suffices for O(1/ε2)O(1/\varepsilon^2)9-approximation of any bounded-GSA concept under Gaussian marginals, improving the long-standing L1L_100 bound and yielding near-optimal SQ-complexity guarantees for agnostically learning PTFs, intersections of halfspaces, and convex sets via L1L_101-polynomial regression. The proof transfers a known Boolean-hypercube construction to the Gaussian setting, and the accompanying analysis of the Hermite truncation of the sign function clarifies why earlier analyses were lossy. The remaining gaps — the extraneous logarithmic factor and the missing matching lower bounds for several classes — delineate the precise boundary of what this work leaves unresolved.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 2 likes about this paper.