Information-computation gap for planted linear regression

Determine whether planted linear regression has an information-computation gap in the regime d+1 \le m \ll d^{3/4}n^{1/4}, specifically whether recovering the coefficient vector \beta exactly or with error asymptotically smaller than the ordinary least-squares rate \sqrt{d/n} requires superpolynomial time despite being information-theoretically possible.

Background

Planted linear regression consists of Gaussian covariates X_i \sim N(0,I_d), with exactly m noiseless labels and the remaining labels having standard Gaussian noise. Once m \ge d+1, the true coefficient vector is information-theoretically identifiable because any hyperplane containing at least d+1 noiseless samples is almost surely the true hyperplane.

The paper gives a polynomial-time estimator above the threshold m \gg d{3/4}n{1/4} and an SQ lower bound below essentially the same threshold. The intermediate regime d+1 \le m \ll d{3/4}n{1/4} therefore remains the proposed computationally hard region, although exact recovery remains statistically possible.

References

In the regime $d + 1 \le m \ll d{3/4}n{1/4}$, recovering $\beta$ is information-theoretically possible, yet we know no polynomial-time algorithm; we conjecture that this is an information-computation gap, and we prove a statistical query (SQ) lower bound in support of this claim (\cref{thm:SQ}).

Algorithms for adaptive and heteroskedastic linear regression at the computational threshold  (2608.18402 - Compton et al., 19 Aug 2026) in Section 1, paragraph “Planted linear regression”; also stated in the Abstract

The work of conjectures that any polynomial-time algorithm requires $m = \Omega(\sqrt{nd})$ to exactly recover $\beta$, and they provide Statistical Query (SQ) lower bounds to support their claim (the SQ lower bounds are not yet tight to their conjecture, this is left as an open problem).

Algorithms for adaptive and heteroskedastic linear regression at the computational threshold  (2608.18402 - Compton et al., 19 Aug 2026) in Section 1, subsection “Related work,” paragraph “Oblivious contamination and noiseless linear regression”

The work of conjectures that any polynomial-time algorithm requires $m = \Omega(\sqrt{nd})$ to exactly recover $\beta$, and they provide Statistical Query (SQ) lower bounds to support their claim (the SQ lower bounds are not yet tight to their conjecture, this is left as an open problem).

Algorithms for adaptive and heteroskedastic linear regression at the computational threshold  (2608.18402 - Compton et al., 19 Aug 2026) in Section 1, subsection “Related work,” paragraph “Oblivious contamination and noiseless linear regression”