Memory–sample lower bounds for noiseless Gaussian regression
The following describes the scope of the Lean formalization related to the following accompanying paper(s):
Scope
In noiseless Gaussian regression, a learner observes inner products of an unknown unit vector in Rd with independent standard Gaussian vectors. The formalization proves that, for every fixed A>0, there is cA>0 such that, in sufficiently large dimension, a learner retaining at most Ad2 bits between observations needs at least cAdlog(1/ε) observations to attain angular error at most ε with probability at least 2/3 for a uniformly random unit signal, for 0<ε≤1/10.
For memory o(d2), the constant can be universal, and the selected statement also permits success at least 2/3 separately for every signal. The linked supporting results include inverse Gram-matrix moments, exact-projection densities, regularization bounds, and sphere- and cube-prior block estimates.
The paper studies how posterior replicas—independent signals drawn conditionally on the same observed data—control information in noiseless Gaussian regression. The formalization establishes the equal-label identity for finite measures, its density and measurable forms, and information bounds for Gaussian, distance, Haar, synthetic, and incidence comparisons. These statements retain exact observation labels, independent side randomness, and randomized transition rules.
For a learner using M(d)=o(d2) bits of memory, the resulting bound says that, for sufficiently large d and 0<ε≤1/10, success probability at least 2/3 under the uniform unit-sphere prior requires T≥cdlog(1/ε) observations, where c>0 is universal. The formalized streaming result uses the 2/3 threshold; the stronger 3/5 variant is outside this scope.
The formalization controls information gained from exact Gaussian observations of a spherical image of the uniform cube. Under the stated entropy and size conditions, repeated localization over t observation blocks produces nested disclosures with mutual information at most C(n+1)t and expected final localization level at most Ct. The selected estimates include inverse-volume control, kernel coercivity, and drift for the actual observation rows.
For a law of finite relative entropy with respect to a uniform cell law in sufficiently large dimension, the regularization procedure terminates almost surely in regular cells. Its expected depth and number of attempts are finite, with at most twice one plus the expected depth in expectation; the remaining relative entropy plus (log2)n/4 times the expected depth is at most the initial relative entropy plus 2/e.
The paper uses projection moments and domination by positive spherical caps to bound the information carried by exact Gaussian observations of an unknown unit vector. The formalization covers integrated frame and sphere moments, cap domination, Riesz-density estimates with fixed offsets, and bounds valid at every radius. It also covers probability bounds for finite-memory learners at a fixed stopping time and after summing over stopping times.
For M(d)=o(d2) bits of memory, sufficiently large d, and 0<ε≤1/10, success probability at least 2/3 under the uniform sphere prior requires T≥cdlog(1/ε) observations for a universal c>0. A corresponding bound holds when M≤Ad2 for each fixed A>0, with the constant allowed to depend on A. The randomized learner model allows independent shared seeds; its reduction to a finite-state model is proved for dimensions at least three, with the auxiliary dimension-two case outside this scope.
The paper compares fresh Gaussian observations with observations coupled to information already held in memory. The formalization covers mixed-moment and actual-row estimates, critical-radius, one-level, and two-label comparisons, and conditional-information bounds across a block of observations. For sufficiently large dimension d, the relevant entropy bounds of order d2 bound the information in the fresh experiment by that in the coupled experiment plus a term of order d; the associated block information increase also has order d.
It also covers two-point identities on observation fibers and the fiber-block bound, including the boundary case of d−2 observed rows. The uniform unit-sphere prior, Gaussian rows, finite memory states, and randomized transition rules remain part of these statements. The separate streaming consequences and several intermediate comparison constructions are outside this selection.
The paper uses subspheres to bound the number of exact Gaussian observations needed to estimate a unit vector with limited memory. The formalization proves that, for M(d)=o(d2) bits of memory, sufficiently large d, and 0<ε≤1/10, success probability at least 2/3 under the uniform sphere prior, or for every unit signal, requires at least 2−16dlog2(1/ε) observations. Under the uniform prior, a lower bound of the same order with a universal positive constant holds at success probability 1/2.
The formalization also covers radius-weighted block and terminal estimates, affine and dimension estimates, and a vector-valued beta-mixture comparison with its boundary cases. These statements retain their dimension, width, radius, and suffix hypotheses, as well as the finite-state learner's independent randomness and restrictions on stopping and output.
Comparator links