Exact Risk Ratios for Weighted Data Selection in Linear Regression
Abstract: Hanneke, Moran, Shlimovich and Yehudayoff (COLT 2025) posed the following open problem. A selector sees a finite dataset , picks at most examples together with nonnegative weights, and hands the weighted least squares objective to the minimum-norm ERM. Writing for the worst-case ratio between the loss of the returned predictor on all of and the optimal loss, they proved for $n<d$, and for , and asked for the value in the open regime $d<n<2d$. We determine this value in several cases. For every we prove , which confirms a claim stated without proof in the original note. We further prove and , the two smallest cells not covered by the endpoint formula. For every intermediate budget we prove the lower bound , where is an explicit harmonic quantity over balanced partitions, and we show that this bound is the exact minimax value over the class of datasets whose whitened gradient systems carry an orthogonal circuit-block structure. All three exact values match , and we conjecture that equality holds throughout the open regime. The upper bound proofs run on a common geometric spine: a rigidity theorem for positive spanning configurations of loss gradients, classifications and structural reductions of small positive bases in and , and a dimension-free extremal-basis argument that converts sign-cone geometry into five-point selections. We also give explicit counterexamples showing that several shorter routes fail, and constructive polynomial-time selection algorithms for all proved cases.
Paper Prompts
Sign up for free to create and run prompts on this paper.