Papers
Topics
Authors
Recent
Search
2000 character limit reached

Minimax Convergence Rates

Updated 4 February 2026
  • Minimax Convergence Rates are metrics that quantify the optimal speed at which any estimator converges to the true parameter value over a class of models.
  • They provide benchmarks for assessing the efficiency of statistical methods in nonparametric, high-dimensional, and privacy-constrained settings.
  • Proof techniques such as Le Cam’s method, Fano’s inequality, and metric entropy are key to establishing tight minimax lower bounds.

A minimax convergence rate quantifies the fundamental speed at which any estimator can approach the true value of a statistical parameter, uniformly over a model class, when the loss is measured in expectation over worst-case data-generating distributions. The notion of minimax optimality is central to statistical decision theory and nonparametric inference, and it provides benchmarks for evaluating the efficiency of statistical procedures in both classical and modern high-dimensional or privacy-constrained regimes.

1. Formal Definition and General Framework

The minimax risk for a parameter class Θ\Theta and loss function L(θ^,θ)L(\hat\theta, \theta) is defined by

Rn=infθ^supθΘEθ[L(θ^n,θ)],R_n^* = \inf_{\hat\theta} \sup_{\theta \in \Theta} \mathbb{E}_{\theta}[L(\hat\theta_n, \theta)],

where θ^n\hat\theta_n ranges over all estimators measurable with respect to the observed nn data points. The minimax convergence rate is the sequence (rn)n1(r_n)_{n\geq 1} such that RnrnR_n^* \asymp r_n (i.e., bounded above and below by constant multiples of rnr_n for all large nn). This rate encapsulates both the statistical complexity of the model class Θ\Theta and the analytic properties of the loss function L(θ^,θ)L(\hat\theta, \theta)0.

In more complex settings, such as those with additional privacy constraints, dependency structures, adversarial perturbations, or partial information, the minimax rate quantifies the exact effect of these features on achievable estimation or prediction accuracy.

2. Canonical Rates and Dependence on Model Complexity

In classical nonparametric estimation and regression, the minimax rate is determined by the interplay between the function class's "smoothness" and the sample size. For example:

  • Gaussian mean, parametric models: L(θ^,θ)L(\hat\theta, \theta)1 for L(θ^,θ)L(\hat\theta, \theta)2.
  • Hölder class regression, sup-norm loss: L(θ^,θ)L(\hat\theta, \theta)3 for L(θ^,θ)L(\hat\theta, \theta)4 on L(θ^,θ)L(\hat\theta, \theta)5, where L(θ^,θ)L(\hat\theta, \theta)6 is the smoothness parameter (Peng et al., 2024).
  • Sobolev class, L(θ^,θ)L(\hat\theta, \theta)7 loss: L(θ^,θ)L(\hat\theta, \theta)8 for L(θ^,θ)L(\hat\theta, \theta)9-smooth functions in Rn=infθ^supθΘEθ[L(θ^n,θ)],R_n^* = \inf_{\hat\theta} \sup_{\theta \in \Theta} \mathbb{E}_{\theta}[L(\hat\theta_n, \theta)],0 dimensions (Zhao et al., 2023).

For estimation of discrete distributions with Rn=infθ^supθΘEθ[L(θ^n,θ)],R_n^* = \inf_{\hat\theta} \sup_{\theta \in \Theta} \mathbb{E}_{\theta}[L(\hat\theta_n, \theta)],1 categories, under no privacy constraints, the minimax risk is Rn=infθ^supθΘEθ[L(θ^n,θ)],R_n^* = \inf_{\hat\theta} \sup_{\theta \in \Theta} \mathbb{E}_{\theta}[L(\hat\theta_n, \theta)],2 in squared error; for multinomial estimation under Rn=infθ^supθΘEθ[L(θ^n,θ)],R_n^* = \inf_{\hat\theta} \sup_{\theta \in \Theta} \mathbb{E}_{\theta}[L(\hat\theta_n, \theta)],3-local differential privacy, the minimax risk is Rn=infθ^supθΘEθ[L(θ^n,θ)],R_n^* = \inf_{\hat\theta} \sup_{\theta \in \Theta} \mathbb{E}_{\theta}[L(\hat\theta_n, \theta)],4, with sharp constants (Duchi et al., 2013).

Minimax rates change in the following settings:

  • Privacy constraints: Effective sample size is scaled by Rn=infθ^supθΘEθ[L(θ^n,θ)],R_n^* = \inf_{\hat\theta} \sup_{\theta \in \Theta} \mathbb{E}_{\theta}[L(\hat\theta_n, \theta)],5 when enforcing Rn=infθ^supθΘEθ[L(θ^n,θ)],R_n^* = \inf_{\hat\theta} \sup_{\theta \in \Theta} \mathbb{E}_{\theta}[L(\hat\theta_n, \theta)],6-local differential privacy (Duchi et al., 2013).
  • Spatial inhomogeneity: Global rates depend on the design density's vanishing at isolated points; for density Rn=infθ^supθΘEθ[L(θ^n,θ)],R_n^* = \inf_{\hat\theta} \sup_{\theta \in \Theta} \mathbb{E}_{\theta}[L(\hat\theta_n, \theta)],7 near Rn=infθ^supθΘEθ[L(θ^n,θ)],R_n^* = \inf_{\hat\theta} \sup_{\theta \in \Theta} \mathbb{E}_{\theta}[L(\hat\theta_n, \theta)],8, minimax Rn=infθ^supθΘEθ[L(θ^n,θ)],R_n^* = \inf_{\hat\theta} \sup_{\theta \in \Theta} \mathbb{E}_{\theta}[L(\hat\theta_n, \theta)],9-risk scales as θ^n\hat\theta_n0 for Besov ball smoothness θ^n\hat\theta_n1 (Antoniadis et al., 2011).
  • Supersmooth deconvolution: Estimation rates become logarithmic in θ^n\hat\theta_n2:

θ^n\hat\theta_n3

for deconvolution with θ^n\hat\theta_n4-supersmooth noise and loss θ^n\hat\theta_n5 (Wasserstein metric) (Dedecker et al., 2013).

3. Representative Results in Key Models

Discrete Probability Estimation under Local Privacy

Given θ^n\hat\theta_n6 privatized samples from an unknown θ^n\hat\theta_n7-dimensional multinomial θ^n\hat\theta_n8, with privacy parameter θ^n\hat\theta_n9, the sharp minimax rate in squared nn0 error is

nn1

achievable by randomized response or Laplace-noise schemes with moment-matching debiasing (Duchi et al., 2013). For small nn2, privacy reduces effective sample size from nn3 to nn4.

Smooth Density Estimation under Local Privacy

For nn5 in a Sobolev class of smoothness nn6 on nn7:

  • Non-private (classical): nn8
  • nn9-locally private: (rn)n1(r_n)_{n\geq 1}0 Thus privacy worsens the polynomial exponent, and minimax estimation becomes strictly slower unless (rn)n1(r_n)_{n\geq 1}1 (Duchi et al., 2013).

Nonparametric Regression with Inhomogeneous Design

If the design density (rn)n1(r_n)_{n\geq 1}2 has a zero of order (rn)n1(r_n)_{n\geq 1}3 at (rn)n1(r_n)_{n\geq 1}4 (i.e., (rn)n1(r_n)_{n\geq 1}5 near (rn)n1(r_n)_{n\geq 1}6), for smoothness index (rn)n1(r_n)_{n\geq 1}7, the minimax global (rn)n1(r_n)_{n\geq 1}8-risk over a Besov ball satisfies

(rn)n1(r_n)_{n\geq 1}9

with logarithmic rates for exponential zeros. Adaptive wavelet thresholding attains these rates up to log factors and reveals that spatial inhomogeneity and function homogeneity jointly determine difficulty (Antoniadis et al., 2011).

Functional Linear Regression in RKHS

In functional linear regression with RKHS-regularized coefficient and general eigenvalue decay, the minimax prediction risk is

RnrnR_n^* \asymp r_n0

and, for polynomial decay RnrnR_n^* \asymp r_n1, RnrnR_n^* \asymp r_n2 (Lian, 2012).

4. Algorithmic Minimax Rates and Statistical-Computational Tradeoffs

For smooth minimax optimization RnrnR_n^* \asymp r_n3 with RnrnR_n^* \asymp r_n4-smooth, strongly convex-concave RnrnR_n^* \asymp r_n5, first-order methods can attain RnrnR_n^* \asymp r_n6 convergence in terms of the primal-dual gap—improved from RnrnR_n^* \asymp r_n7 by combining Mirror-Prox and Nesterov AGD (Thekumparampil et al., 2019). In nonconvex-concave settings, a proximal-point-based algorithm yields RnrnR_n^* \asymp r_n8 to first-order stationarity, sharpening the previous RnrnR_n^* \asymp r_n9 best-known rate.

5. Minimax Rates in Complex and Structured Models

Crowdsourced Binary Label Inference

In the Dawid–Skene one-coin model with rnr_n0 workers and rnr_n1 items, the minimax error in estimating the true label vector is exponentially small in the effective crowd ability: rnr_n2 where rnr_n3, rnr_n4 summarize collective ability (Gao et al., 2013). The projected EM algorithm attains these exponents, matching minimax lower bounds.

Reinforcement Learning/Off-Policy Evaluation

In OPE with function approximation, minimax optimal rates for marginal importance weight and Q-function estimation under completeness and realizability align with the critical local Rademacher complexity of the underlying function class: rnr_n5 where rnr_n6 scale as rnr_n7 for finite VC-dimension and as rnr_n8 for metric entropy exponent rnr_n9 (Uehara et al., 2021). Doubly robust OPE estimators achieve the semiparametric efficiency bound.

6. Extensions: Robustness, Privacy, and Time-Adaptive Frameworks

Adversarial and Privacy Effects

When regression is exposed to adversarial input perturbations of size nn0, the minimax sup-norm rate becomes the sum of the non-adversarial rate and the maximum variation induced by the perturbation: nn1 where nn2 (Peng et al., 2024). For nn3 classes, nn4.

Under sample-size uncertainty ("time-robust" minimaxity), the adversarial minimax risk is inflated by at most a logarithmic or iterated logarithmic factor in nn5 relative to the classical rate, e.g., nn6 for Gaussian mean estimation (Kirichenko et al., 2020).

Statistical Models with Partial Derivative Observations

Derivative observations in ANOVA/RKHS models can reduce the effective interaction order and accelerate minimax rates. For nn7 covariates with observed first partial derivatives in a nn8-way interaction model (order nn9), the minimax rate for function estimation matches that of a Θ\Theta0-way interaction model without derivatives: Θ\Theta1 (Dai et al., 2017). For Θ\Theta2 (all partials), the rate is parametric Θ\Theta3.

7. Proof Techniques and Key Lower Bound Constructions

The proofs of minimax lower rates in the cited literature predominantly employ:

  • Le Cam's two-point and Fano's multiple-hypothesis testing: Reduction from estimation to testing over well-separated parameter packings.
  • Information contraction (under privacy): KL divergence can be controlled under post-processing, significantly reducing effective sample size in private estimation (Duchi et al., 2013).
  • Metric entropy and covering arguments: Complexity of function/density class is measured under relevant loss/geometric properties (Hausdorff, Θ\Theta4, sup-norm, Wasserstein), directly determining optimal separation size and rates (Genovese et al., 2010, Dedecker et al., 2013, Zhao et al., 2023).
  • Specialized analysis for complex structures: E.g., dyadic dependence leads to pointwise minimax risk scaling in terms of Θ\Theta5 instead of Θ\Theta6 for Θ\Theta7 agents—owing to shared agent dependence (Graham et al., 2020).

8. Impact and Implications Across Statistical Domains

Minimax convergence rates serve as fundamental performance limits and guide both the development of estimation algorithms and theoretical hypotheses about statistical hardness under realistic constraints (privacy, adversarial robustness, function complexity, nonstandard designs, partial observations). They also catalyze algorithmic work at the interface with convex optimization, empirical risk minimization in modern machine learning, and information theory.

The rates tabulated above enable practitioners to directly compare the impact of model complexity, regularity, privacy, and robustness constraints, and to benchmark the performance of practical estimators and learning algorithms. In modern domains—privacy-preserving data analysis, crowdsourced inference, time-adaptive decision making, RL/OPE, high-dimensional graphical modeling—the precise identification of minimax exponents and constants continues to play a central role in statistical methodology and the theory of learning.


References

Definition Search Book Streamline Icon: https://streamlinehq.com
References (17)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Minimax Convergence Rates.