---
title: Catoni-Type Robust Estimators
url: https://www.emergentmind.com/topics/catoni-type-robust-estimators
type: topic
---

# Catoni-Type Robust Estimators

Catoni-type robust estimators are influence-function-based procedures for mean estimation, regression, and sequential inference that replace raw empirical averages or score equations by monotone soft truncations satisfying logarithmic envelope bounds. In the canonical finite-variance form, the estimator is defined as a root of
\[
\sum_{i=1}^n \varphi\!\big(\beta(X_i-\hat\theta)\big)=0,
\]
with \(\varphi\) chosen so that
\[
-\log\!\left(1-x+\frac{x^2}{2}\right)\le \varphi(x)\le \log\!\left(1+x+\frac{x^2}{2}\right).
\]
This construction yields high-probability control under weak moment assumptions and has since been extended to heavier-tailed regimes, time-uniform confidence sequences, high-dimensional mean estimation, regression, time series, and online decision problems [2202.01250][2010.05008][2311.13010][2602.12589].

## 1. Canonical form and influence-function architecture

The defining feature of a Catoni-type estimator is the replacement of the empirical score \(X_i-\theta\) by a transformed score \(\varphi(\beta(X_i-\theta))\), where \(\varphi\) is nondecreasing and obeys a logarithmic upper and lower envelope. In the mean-estimation setting, one works with
\[
r(\theta)=\frac{1}{\beta n}\sum_{i=1}^n \varphi\!\big(\beta(X_i-\theta)\big),
\]
so that \(r(\theta)\) is nonincreasing in \(\theta\), and \(\hat\theta\) is any root of \(r(\hat\theta)=0\). The proof strategy is typically based on exponential-moment inequalities for \(r(\theta)\), deterministic bracketing of the root, and high-probability control derived from the Catoni envelope [2010.05008].

In sequential settings, the same architecture is expressed through supermartingales. For predictable \(\lambda_t\), Catoni-style confidence sequences use products of the form
\[
M_t^{+} = \prod_{i=1}^t \exp\!\left(\phi(\lambda_i(X_i-\mu))-\frac{\lambda_i^2\sigma^2}{2}\right),\qquad
M_t^{-} = \prod_{i=1}^t \exp\!\left(-\phi(\lambda_i(X_i-\mu))-\frac{\lambda_i^2\sigma^2}{2}\right),
\]
or their heavy-tail analogues, and then invert Ville’s inequality to obtain time-uniform confidence sets for the mean [2202.01250][2409.04198].

The same formal pattern reappears in regression. Catoni-type robust regression replaces raw empirical loss minimization by truncated objectives such as
\[
\widehat R_{\varphi,\ell_1}(\theta)=\frac{1}{n\beta}\sum_{i=1}^n \varphi\!\big(\beta |y_i-\mathbf{x}_i^\top\theta|\big),
\]
or matrix-valued high-dimensional variants based on \(\psi_\alpha(\lambda |Y_i-\theta X_i|)\). A common interpretation is that the estimator preserves local linear behavior near zero while suppressing the contribution of extreme residuals in the tails [2010.05008][2411.05217].

## 2. Finite-variance theory and Gaussian-optimal behavior

In one dimension under finite variance, Catoni’s estimator is the canonical sub-Gaussian robust mean estimator. A local Newton-step-style variant analyzed for covariance-bounded data satisfies
\[
|\hat\mu-\mu| \le \left(1+O\!\left(\frac{\log(1/\delta)}{n}\right)\right)\sigma\sqrt{\frac{2\log(2/\delta)}{n}}
\]
with probability at least \(1-\delta\), provided \(n\gg \log(1/\delta)\). The leading constant is the Gaussian-optimal constant \(1\) in one dimension up to a vanishing correction term [2311.13010].

The sequential counterpart is the Catoni-style confidence sequence under the assumption
\[
\mathbb E[(X_t-\mu)^2\mid \mathcal F_{t-1}] \le \sigma^2.
\]
Its confidence set at time \(t\) is
\[
CI_t = \Bigl\{ m\in\mathbb R: -\frac{\sigma^2\sum_{i=1}^t\lambda_i^2}{2}-\log(2/\alpha) < \sum_{i=1}^t \phi(\lambda_i(X_i-m)) < \frac{\sigma^2\sum_{i=1}^t\lambda_i^2}{2}+\log(2/\alpha) \Bigr\}.
\]
This construction is anytime-valid, remains valid at arbitrary stopping times, and after stitching attains width of order
\[
\Theta\!\left(\sqrt{\frac{\log\log t}{t}}\right)
\]
up to constants and lower-order terms, matching the law-of-the-iterated-logarithm lower bound stated in that setting [2202.01250].

A recurrent issue in the finite-variance literature is scale calibration. Some Catoni procedures require a known variance proxy or related scale information, whereas later developments address unknown variance through self-normalization or joint estimation. A plausible implication is that Catoni-type methodology bifurcates into two design philosophies: externally tuned robustification versus internally estimated scale normalization [2311.13010][2511.11054][2602.12589].

## 3. Extensions to heavier tails and infinite-variance regimes

A direct heavy-tail generalization replaces the quadratic term \(x^2/2\) in the Catoni envelope by \(|x|^\alpha/\alpha\) for \(\alpha\in(1,2)\). The generalized influence function is any nondecreasing \(\varphi\) satisfying
\[
-\log\!\left(1-x+\frac{|x|^\alpha}{\alpha}\right)\le \varphi(x)\le \log\!\left(1+x+\frac{|x|^\alpha}{\alpha}\right),
\]
with the widest possible choice
\[
\varphi(x)=
\begin{cases}
\log\!\left(1+x+\frac{x^\alpha}{\alpha}\right), & x\ge 0,\\[4pt]
-\log\!\left(1-x+\frac{|x|^\alpha}{\alpha}\right), & x<0.
\end{cases}
\]
The resulting estimator assumes only \(v=\mathbb E|X-m|^\alpha<\infty\) and achieves deviation rate
\[
|\hat\theta-m| \asymp \left(\frac{\log(1/\epsilon)}{n}\right)^{(\alpha-1)/\alpha},
\]
up to constants depending on \(v\) and \(\alpha\). As \(\alpha\uparrow 2\), this recovers the usual Catoni rate \(O(\sqrt{\log(1/\epsilon)/n})\); as \(\alpha\downarrow 1\), the rate slows, reflecting the heavier tails. The same paper reports that the generalized Catoni estimator performs better than the empirical mean estimator and that the advantage becomes more pronounced as \(\alpha\) decreases [2010.05008].

A sequential infinite-variance analogue uses a \((1+\alpha)\)-moment bound with \(\alpha\in(0,1]\). In that parameterization, the influence function \(\psi_\alpha\) is required to satisfy
\[
-\log\big(1-x + C_\alpha |x|^{1+\alpha}\big) \le \psi_\alpha(x) \le \log\big(1+x + C_\alpha |x|^{1+\alpha}\big),
\]
where
\[
C_\alpha=\frac{12^\alpha}{1+\alpha}\left(\frac{1-\alpha}{\alpha}\right)^{1-\alpha}
\]
is identified as the tightest admissible coefficient. The associated confidence sequence is valid under only a known upper bound on the \((1+\alpha)\)-th central moment, and the stitched version has width
\[
\mathcal O\!\left(\left(\frac{\log \log t}{t}\right)^{\alpha/(1+\alpha)}\right)
\]
as \(t\) grows and
\[
\mathcal O\!\left((\log(1/\delta))^{\alpha/(1+\alpha)}\right)
\]
as \(\delta\downarrow 0\). The paper explicitly states that this improves the earlier Catoni-type confidence-sequence bounds and recovers the finite-variance case when \(\alpha=1\) [2409.04198].

These two heavy-tail extensions use different parameterizations of tail strength, but both preserve the same mechanism: logarithmic soft truncation replaces unavailable quadratic control. This suggests that the essential Catoni principle is not the quadratic form itself, but the envelope-based conversion of low-order moment assumptions into exponential-type deviation inequalities [2010.05008][2409.04198].

## 4. Multivariate mean estimation, regression, and dependent data

In higher dimensions, a standard Catoni-style strategy estimates every one-dimensional projection, intersects the resulting confidence slabs, and outputs the center of the minimum enclosing ball of the feasible set. By Jung’s theorem, this introduces the factor
\[
JUNG_d := \sqrt{\frac{2d}{d+1}}.
\]
For covariance-bounded heavy-tailed mean estimation with \(\Sigma \preccurlyeq \sigma^2 I_d\), the naive multidimensional Catoni lift therefore gives radius
\[
JUNG_d\,(1+o(1))\,\sigma\sqrt{\frac{2\log(2/\delta)}{n}}.
\]
A sharper analysis shows that this Jung-factor loss is not necessary in the uncontaminated heavy-tailed setting: for \(d\ge 2\) and \(n \ge C\log(1/\delta)\ge C^2 d\), there exists an estimator achieving
\[
\|\hat\mu-\mu\| \le (1-\tau)\,JUNG_d\,\sigma\sqrt{\frac{2\log(1/\delta)}{n}}
\]
for some universal constant \(\tau>0\). By contrast, in the adversarial contamination setting, the same paper proves that the Jung-factor loss is optimal in the infinite-sample limit [2311.13010].

Catoni-type robustification extends naturally to regression. For \(\ell_1\) regression under finite \(\alpha\)-moment assumptions with \(\alpha\in(1,2)\), the robust empirical objective
\[
\widehat R_{\varphi,\ell_1}(\theta) = \frac{1}{n\beta}\sum_{i=1}^n \varphi\!\big(\beta |y_i-\mathbf{x}_i^\top\theta|\big)
\]
replaces direct minimization of empirical absolute loss. Under total boundedness of \(\Theta\), finite \(\alpha\)-moments of \(\mathbf x\), and \(\sup_{\theta\in\Theta}\mathbb E|y-\mathbf{x}^\top\theta|^\alpha<\infty\), the resulting excess-risk guarantee scales as
\[
O\!\left(\left(\frac{d\log n}{n}\right)^{\frac{\alpha-1}{\alpha}}\right)
\]
under a radius bound on \(\Theta\). The stated benefit is that the procedure remains valid under infinite variance while retaining convergence rate \(n^{-(\alpha-1)/\alpha}\), which interpolates to the classical \(n^{-1/2}\) regime as \(\alpha\to 2\) [2010.05008].

For dependent heavy-tailed data, Catoni-type truncation has been integrated into high-dimensional least absolute deviation regression for exponentially \(\beta\)-mixing time series. The estimator minimizes
\[
\widehat\theta = \arg\min_{\theta\in\Theta} \left\{ \widehat R_{\psi_\alpha,\ell_1}(\theta)+\gamma\|\theta\|_{1,1} \right\},\qquad
\widehat R_{\psi_\alpha,\ell_1}(\theta) = \frac{1}{n\lambda}\sum_{i=1}^n \psi_\alpha\!\left(\lambda\,|Y_i-\theta X_i|\right),
\]
and achieves excess risk of order
\[
\mathcal{O}\!\left(\left(\frac{(d_1+d_2)(d_1\wedge d_2)\log^2 n}{n}\right)^{(\alpha-1)/\alpha}\right)
\]
up to mixing constants and covering factors. The same framework is applied to high-dimensional VAR regression, and the reported simulations indicate that the truncated estimator is essential because classical LAD can have risk with a tendency to blow up under heavy tails [2411.05217].

## 5. Sequential inference, bandits, and online adaptation

Catoni-type estimators have become central in sequential heavy-tail problems because the same influence-function machinery can be embedded in test supermartingales and confidence sequences. Under finite variance, Catoni-style confidence sequences are anytime-valid and robust to unbounded observations [2202.01250]. Under only a \((1+\alpha)\)-moment bound, improved \(\alpha\)-Catoni supermartingales produce tighter time-uniform confidence sets with better asymptotic width than earlier constructions [2409.04198].

In non-stationary heavy-tailed bandits, Catoni-style confidence sequences are used as change-point detectors. The Catoni-FCS-detector initializes repeated confidence sequences and declares a change when the intersection of active Catoni confidence sets becomes empty. Under the heavy-tail condition
\[
\mathbb{E}_{\nu}\!\left[|X-\mathbb{E}_{\nu}[X]|^{1+\epsilon}\right]\le v,\qquad \epsilon\in(0,1],
\]
the detector admits a finite-time guarantee with \(\gamma=2/T^3\) and predictable \(\{\lambda_i\}\):
\[
\mathbb{P}_{t_c}\!\left((\tau-t_c)^+ \le \mathcal{O}\!\left( v^{1/\epsilon}\, \frac{\log(T)}{\delta^{(1+\epsilon)/\epsilon}} \right)\right) \ge 1-\frac{14}{T},\qquad
\mathbb{P}_{t_c}(\tau<t_c)\le \frac{14}{T}.
\]
Embedded inside Robust-CPD-UCB, this leads to a near-optimal regret bound
\[
\widetilde{\mathcal{O}\!\left( (K\Upsilon)^{\epsilon/(1+\epsilon)}(vT)^{1/(1+\epsilon)} \right)},
\]
matching the paper’s lower bound up to logarithmic factors and constants [2505.20051].

In contextual bandits with general function approximation, Catoni’s estimator is used in a different role: it robustifies the excess-loss cross term inside a variance-weighted optimistic regression objective rather than directly estimating the reward mean. In the known-variance case, the regret bound depends on the cumulative reward variance and only logarithmically on the reward range \(R\); in the unknown-variance case, a peeling-based algorithm uses Catoni both for excess-loss estimation and for aggregate variance estimation. The leading-order variance dependence is shown to be minimax optimal through a matching lower bound of order
\[
\Omega\!\left(\sqrt{\mathbb E\sum_{t=1}^T \sigma_t^2}\right).
\]
This suggests that Catoni-type robustification is compatible not only with heavy-tail protection, but also with variance-sensitive online learning [2502.02486].

## 6. Relations to adjacent methods, unknown variance, and asymptotic limit theory

Catoni-type estimators are often discussed alongside median-of-means and trimming-based methods, but the equivalence is only partial. One robust empirical mean construction based on the median of block averages explicitly contrasts itself with Catoni’s smooth truncation: it targets the same sub-Gaussian deviation goal under finite variance, but does not require prior knowledge of a variance proxy and instead depends on the number of blocks \(V\) chosen from the confidence level [1112.3914]. Another framework interpolates between the two paradigms: when \(n=1\) and \(\Delta\propto \sigma(F)\sqrt{N}\), it recovers Catoni’s estimator, while large block size and \(\Delta\propto \sigma(F)\) lead to a MOM-type estimator. That interpolation also yields uniform deviation bounds over function classes and optimal contamination dependence in multivariate mean estimation [1812.03523].

The same distinction appears in covariance estimation. A trimmed quadratic-form estimator that removes the largest \(\langle Y_i,v\rangle^2\) values in each direction is explicitly described as not Catoni’s influence-function estimator in form, but as conceptually analogous: Catoni uses soft truncation through a bounded influence function, whereas trimming discards the largest values directly. Under a bounded \(L^p\)-\(L^2\) marginal condition for \(p\ge 4\), this trimming-based covariance estimator attains Gaussian-like operator-norm rates and optimal robustness to adversarial contamination, but it remains Catoni-adjacent rather than literally Catoni-type [2209.13485].

Unknown variance has motivated two further lines of development. One is self-normalization: a Catoni estimator with empirical standard deviation \(\hat\sigma\) in the denominator admits Berry–Esseen and moderate-deviation bounds under heavy tails, including a self-normalized Berry–Esseen bound and Cramér-type moderate deviations [2602.12589]. The other is joint estimation of location or regression parameters together with the error variance by solving two coupled Catoni-type equations. That joint framework is explicitly not the gradient of any scalar loss in general, is described as tuning-free, and yields \((1-\epsilon)\)-confidence interval length of order
\[
O\!\left(\sigma \sqrt{\frac{\log(\epsilon^{-1})}{n}}\right)
\]
for mean estimation while also covering linear and \(\ell_2\)-penalized regression [2511.11054].

Recent asymptotic theory clarifies that Catoni-type estimators are not only nonasymptotically robust. For mean estimation with tuning \(\alpha_n=O(n^{-1/2})\), Berry–Esseen bounds and moderate deviation principles have been established, with centering at an implicit target \(u_n\) satisfying
\[
\mathbb E[\varphi(\alpha_n(X_1-u_n))]=0,
\]
and with regression analogues giving a multivariate Berry–Esseen bound for the Catoni-type score equation
\[
\sum_{i=1}^n x_i\,\varphi\!\left(\alpha_n(y_i-x_i^\top\beta)\right)=0.
\]
A common misconception is that Catoni-type procedures are only finite-sample concentration devices; the available results indicate that they also support quantitative central limit theory and moderate deviation analysis under weak moment assumptions [2602.12589].

Source: https://www.emergentmind.com/topics/catoni-type-robust-estimators