---
title: 'FlyPrompt: Catoni-Type Robust Estimators'
url: https://www.emergentmind.com/topics/flyprompt
type: topic
---

# FlyPrompt: Catoni-Type Robust Estimators

Catoni-type robust estimators are influence-function-based procedures for estimating means, risks, and regression parameters under heavy-tailed sampling, typically by replacing raw residuals with a monotone score function whose growth is controlled by logarithmic envelope inequalities. In their canonical form, they are defined as roots of estimating equations rather than as empirical averages, and they are designed to retain high-probability behavior close to Gaussian or sub-Gaussian benchmarks under substantially weaker moment assumptions than those required by classical estimators [2202.01250]. Subsequent work has extended the Catoni paradigm from finite-variance mean estimation to finite \(\alpha\)-moment settings with \(\alpha\in(1,2)\), to sequential confidence sequences, to high-dimensional regression and time series, to contextual and non-stationary bandits, and to asymptotic normal approximation via Berry–Esseen bounds and moderate deviations [2010.05008], [2411.05217], [2502.02486], [2505.20051], [2602.12589].

## 1. Canonical form and influence-function construction

Catoni’s original mean estimator targets \(m=\mathbb E[X]\) through the root of
\[
\sum_{i=1}^n \varphi\!\big(\beta(X_i-\hat\theta)\big)=0,
\]
where \(\beta>0\) is a tuning parameter and \(\varphi\) is nondecreasing and satisfies the envelope
\[
-\log\!\left(1-x+\frac{x^2}{2}\right)\le \varphi(x)\le \log\!\left(1+x+\frac{x^2}{2}\right).
\]
A standard explicit choice is
\[
\varphi(x)=
\begin{cases}
\log(1+x+x^2/2), & x\ge 0,\\
-\log(1-x+x^2/2), & x<0.
\end{cases}
\]
This construction is the basic one-dimensional Catoni estimator under finite variance and is the benchmark against which many later developments are formulated [2202.01250].

The defining feature of the Catoni envelope is that it preserves approximate linearity near zero while suppressing the contribution of large observations. In the mean-estimation literature, this yields high-probability concentration with only a variance bound rather than boundedness or sub-Gaussian mgf assumptions. In sequential form, the same envelope generates nonnegative supermartingales of the form
\[
M_t^{+} = \prod_{i=1}^t \exp\!\left(\phi(\lambda_i(X_i-\mu))-\frac{\lambda_i^2\sigma^2}{2}\right), \qquad
M_t^{-} = \prod_{i=1}^t \exp\!\left(-\phi(\lambda_i(X_i-\mu))-\frac{\lambda_i^2\sigma^2}{2}\right),
\]
which are then converted into time-uniform confidence sequences by Ville’s inequality [2202.01250].

A later asymptotic treatment formalizes the same structure through estimating equations
\[
\sum_{i=1}^n \varphi\bigl(\alpha_n(X_i-\theta)\bigr)=0
\]
for mean estimation and
\[
\sum_{i=1}^n x_i\,\varphi\!\left(\alpha_n(y_i-x_i^\top\beta)\right)=0
\]
for regression, with \(\varphi\) continuous, monotone, and satisfying the same Catoni envelope. That formulation makes explicit that Catoni-type estimators are robust \(M\)-estimators whose score is a smooth truncation of the residual rather than a hard cutoff [2602.12589].

## 2. Heavy tails and finite-moment generalizations

A central development is the extension of Catoni’s finite-variance estimator to settings where only a finite \(\alpha\)-th central moment exists, with \(\alpha\in(1,2)\). In that setting, the quadratic term \(x^2/2\) in the envelope is replaced by \(|x|^\alpha/\alpha\), giving the generalized condition
\[
-\log\!\left(1-x+\frac{|x|^\alpha}{\alpha}\right)\le \varphi(x)\le \log\!\left(1+x+\frac{|x|^\alpha}{\alpha}\right),
\]
with the widest possible choice
\[
\varphi(x)=
\begin{cases}
\log\!\left(1+x+\frac{x^\alpha}{\alpha}\right), & x\ge 0,\\[4pt]
-\log\!\left(1-x+\frac{|x|^\alpha}{\alpha}\right), & x<0.
\end{cases}
\]
The estimator remains the root of the Catoni score equation, but the proof substitutes \(\alpha\)-moment control for variance control [2010.05008].

For i.i.d. data with
\[
m=\mathbb E[X_1], \qquad v=\mathbb E|X_1-m|^\alpha<\infty,
\]
one defines
\[
r(\theta)=\frac{1}{\beta n}\sum_{i=1}^n \varphi\!\big(\beta(X_i-\theta)\big).
\]
Since \(\varphi\) is nondecreasing, \(r(\theta)\) is nonincreasing in \(\theta\), and the estimator \(\hat\theta\) is any solution of \(r(\hat\theta)=0\). The key exponential-moment lemma gives
\[
\mathbb{E}\!\left[\exp\!\big(\beta n r(\theta)\big)\right]
\le
\exp\!\left( n\beta(m-\theta) + 2^{\alpha-1}n\beta^\alpha \frac{v+|m-\theta|^\alpha}{\alpha} \right),
\]
together with the corresponding lower-tail bound. The proof uses the pointwise truncation inequalities, independence, \(1+x\le e^x\), and
\[
|a+b|^\alpha \le 2^{\alpha-1}\big(|a|^\alpha+|b|^\alpha\big),\qquad \alpha\in(1,2)
\]
[2010.05008].

Under a suitable choice of \(\beta\), the resulting deviation guarantee is
\[
|\hat\theta-m| = O\!\left( \big(\log(\epsilon^{-1})\big)^{\frac{\alpha-1}{\alpha}} n^{-\frac{\alpha-1}{\alpha}} \right)
\quad\text{with probability at least }1-2\epsilon.
\]
Equivalently, up to constants depending on \(v\) and \(\alpha\),
\[
|\hat\theta-m| \asymp \left(\frac{\log(1/\epsilon)}{n}\right)^{(\alpha-1)/\alpha}.
\]
As \(\alpha\uparrow 2\), the exponent \((\alpha-1)/\alpha\) tends to \(1/2\), recovering the classical Catoni rate \(O(\sqrt{\log(1/\epsilon)/n})\); as \(\alpha\downarrow 1\), the rate deteriorates, reflecting heavier tails [2010.05008].

The sequential confidence-sequence literature develops an analogous extension for infinite-variance settings by assuming a known \((1+\alpha)\)-th central moment bound with \(\alpha\in(0,1]\). The influence function \(\psi_\alpha\) then satisfies
\[
-\log\big(1-x + C_\alpha |x|^{1+\alpha}\big) \le \psi_\alpha(x) \le \log\big(1+x + C_\alpha |x|^{1+\alpha}\big),
\]
where
\[
C_\alpha=\frac{12^\alpha}{1+\alpha}\left(\frac{1-\alpha}{\alpha}\right)^{1-\alpha},
\]
and this coefficient is stated to be the tightest one; it recovers Catoni’s constant \(1/2\) when \(\alpha=1\) [2409.04198].

## 3. Deviation theory, sharp constants, and asymptotics

In one dimension, Catoni’s estimator is the canonical sub-Gaussian mean estimator under finite variance. A Newton-step-style variant
\[
\hat\mu \;=\; \mu_0 + \frac{T}{n}\sum_{i=1}^n \psi\!\left(\frac{x_i-\mu_0}{T}\right), \qquad
T=\sigma\sqrt{\frac{n}{2\log(2/\delta)}}
\]
satisfies
\[
|\hat\mu-\mu| \le \left(1+O\!\left(\frac{\log(1/\delta)}{n}\right)\right)\sigma\sqrt{\frac{2\log(2/\delta)}{n}}
\]
with probability at least \(1-\delta\), which gives the Gaussian-optimal constant \(1\) in one dimension up to a vanishing correction [2311.13010].

A central issue in higher dimensions is whether one can estimate each projection \(\langle u,\mu\rangle\) by one-dimensional Catoni procedures, intersect the resulting confidence slabs, and output the center of the minimum enclosing ball without losing a geometric constant. The standard approach incurs the Jung factor
\[
JUNG_d := \sqrt{\frac{2d}{d+1}},
\]
yielding the natural Catoni-lifted radius
\[
\|\hat\mu-\mu\| \le JUNG_d\cdot (1+o(1))\, \sigma\sqrt{\frac{2\log(2/\delta)}{n}}.
\]
A sharper analysis shows that in the heavy-tailed but uncontaminated covariance-bounded setting this loss is not necessary: for \(d\ge 2\), if \(n \ge C\log(1/\delta)\ge C^2 d\), there is an estimator such that
\[
\|\hat\mu-\mu\| \le (1-\tau)\,JUNG_d \,\sigma\sqrt{\frac{2\log(1/\delta)}{n}}
\]
for some universal constant \(\tau>0\) [2311.13010].

That improvement is achieved by regime adaptivity. The estimator distinguishes “inlier-light” from “outlier-light” regimes. In the former, it uses a sharpened Catoni step with envelope
\[
-\log\!\left(1-x+(1-\eta)\frac{x^2}{2}\right) \le \psi(x)\le \log\!\left(1+x+(1-\eta)\frac{x^2}{2}\right),
\]
which improves the constant in the high-probability term. In the latter, it uses a trimmed mean after discarding samples whose coordinates exceed \(\sqrt{\beta}T\), thereby approaching the Gaussian constant \(1\) while avoiding the Jung-factor loss on that branch [2311.13010]. The same paper establishes a sharp distinction between heavy-tailed estimation and adversarial contamination: in the contamination model, the Jung-factor loss becomes optimal in the infinite-sample limit.

Beyond nonasymptotic deviation inequalities, asymptotic distribution theory has also been developed. For i.i.d. mean estimation with tuning \(\alpha_n=a_n/\sigma\), an implicit centering \(u_n\) is defined by
\[
\mathbb E\bigl[\varphi(\alpha_n(X_1-u_n))\bigr]=0.
\]
The estimator \(\hat\theta\) solving
\[
\sum_{i=1}^n \varphi\bigl(\alpha_n(X_i-\theta)\bigr)=0
\]
then satisfies the Berry–Esseen bound
\[
\sup_{z\in\mathbb R}\left| P\!\left(\frac{\sqrt n(\hat\theta-u_n)}{\sigma}\le z\right)-\Phi(z) \right| \le C\beta_2 + C\beta_3,
\]
where \(\beta_2\) and \(\beta_3\) are the truncated second- and third-moment terms defined in the theorem. Under \(E|X_1-u|^{2+\delta}<\infty\), this becomes an \(n^{-\delta/2}\) bound. The same work proves moderate deviations for the known-variance estimator and a self-normalized version using \(\hat\sigma\) when the variance is unknown [2602.12589].

## 4. Sequential and anytime-valid Catoni procedures

Catoni-type methods have a natural sequential incarnation because the logarithmic envelope yields test supermartingales. Under the martingale-type assumptions
\[
\mathbb E[X_t\mid \mathcal F_{t-1}] = \mu,\qquad
\mathbb E[(X_t-\mu)^2\mid \mathcal F_{t-1}] \le \sigma^2,
\]
the Catoni-style confidence sequence is
\[
CI_t = \Bigl\{ m\in\mathbb R: -\frac{\sigma^2\sum_{i=1}^t\lambda_i^2}{2}-\log(2/\alpha) < \sum_{i=1}^t \phi(\lambda_i(X_i-m)) < \frac{\sigma^2\sum_{i=1}^t\lambda_i^2}{2}+\log(2/\alpha) \Bigr\},
\]
where \((\lambda_t)\) is predictable. The sequence satisfies
\[
\Pr\bigl[\forall t\ge 1,\ \mu\in CI_t\bigr]\ge 1-\alpha,
\]
so it is valid at arbitrary stopping times [2202.01250].

The same framework extends to finite \(p\)-th moments with \(1<p<2\). If
\[
\mathbb E[|X_t-\mu|^p\mid \mathcal F_{t-1}] \le v_t,
\]
one uses a \(p\)-Catoni influence function \(\phi_p\) satisfying
\[
-\log\!\left(1-x+\frac{|x|^p}{p}\right)\le \phi_p(x)\le \log\!\left(1+x+\frac{|x|^p}{p}\right),
\]
and obtains a confidence sequence
\[
CI_t^p= \left\{ m\in\mathbb R: -\frac{\sum_{i=1}^t v_i\lambda_i^p}{p}-\log(2/\alpha) < \sum_{i=1}^t \phi_p(\lambda_i(X_i-m)) < \frac{\sum_{i=1}^t v_i\lambda_i^p}{p}+\log(2/\alpha) \right\}.
\]
For i.i.d. \(p\)-moment data with \(\lambda_t\asymp t^{-1/p}\), the width rate is
\[
|CI_t^p|=\tilde{\mathcal O}\bigl(t^{-(p-1)/p}\bigr),
\]
matching known minimax lower bounds up to logarithmic factors [2202.01250].

A stitched Catoni-style construction further achieves the law-of-the-iterated-logarithm rate. With epochs \(t_j=2^j\) and error budgets \(\alpha_j=\alpha/(j+2)^2\), the resulting width satisfies
\[
|CI_t| \lesssim \sqrt{\frac{\log(2/\alpha)+2\log\log^2 t}{t}},
\]
which corresponds to the optimal \(\Theta(\sqrt{\log\log t/t})\) scaling up to constants and lower-order logarithmic terms [2202.01250].

For infinite-variance models with known \((1+\alpha)\)-moment bound \(\nu_\alpha\), improved Catoni-type confidence sequences are built from the supermartingales
\[
M_t^+ = \prod_{i=1}^t \exp\Big\{ \psi_\alpha\!\big(\theta_i(X_i-\mu)\big) -\theta_i^{1+\alpha} C_\alpha \nu_\alpha \Big\},
\qquad
M_t^- = \prod_{i=1}^t \exp\Big\{ -\psi_\alpha\!\big(\theta_i(X_i-\mu)\big) -\theta_i^{1+\alpha} C_\alpha \nu_\alpha \Big\},
\]
leading to
\[
\mathrm{CI}_t = \left\{ m\in\mathbb R: \sum_{i=1}^t \psi_\alpha\!\big(\theta_i(X_i-m)\big) \le C_\alpha \nu_\alpha \sum_{i=1}^t \theta_i^{1+\alpha} + \log\frac{2}{\delta} \right\}.
\]
After stitching, the width obeys
\[
|\mathrm{CI}_t| = \mathcal O\!\left(\left(\frac{\log\log t}{t}\right)^{\alpha/(1+\alpha)}\right),
\qquad
|\mathrm{CI}_t| = \mathcal O\!\left((\log(1/\delta))^{\alpha/(1+\alpha)}\right),
\]
improving earlier Catoni-type sequential bounds for \(\alpha<1\) and recovering the optimized finite-variance case when \(\alpha=1\) [2409.04198].

## 5. Regression, time series, and high-dimensional structured estimation

Catoni-type robustification extends naturally from scalar mean estimation to empirical-risk minimization. For \(\ell_1\) regression with parameter space \(\Theta\), the population target is
\[
\theta^* \in \arg\min_{\theta\in\Theta} R_{\ell_1}(\theta), \qquad
R_{\ell_1}(\theta)=\mathbb{E}_{(\mathbf{x},y)\sim\Pi}\big[\,| \mathbf{x}^\top\theta - y|\,\big],
\]
and the robust empirical objective is
\[
\widehat R_{\varphi,\ell_1}(\theta) = \frac{1}{n\beta}\sum_{i=1}^n \varphi\!\big(\beta |y_i-\mathbf{x}_i^\top\theta|\big),
\]
where \(\varphi\) is the \(\alpha\)-generalized influence function. Under total boundedness of \(\Theta\), \(\mathbb E|\mathbf{x}|^\alpha<\infty\), and \(\sup_{\theta\in\Theta}\mathbb E|y-\mathbf{x}^\top\theta|^\alpha<\infty\), the estimator minimizing \(\widehat R_{\varphi,\ell_1}\) satisfies an excess-risk bound of order
\[
R_{\ell_1}(\hat\theta)-R_{\ell_1}(\theta^*) =O\!\left(\left(\frac{d\log n}{n}\right)^{\frac{\alpha-1}{\alpha}}\right)
\]
under an additional radius bound on \(\Theta\subseteq\mathbb R^d\) [2010.05008].

For high-dimensional dependent data, the same strategy appears in a Catoni type truncated minimization framework for LAD regression. With stationary exponentially \(\beta\)-mixing time series \(\{(X_i,Y_i)\}\), matrix parameter \(\theta\in\mathbb R^{d_2\times d_1}\), and penalized objective
\[
\widehat\theta = \arg\min_{\theta\in\Theta} \left\{ \widehat R_{\psi_\alpha,\ell_1}(\theta)+\gamma\|\theta\|_{1,1} \right\},
\qquad
\widehat R_{\psi_\alpha,\ell_1}(\theta) = \frac{1}{n\lambda}\sum_{i=1}^n \psi_\alpha\!\left(\lambda\,|Y_i-\theta X_i|\right),
\]
one obtains excess risk of order
\[
\mathcal{O}\!\left(\left(\frac{(d_1+d_2)(d_1\wedge d_2)\log^2 n}{n}\right)^{(\alpha - 1)/\alpha}\right)
\]
under finite \(\alpha\)-moments, total boundedness, low-rank boundedness, and exponential \(\beta\)-mixing [2411.05217]. The same paper applies the method to high-dimensional VAR regression and reports that classical LAD has a tendency to blow up under heavy-tailed innovations, whereas the truncated estimator is stabilized by the Catoni transform.

A distinct line of work studies joint robust estimation when both the trend parameter and the error variance are unknown. In that framework, mean estimation is based on the coupled equations
\[
\begin{cases}
f_{1}(\theta,v)=\frac{1}{n}\sum_{i=1}^{n} \psi_{1}\!\left(\alpha_{1} \frac{X_{i}-\theta}{v}  \right)=0,\\[4pt]
f_{2}(\theta,v)=\frac{1}{n}\sum_{i=1}^{n} \psi_{2}\!\left(\alpha_{2}\left( \frac{(X_{i}-\theta)^{2}}{v^{2}}-1 \right)\right)=0,
\end{cases}
\]
and regression is handled by analogous coupled equations in \((\theta,v)\). The key conceptual point is that these equations generally cannot be written as the gradient of a single scalar function; existence is proved by a Poincaré–Miranda argument on suitable geometric regions [2511.11054].

## 6. Bandits, robustness comparisons, and conceptual boundaries

Catoni-type estimators have become algorithmic primitives in online learning. In contextual bandits with general function approximation, Catoni’s estimator is used not on rewards directly but on the excess-loss cross term inside a variance-weighted regression objective:
\[
f_t = \arg\min_{f\in F}\max_{f'\in F} L_t(f,f'),
\]
with
\[
L_t(f,f') := \sum_{i\in[t]}\frac{1}{\sigma_i^2}\bigl(f'(x_i)-f(x_i)\bigr)^2 + 2t\, Catoni_{\theta_t(f,f')}\bigl(\{Z_i(f,f')\}_{i\in[t]}\bigr),
\]
and
\[
Z_i(f,f'):= \frac{1}{\sigma_i^2}\bigl(f(x_i)-f'(x_i)\bigr)\bigl(f'(x_i)-y_i\bigr).
\]
In the known-variance case, the resulting regret bound depends on the cumulative reward variance and only logarithmically on the reward range \(R\), and the leading variance term is accompanied by a matching lower bound \(\Omega\!\left(\sqrt{\mathbb E\sum_{t=1}^T \sigma_t^2}\right)\) [2502.02486].

In heavy-tailed piecewise-stationary bandits, Catoni-style confidence sequences are used for change-point detection. The Catoni confidence sequence is
\[
CI_t^\phi = \left\{ m\in\mathbb{R}:\; \sum_{i=1}^t \phi_\epsilon(\lambda_i(X_i-m)) \in \left[\mp \frac{v}{2}\sum_{i=1}^t \lambda_i^{1+\epsilon}\pm \log\!\left(\frac{2}{\gamma}\right)\right] \right\},
\]
and repeated initialization of such sequences yields a detector whose stopping time is the first time the intersection of all active sequences becomes empty. This is used inside Robust-CPD-UCB for regret minimization under heavy tails and unknown changes in arm means [2505.20051].

The methodological boundary between Catoni-type estimators and other robust procedures is important. Median-of-means and block-median estimators aim at the same objective—sub-Gaussian deviation under finite variance—but they robustify by sample splitting and aggregation rather than by smooth influence functions. A representative construction splits data into blocks \(B_1,\dots,B_V\) and defines
\[
P_B f := \med\bigl(P_{B_K}f,\;K=1,\dots,V\bigr),
\]
achieving deviation control without boundedness of \(f\) and without prior knowledge of a variance proxy, at the cost of fixing the confidence level in advance [1112.3914]. Another line constructs estimators that interpolate between Catoni and median-of-means: with block size \(n=1\) and \(\Delta\propto \sigma(F)\sqrt{N}\) one recovers Catoni’s estimator, while larger blocks lead to MOM-type behavior [1812.03523].

A related misconception is that every robust heavy-tail estimator is Catoni-type. That is not the case. Some covariance estimators achieve Catoni-like goals—Gaussian/sub-Gaussian operator-norm rates under heavy tails and robustness to contamination—through hard trimming rather than influence-function truncation. The trimmed covariance estimator based on
\[
\widehat{\mu}_k(v) := \inf_{S\subset [n],\,|S|=n-k}\frac{1}{n-k}\sum_{i\in S}\langle Y_i,v\rangle^2
\]
is explicitly described as conceptually analogous to Catoni-type methods but not literally a Catoni estimator in form [2209.13485].

Taken together, these developments show that “Catoni-type” denotes a precise robust-estimation architecture: monotone influence transformation, logarithmic envelope control, and estimation by score balancing or test-supermartingale inversion. What varies across the literature is the moment regime, the geometry, and the inferential target. In one dimension, Catoni’s estimator remains the baseline for sharp constants; in higher dimensions, geometry can sometimes be improved beyond naive projection-lifting; in contamination models, the same geometry can become an information-theoretic barrier; and in sequential, dependent, or online settings, the Catoni envelope serves as a reusable robust primitive rather than merely a mean estimator [2311.13010].

Source: https://www.emergentmind.com/topics/flyprompt