Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sigmoid-FTRL: Adaptive ATE Estimation

Updated 26 November 2025
  • Sigmoid-FTRL is an adaptive online design strategy that minimizes Neyman regret for average treatment effect estimation using AIPW estimators.
  • It decomposes a nonconvex variance minimization problem into two convex learning tasks solved via FTRL updates over treatment probabilities and linear predictors.
  • The method achieves asymptotic optimality and enables valid inference through consistently conservative variance estimation and adaptive ridge regression.

Sigmoid-FTRL is an adaptive online experimental design strategy for minimizing variance (Neyman regret) in the estimation of average treatment effects using Augmented Inverse Probability Weighting (AIPW) estimators, explicitly within the design-based potential outcomes framework where both outcomes and covariates are deterministic. The method unifies online convex optimization and adaptive Neyman allocation via a decomposition of a nonconvex variance-minimization problem into two convex online learning problems, efficiently addressed through Follow-the-Regularized-Leader (FTRL) updates over both treatment probabilities and linear predictors. Sigmoid-FTRL establishes asymptotic optimality, supports consistently conservative variance estimation, and enables construction of valid confidence intervals under broad regularity conditions (Chen et al., 25 Nov 2025).

1. Design-Based Setting and Problem Formulation

Consider TT observed units indexed by t=1,…,Tt = 1, \ldots, T, each with a covariate vector xt∈Rdx_t \in \mathbb{R}^d bounded in norm (∥xt∥≤R\|x_t\| \leq R) and deterministic potential outcomes yt(1),yt(0)∈Ry_t(1), y_t(0) \in \mathbb{R}. The goal is estimation of the average treatment effect (ATE),

τ:=1T∑t=1T[yt(1)−yt(0)]\tau := \frac{1}{T}\sum_{t=1}^T[y_t(1) - y_t(0)]

using only randomized assignment. At each round tt, the procedure selects:

  • Assignment probability pt∈(0,1)p_t \in (0,1) as a function of history Ft−1\mathcal{F}_{t-1}
  • Linear predictor coefficients βt(1),βt(0)∈Rd\beta_t(1), \beta_t(0) \in \mathbb{R}^d

Treatment t=1,…,Tt = 1, \ldots, T0 is randomized, generating observed outcome t=1,…,Tt = 1, \ldots, T1. For each arm t=1,…,Tt = 1, \ldots, T2, online ridge-regression is used to fit

t=1,…,Tt = 1, \ldots, T3

with t=1,…,Tt = 1, \ldots, T4 an adaptive regularization parameter.

The adaptive AIPW estimator is

t=1,…,Tt = 1, \ldots, T5

which is unbiased, and whose variance (and thus regret relative to the oracle design) admits closed-form analysis (Chen et al., 25 Nov 2025).

2. Neyman Regret and Oracle Design Benchmark

The “oracle” nonadaptive design fixes both linear predictors t=1,…,Tt = 1, \ldots, T6 (by armwise OLS on all t=1,…,Tt = 1, \ldots, T7 units) and probability t=1,…,Tt = 1, \ldots, T8, minimizing the expected variance:

t=1,…,Tt = 1, \ldots, T9

where xt∈Rdx_t \in \mathbb{R}^d0 is the residual variance for potential outcomes under arm xt∈Rdx_t \in \mathbb{R}^d1.

The benchmark variance is

xt∈Rdx_t \in \mathbb{R}^d2

The Neyman regret of any adaptive policy xt∈Rdx_t \in \mathbb{R}^d3 is

xt∈Rdx_t \in \mathbb{R}^d4

highlighting the additional variance incurred by adaptation relative to the nonadaptive oracle (Chen et al., 25 Nov 2025).

3. Algorithmic Formulation: Decomposition and Convexification

Direct minimization of Neyman regret is nonconvex in the triple xt∈Rdx_t \in \mathbb{R}^d5. Sigmoid-FTRL circumvents this by decomposing the regret into two convex sequences:

  • Probability Regret: For fixed predictors,

xt∈Rdx_t \in \mathbb{R}^d6

and

xt∈Rdx_t \in \mathbb{R}^d7

with xt∈Rdx_t \in \mathbb{R}^d8 convex on xt∈Rdx_t \in \mathbb{R}^d9.

  • Prediction Regret: For fixed ∥xt∥≤R\|x_t\| \leq R0,

∥xt∥≤R\|x_t\| \leq R1

and

∥xt∥≤R\|x_t\| \leq R2

which is jointly convex in ∥xt∥≤R\|x_t\| \leq R3.

Lemma 3.3 asserts that

∥xt∥≤R\|x_t\| \leq R4

enabling separate convex-optimizable updates (Chen et al., 25 Nov 2025).

4. Sigmoid-FTRL Mechanism

The algorithm maintains parameter ∥xt∥≤R\|x_t\| \leq R5 such that ∥xt∥≤R\|x_t\| \leq R6 for a differentiable sigmoid ∥xt∥≤R\|x_t\| \leq R7 with properties: monotonicity, ∥xt∥≤R\|x_t\| \leq R8, as well as specific convexity and derivative decay conditions. Examples include ∥xt∥≤R\|x_t\| \leq R9 or yt(1),yt(0)∈Ry_t(1), y_t(0) \in \mathbb{R}0.

For probability updates:

  • Define yt(1),yt(0)∈Ry_t(1), y_t(0) \in \mathbb{R}1.
  • Use FTRL with regularizer yt(1),yt(0)∈Ry_t(1), y_t(0) \in \mathbb{R}2:

yt(1),yt(0)∈Ry_t(1), y_t(0) \in \mathbb{R}3

where yt(1),yt(0)∈Ry_t(1), y_t(0) \in \mathbb{R}4 is an IPW-estimator of yt(1),yt(0)∈Ry_t(1), y_t(0) \in \mathbb{R}5.

For linear predictor updates:

  • For each arm, solve

yt(1),yt(0)∈Ry_t(1), y_t(0) \in \mathbb{R}6

Regularization is adaptive: yt(1),yt(0)∈Ry_t(1), y_t(0) \in \mathbb{R}7, with yt(1),yt(0)∈Ry_t(1), y_t(0) \in \mathbb{R}8.

Sequential steps (summarized):

Step Description Complexity
Prediction update Ridge regression by arm yt(1),yt(0)∈Ry_t(1), y_t(0) \in \mathbb{R}9 per step
Probability update 1D convex minimization in τ:=1T∑t=1T[yt(1)−yt(0)]\tau := \frac{1}{T}\sum_{t=1}^T[y_t(1) - y_t(0)]0 τ:=1T∑t=1T[yt(1)−yt(0)]\tau := \frac{1}{T}\sum_{t=1}^T[y_t(1) - y_t(0)]1
Residuals Estimate armwise via IPW sums τ:=1T∑t=1T[yt(1)−yt(0)]\tau := \frac{1}{T}\sum_{t=1}^T[y_t(1) - y_t(0)]2 (overall)

No hyperparameter tuning is required; all regularization is data-adaptive (Chen et al., 25 Nov 2025).

5. Theoretical Guarantees

Convergence of Neyman Regret

Sigmoid-FTRL achieves

τ:=1T∑t=1T[yt(1)−yt(0)]\tau := \frac{1}{T}\sum_{t=1}^T[y_t(1) - y_t(0)]3

assuming bounded moments and well-conditioned Gram matrices after initial τ:=1T∑t=1T[yt(1)−yt(0)]\tau := \frac{1}{T}\sum_{t=1}^T[y_t(1) - y_t(0)]4 samples, for any τ:=1T∑t=1T[yt(1)−yt(0)]\tau := \frac{1}{T}\sum_{t=1}^T[y_t(1) - y_t(0)]5 as above. This matches the lower bound:

τ:=1T∑t=1T[yt(1)−yt(0)]\tau := \frac{1}{T}\sum_{t=1}^T[y_t(1) - y_t(0)]6

where no algorithm can improve on the τ:=1T∑t=1T[yt(1)−yt(0)]\tau := \frac{1}{T}\sum_{t=1}^T[y_t(1) - y_t(0)]7 rate under analogous regularity assumptions (demonstrated via a noisy two-armed construction) (Chen et al., 25 Nov 2025).

Distributional Asymptotics and Inference

Under non-superefficiency (τ:=1T∑t=1T[yt(1)−yt(0)]\tau := \frac{1}{T}\sum_{t=1}^T[y_t(1) - y_t(0)]8):

τ:=1T∑t=1T[yt(1)−yt(0)]\tau := \frac{1}{T}\sum_{t=1}^T[y_t(1) - y_t(0)]9

facilitating Wald-type inference.

Consistent Conservative Variance Estimator

A variance bound estimator,

tt0

with tt1 defined by armwise IPW residuals, is consistent:

tt2

Both the variance estimator and the estimator tt3 enable construction of asymptotically accurate tt4 Wald-type confidence intervals,

tt5

with coverage tending to tt6 as tt7 (Chen et al., 25 Nov 2025).

6. Implementation Guidelines

  • Sigmoid choice: Recommended tt8 includes either tt9 or pt∈(0,1)p_t \in (0,1)0.
  • Adaptive regularization: Use pt∈(0,1)p_t \in (0,1)1, where pt∈(0,1)p_t \in (0,1)2 tracks the maximum covariate norm to date; scaling with known pt∈(0,1)p_t \in (0,1)3 is possible.
  • Complexity: Total algorithmic run time is pt∈(0,1)p_t \in (0,1)4.
  • No additional tuning required: There are no separate step-size or clipping parameters beyond the inherent adaptivity and regularization.

A plausible implication is that Sigmoid-FTRL offers a turn-key approach for optimal assignment in design-based adaptive experiments using AIPW estimators.

7. Broader Context and Implications

Sigmoid-FTRL extends the literature connecting Neyman allocation and online convex optimization (OCO) beyond the Horvitz-Thompson estimator, addressing nonconvexity via convex decomposition and FTRL dynamics. It establishes sharp upper and lower regret bounds and supports practical confidence interval construction for deterministic potential outcomes, which is especially relevant for design-based inference in randomized controlled trials and sequential experimentation. The method’s adaptivity and lack of tuning requirements suggest applicability in practical online experiment pipelines without additional complexity (Chen et al., 25 Nov 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Sigmoid-FTRL.