Papers
Topics
Authors
Recent
Search
2000 character limit reached

Regression Adjustment Methods

Updated 12 July 2026
  • Regression adjustment is a statistical method that leverages auxiliary covariates to reduce variance without requiring a fully specified outcome model.
  • It improves the precision of treatment effect estimators in randomized experiments and observational studies using techniques such as OLS, debiasing, and cross‐fitting.
  • Recent advances include design-based estimators, high-dimensional adjustments, and methods addressing leverage bias, broadening its applicability in causal inference.

Regression adjustment denotes a family of procedures that use auxiliary information to remove predictable variation from an estimator or fitted object. In randomized experiments, its canonical role is to improve the precision of treatment-effect estimation without requiring a correctly specified outcome model; in that setting it is a design-based variance reduction device rather than a model-based causal model. In other domains, the same label refers to post-processing or second-stage correction steps, such as adjusting accepted draws in Approximate Bayesian Computation, smoothing separately fitted quantiles to enforce noncrossing, or subtracting pre-experiment predictability in online experimentation (Lei et al., 2018, Reluga et al., 2022, Blum, 2017, Rodrigues et al., 2015, Ting et al., 2023).

1. Core statistical idea

A standard randomized-experiment representation writes the adjusted average treatment effect estimator as

τ^(β1,β0)=n11i=1nZi(Yiβ1wi)n01i=1n(1Zi)(Yiβ0wi),\hat{\tau}(\bm{\beta}_1,\bm{\beta}_0) = n_1^{-1}\sum_{i=1}^n Z_i\bigl(Y_i-\bm{\beta}_1'\bm w_i\bigr) - n_0^{-1}\sum_{i=1}^n (1-Z_i)\bigl(Y_i-\bm{\beta}_0'\bm w_i\bigr),

or equivalently

τ^(β1,β0)=τ^(r0β1+r1β0)τ^w.\hat{\tau}(\bm{\beta}_1,\bm{\beta}_0) = \hat\tau - (r_0\bm\beta_1+r_1\bm\beta_0)'\hat{\bm\tau}_{\bm w}.

The unadjusted difference-in-means is the special case τ^(0,0)\hat\tau(\bm 0,\bm 0). In arm-wise ordinary least squares (OLS) adjustment, one fits separate regressions in treatment and control and uses τ^OLS=μ^1μ^0\hat\tau_{\mathrm{OLS}}=\hat\mu_1-\hat\mu_0 (Li et al., 2019, Lei et al., 2018).

A more general design-based abstraction is the generalized regression estimator

δ^GREG=δ^HTδ^xb^=n112nπ1Ru^+n112nxb^,\widehat{\delta}^{GREG} = \widehat{\delta}^{HT}-\widehat{\delta}_x'\widehat b = n^{-1}1'_{2n}\pi^{-1}R\widehat u+n^{-1}1'_{2n}x\widehat b,

with u^=yxb^\widehat u=y-x\widehat b. This formulation subsumes common regression adjustments, including OLS, and makes explicit that the adjustment is a covariate-based correction to a Horvitz–Thompson baseline rather than a superpopulation outcome model (Middleton, 2018).

Outside treatment-effect estimation, the same logic appears as an explicit correction step. In Approximate Bayesian Computation, accepted parameter draws are adjusted from simulated summaries s(i)s^{(i)} to the observed summaries sobss_{\rm obs} by

θc(i)=m^(sobs)+(θ(i)m^(s(i))),\theta_c^{(i)} = \hat m(s_{\rm obs}) + \bigl(\theta^{(i)} - \hat m(s^{(i)})\bigr),

with a heteroscedastic version that additionally rescales residuals. In Bayesian quantile regression, separately fitted quantiles are subsequently adjusted across τ\tau by Gaussian process regression to produce a monotone, noncrossing quantile process (Blum, 2017, Rodrigues et al., 2015).

2. Canonical theory in randomized experiments

Under Neyman’s finite-population model for complete randomization, potential outcomes are fixed, randomness comes only from treatment assignment, and the average treatment effect is

τ^(β1,β0)=τ^(r0β1+r1β0)τ^w.\hat{\tau}(\bm{\beta}_1,\bm{\beta}_0) = \hat\tau - (r_0\bm\beta_1+r_1\bm\beta_0)'\hat{\bm\tau}_{\bm w}.0

The difference-in-means estimator is unbiased, but regression adjustment can reduce variance by removing outcome variation explained by pre-treatment covariates. This gain does not require a correctly specified linear outcome model; the regression serves as an estimating device (Lei et al., 2018).

Within linear regression adjustment, the class

τ^(β1,β0)=τ^(r0β1+r1β0)τ^w.\hat{\tau}(\bm{\beta}_1,\bm{\beta}_0) = \hat\tau - (r_0\bm\beta_1+r_1\bm\beta_0)'\hat{\bm\tau}_{\bm w}.1

contains ANOVA, ANCOVA, and ANHECOVA. For completely randomized binary treatment, the unified asymptotic comparison shows that ANHECOVA dominates both ANOVA and ANCOVA, whereas there is no guaranteed ordering between ANOVA and ANCOVA. The underlying asymptotic variance calculation relies on classical linear regression theory with heteroscedastic errors and therefore does not assume that the postulated linear model is correct (Reluga et al., 2022).

A broader design-based theory extends these results beyond complete randomization. For arbitrary identified designs, the exact variance-minimizing fixed coefficient within the generalized regression class is

τ^(β1,β0)=τ^(r0β1+r1β0)τ^w.\hat{\tau}(\bm{\beta}_1,\bm{\beta}_0) = \hat\tau - (r_0\bm\beta_1+r_1\bm\beta_0)'\hat{\bm\tau}_{\bm w}.2

This perspective yields the 3HT! estimator, which estimates the optimal coefficient by Horvitz–Thompson methods, and the recursive 2R! estimator, described as regression-adjusted regression adjustment. In complete randomization, this framework recovers the familiar result that OLS with treatment-by-covariate interactions is asymptotically optimal (Middleton, 2018).

Regression adjustment also interacts systematically with design improvements. Under rerandomization, the asymptotic theory shows that rerandomization never hurts either the sampling precision or the estimated precision of a given estimator, and that the recommended implementation is rerandomization in the design, regression adjustment in the analysis, followed by the Huber–White robust standard error (Li et al., 2019).

An alternative weighted implementation is tyranny-of-the-minority (ToM) regression adjustment, defined in complete randomization as the weighted least-squares coefficient of τ^(β1,β0)=τ^(r0β1+r1β0)τ^w.\hat{\tau}(\bm{\beta}_1,\bm{\beta}_0) = \hat\tau - (r_0\bm\beta_1+r_1\bm\beta_0)'\hat{\bm\tau}_{\bm w}.3 in

τ^(β1,β0)=τ^(r0β1+r1β0)τ^w.\hat{\tau}(\bm{\beta}_1,\bm{\beta}_0) = \hat\tau - (r_0\bm\beta_1+r_1\bm\beta_0)'\hat{\bm\tau}_{\bm w}.4

ToM is asymptotically equivalent to Lin’s interacted regression in completely randomized experiments, is optimal in the class of linearly adjusted estimators, and is extended in the source to stratified randomized experiments, completely randomized survey experiments, and cluster randomized experiments (Lu et al., 2022).

3. High-dimensional covariates, leverage, and debiasing

When the number of covariates τ^(β1,β0)=τ^(r0β1+r1β0)τ^w.\hat{\tau}(\bm{\beta}_1,\bm{\beta}_0) = \hat\tau - (r_0\bm\beta_1+r_1\bm\beta_0)'\hat{\bm\tau}_{\bm w}.5 diverges with τ^(β1,β0)=τ^(r0β1+r1β0)τ^w.\hat{\tau}(\bm{\beta}_1,\bm{\beta}_0) = \hat\tau - (r_0\bm\beta_1+r_1\bm\beta_0)'\hat{\bm\tau}_{\bm w}.6, the classical low-dimensional asymptotics fail because the OLS-adjusted estimator develops a leverage-driven bias. In a completely randomized treatment-control experiment with centered covariates, the maximum leverage score

τ^(β1,β0)=τ^(r0β1+r1β0)τ^w.\hat{\tau}(\bm{\beta}_1,\bm{\beta}_0) = \hat\tau - (r_0\bm\beta_1+r_1\bm\beta_0)'\hat{\bm\tau}_{\bm w}.7

governs the bias. A design-based debiased estimator is

τ^(β1,β0)=τ^(r0β1+r1β0)τ^w.\hat{\tau}(\bm{\beta}_1,\bm{\beta}_0) = \hat\tau - (r_0\bm\beta_1+r_1\bm\beta_0)'\hat{\bm\tau}_{\bm w}.8

where τ^(β1,β0)=τ^(r0β1+r1β0)τ^w.\hat{\tau}(\bm{\beta}_1,\bm{\beta}_0) = \hat\tau - (r_0\bm\beta_1+r_1\bm\beta_0)'\hat{\bm\tau}_{\bm w}.9 is the sample analogue of the leverage-weighted residual term. Under favorable leverage scores of order τ^(0,0)\hat\tau(\bm 0,\bm 0)0, ordinary regression adjustment is asymptotically normal when τ^(0,0)\hat\tau(\bm 0,\bm 0)1, whereas the debiased estimator is asymptotically normal when

τ^(0,0)\hat\tau(\bm 0,\bm 0)2

The same analysis recommends HC3 variance estimation and notes that trimming covariates to reduce leverage can improve finite-sample behavior (Lei et al., 2018).

Cross-fitting sharpens this picture by replacing each unit’s fitted coefficient with a leave-one-out coefficient from the same arm. The resulting cross-fitted regression adjustment admits a higher-order expansion with the same leading linear and second-order terms as bias-corrected OLS but a smaller remainder. Under favorable leverage designs with τ^(0,0)\hat\tau(\bm 0,\bm 0)3, asymptotic normality holds when

τ^(0,0)\hat\tau(\bm 0,\bm 0)4

and the paper proposes a modified HC3 variance estimator that corrects the second-order terms missed by conventional HC0 and HC3 (Chiang et al., 2023).

A different route is decorrelation. Instead of fitting adjustment functions on exactly the same units used in the final residual average, overlapping randomized subsets τ^(0,0)\hat\tau(\bm 0,\bm 0)5 and τ^(0,0)\hat\tau(\bm 0,\bm 0)6 are constructed so that the fitting indicators and averaging indicators are independent by design. The decorrelated estimator τ^(0,0)\hat\tau(\bm 0,\bm 0)7 is then asymptotically normal under mere consistency of the first-stage regression fits,

τ^(0,0)\hat\tau(\bm 0,\bm 0)8

without any explicit τ^(0,0)\hat\tau(\bm 0,\bm 0)9-type rate. The source gives oracle, nonasymptotic, and asymptotic guarantees for OLS, sparse high-dimensional linear regression, and nonparametric constrained least squares (Su et al., 2023).

A further refinement expands the OLS remainder by a Neumann series. Writing

τ^OLS=μ^1μ^0\hat\tau_{\mathrm{OLS}}=\hat\mu_1-\hat\mu_00

one obtains a hierarchy of degree-τ^OLS=μ^1μ^0\hat\tau_{\mathrm{OLS}}=\hat\mu_1-\hat\mu_01 corrected estimators,

τ^OLS=μ^1μ^0\hat\tau_{\mathrm{OLS}}=\hat\mu_1-\hat\mu_02

Under mild regularity and diffuse leverage, the degree-τ^OLS=μ^1μ^0\hat\tau_{\mathrm{OLS}}=\hat\mu_1-\hat\mu_03 estimator is asymptotically normal whenever

τ^OLS=μ^1μ^0\hat\tau_{\mathrm{OLS}}=\hat\mu_1-\hat\mu_04

Thus τ^OLS=μ^1μ^0\hat\tau_{\mathrm{OLS}}=\hat\mu_1-\hat\mu_05 yields the τ^OLS=μ^1μ^0\hat\tau_{\mathrm{OLS}}=\hat\mu_1-\hat\mu_06 regime, τ^OLS=μ^1μ^0\hat\tau_{\mathrm{OLS}}=\hat\mu_1-\hat\mu_07 yields τ^OLS=μ^1μ^0\hat\tau_{\mathrm{OLS}}=\hat\mu_1-\hat\mu_08, and higher τ^OLS=μ^1μ^0\hat\tau_{\mathrm{OLS}}=\hat\mu_1-\hat\mu_09 enlarges the admissible growth region (Song, 11 Nov 2025).

Recent work also provides a non-asymptotic, design-based analysis of arm-wise OLS adjustment under complete randomization, including the case δ^GREG=δ^HTδ^xb^=n112nπ1Ru^+n112nxb^,\widehat{\delta}^{GREG} = \widehat{\delta}^{HT}-\widehat{\delta}_x'\widehat b = n^{-1}1'_{2n}\pi^{-1}R\widehat u+n^{-1}1'_{2n}x\widehat b,0. The key finite-sample confidence interval has radius

δ^GREG=δ^HTδ^xb^=n112nπ1Ru^+n112nxb^,\widehat{\delta}^{GREG} = \widehat{\delta}^{HT}-\widehat{\delta}_x'\widehat b = n^{-1}1'_{2n}\pi^{-1}R\widehat u+n^{-1}1'_{2n}x\widehat b,1

where δ^GREG=δ^HTδ^xb^=n112nπ1Ru^+n112nxb^,\widehat{\delta}^{GREG} = \widehat{\delta}^{HT}-\widehat{\delta}_x'\widehat b = n^{-1}1'_{2n}\pi^{-1}R\widehat u+n^{-1}1'_{2n}x\widehat b,2 and δ^GREG=δ^HTδ^xb^=n112nπ1Ru^+n112nxb^,\widehat{\delta}^{GREG} = \widehat{\delta}^{HT}-\widehat{\delta}_x'\widehat b = n^{-1}1'_{2n}\pi^{-1}R\widehat u+n^{-1}1'_{2n}x\widehat b,3 are variance-adaptive martingale quantities and δ^GREG=δ^HTδ^xb^=n112nπ1Ru^+n112nxb^,\widehat{\delta}^{GREG} = \widehat{\delta}^{HT}-\widehat{\delta}_x'\widehat b = n^{-1}1'_{2n}\pi^{-1}R\widehat u+n^{-1}1'_{2n}x\widehat b,4 is a Stein-type bias bound. The proof uses the Johnson graph, a variance-adaptive Doob martingale, Freedman’s inequality, and Stein’s method of exchangeable pairs; the resulting analysis highlights leverages and cross-leverages as the geometric quantities governing stability (Song, 19 Nov 2025).

4. Beyond the average treatment effect under complete randomization

In covariate-adaptive randomization with imperfect compliance, regression adjustment becomes a doubly robust augmentation of a local average treatment effect estimator. The central score uses two regression adjustments, one for δ^GREG=δ^HTδ^xb^=n112nπ1Ru^+n112nxb^,\widehat{\delta}^{GREG} = \widehat{\delta}^{HT}-\widehat{\delta}_x'\widehat b = n^{-1}1'_{2n}\pi^{-1}R\widehat u+n^{-1}1'_{2n}x\widehat b,5 and one for δ^GREG=δ^HTδ^xb^=n112nπ1Ru^+n112nxb^,\widehat{\delta}^{GREG} = \widehat{\delta}^{HT}-\widehat{\delta}_x'\widehat b = n^{-1}1'_{2n}\pi^{-1}R\widehat u+n^{-1}1'_{2n}x\widehat b,6,

δ^GREG=δ^HTδ^xb^=n112nπ1Ru^+n112nxb^,\widehat{\delta}^{GREG} = \widehat{\delta}^{HT}-\widehat{\delta}_x'\widehat b = n^{-1}1'_{2n}\pi^{-1}R\widehat u+n^{-1}1'_{2n}x\widehat b,7

where δ^GREG=δ^HTδ^xb^=n112nπ1Ru^+n112nxb^,\widehat{\delta}^{GREG} = \widehat{\delta}^{HT}-\widehat{\delta}_x'\widehat b = n^{-1}1'_{2n}\pi^{-1}R\widehat u+n^{-1}1'_{2n}x\widehat b,8 and δ^GREG=δ^HTδ^xb^=n112nπ1Ru^+n112nxb^,\widehat{\delta}^{GREG} = \widehat{\delta}^{HT}-\widehat{\delta}_x'\widehat b = n^{-1}1'_{2n}\pi^{-1}R\widehat u+n^{-1}1'_{2n}x\widehat b,9 combine inverse-probability weighting with stratum-specific outcome and compliance regressions. This estimator is consistent and asymptotically normal under heterogeneous assignment probabilities and misspecified working regressions, the optimal linear estimator is most efficient among linearly adjusted LATE estimators, and sufficiently rich nonparametric or regularized adjustments attain the semiparametric efficiency bound (Jiang et al., 2022).

Regression adjustment can also target the entire outcome distribution. In randomized controlled trials with u^=yxb^\widehat u=y-x\widehat b0 treatment arms, the distributional treatment effect is

u^=yxb^\widehat u=y-x\widehat b1

and the adjusted estimator replaces the empirical arm-specific cumulative distribution function with an average of estimated conditional distributions,

u^=yxb^\widehat u=y-x\widehat b2

Under canonical-link distributional regression, the oracle variance of the adjusted estimator is no larger than that of the simple empirical estimator, and under correct specification the adjusted DTE estimator attains the semiparametric efficiency bound at each fixed u^=yxb^\widehat u=y-x\widehat b3; the same framework extends to probability treatment effects and quantile treatment effects (Oka et al., 2024).

Under network interference, outcomes depend on own treatment and treated-neighbor exposure, yet the adjustment principle persists. For the average direct effect

u^=yxb^\widehat u=y-x\widehat b4

the linear regression-adjusted estimator is built from arm-specific OLS coefficients,

u^=yxb^\widehat u=y-x\widehat b5

and satisfies a central limit theorem with asymptotic variance u^=yxb^\widehat u=y-x\widehat b6. Within the class of linear adjustments, u^=yxb^\widehat u=y-x\widehat b7 is minimal, and a kernel-and-trimming nonparametric extension is asymptotically optimal in a broader nonlinear class (Fan et al., 9 Feb 2025).

Covariate-adaptive randomization with missing covariates introduces a second layer of adjustment. The analyzed procedures combine missingness handling—CCOV, IMP, and MIM—with Fisher, Lin, or ToM regression. For unequal allocation, Lin and ToM are asymptotically normal with consistent variance estimators, and the variance ordering

u^=yxb^\widehat u=y-x\widehat b8

shows that MIM is asymptotically most efficient among the three missing-data strategies considered. The source recommends Fisher regression when u^=yxb^\widehat u=y-x\widehat b9, ToM regression when s(i)s^{(i)}0, and stratum-specific adjustment when strata matter (Fu et al., 13 Aug 2025).

5. Other methodological roles of regression adjustment

In difference-in-differences, regression adjustment is not merely the inclusion of covariate main effects. A covariate is a confounder only if it helps explain why untreated outcome trends differ between treated and comparison groups over time. For a time-invariant confounder with a time-varying outcome effect, the appropriate regression includes a covariate-by-time interaction,

s(i)s^{(i)}1

rather than only a main effect for s(i)s^{(i)}2. When a time-varying covariate is affected by treatment, standard regression adjustment may estimate only a direct effect rather than the ATT, and the source explicitly points to g-methods and inverse probability weighting as more appropriate tools (Zeldow et al., 2019).

In online experimentation, regression adjustment is often called CUPED. If s(i)s^{(i)}3 is the pre-experiment value of the goal metric and s(i)s^{(i)}4, then the basic variance reduction factor is

s(i)s^{(i)}5

For an engineered pre-treatment covariate s(i)s^{(i)}6 with s(i)s^{(i)}7, the advanced factor is s(i)s^{(i)}8. Under the temporal predictability assumption s(i)s^{(i)}9, the source proves

sobss_{\rm obs}0

implying that sophisticated safe feature engineering can yield at most an additional sobss_{\rm obs}1 variance reduction, or equivalently at most a sobss_{\rm obs}2 further reduction in confidence-interval width, relative to basic CUPED (Ting et al., 2023).

In Approximate Bayesian Computation, regression adjustment is a post-processing device that corrects the imperfect match between simulated and observed summary statistics after rejection sampling. The method is motivated by the curse of dimensionality in the summary vector sobss_{\rm obs}3, and the adjusted draws typically shrink posterior spread relative to rejection ABC. The source gives both homoscedastic and heteroscedastic corrections, proves bias and variance expansions for the resulting posterior density estimator, and reports that simple linear adjustment reduced error compared to rejection in a population-genetics application (Blum, 2017).

In Bayesian quantile regression, a two-stage regression adjustment is used to remove quantile crossing. After fitting each quantile separately with an asymmetric Laplace likelihood, induced quantiles from neighboring auxiliary quantile levels are treated as noisy observations in a Gaussian process regression with kernel

sobss_{\rm obs}4

The adjusted estimate is a weighted average of induced quantiles, and there exists a bandwidth sobss_{\rm obs}5 such that the adjusted quantile curve is noncrossing (Rodrigues et al., 2015).

The label can also be used differently in machine vision. In imbalanced regression, "hierarchical classification adjustment" constructs multiple classifiers at different target granularities and combines coarse and fine predictions through additive or log-space adjustment, together with a range-preserving distillation step, in order to reduce the trade-off between class balance and quantization error (Xiong et al., 2023).

6. Interpretation, limits, and recurrent misconceptions

A central misconception is to treat regression adjustment as the source of causal identification. In randomized experiments, identification comes from randomization; adjustment is used to improve efficiency, and its theoretical justification can be entirely randomization-based and robust to misspecification of the working linear model (Reluga et al., 2022, Lei et al., 2018).

A second misconception is that residuals automatically recover the latent object one wants to compare. In benchmarking or data-adjustment problems with

sobss_{\rm obs}6

using least-squares residuals sobss_{\rm obs}7 as adjusted performance measures is problematic because residuals recover only the component of sobss_{\rm obs}8 that is orthogonal to the fitted covariate space. Systematic unit differences correlated with sobss_{\rm obs}9 are absorbed into θc(i)=m^(sobs)+(θ(i)m^(s(i))),\theta_c^{(i)} = \hat m(s_{\rm obs}) + \bigl(\theta^{(i)} - \hat m(s^{(i)})\bigr),0, not left in the residuals, and high leverage shrinks residuals toward zero through

θc(i)=m^(sobs)+(θ(i)m^(s(i))),\theta_c^{(i)} = \hat m(s_{\rm obs}) + \bigl(\theta^{(i)} - \hat m(s^{(i)})\bigr),1

The residual is therefore a measure of partial performance, not a generally valid measure of true adjusted performance (Duembgen, 2012).

The limits are also design-specific. Uniform variance dominance is a special feature of classical linear adjustment under complete randomization; it generally fails when treatment assignment depends on covariates, and it does not generally extend to generalized linear models because treatment coefficients can cease to target the same estimand across specifications (Reluga et al., 2022). In difference-in-differences, main-effect adjustment can be inadequate because time-varying confounding requires time-varying regression structure, and adjustment for post-treatment covariates can bias the ATT (Zeldow et al., 2019). In online experimentation, the temporal predictability bound implies that much of the feasible precision gain is already captured by the pre-experiment value of the goal metric itself (Ting et al., 2023).

Current theory also makes the remaining frontier explicit. High-dimensional randomized-experiment work now supports bias correction, cross-fitting, decorrelation, Neumann-series corrections, and finite-sample oracle confidence intervals, but the automatic selection of correction degree and fully rigorous data-driven finite-sample standard errors remain open problems in the cited analyses (Song, 11 Nov 2025, Song, 19 Nov 2025). A plausible implication is that future work will continue to treat regression adjustment less as a single estimator and more as a design-aware family of orthogonalization, residualization, and bias-control techniques whose validity depends on how the auxiliary information enters the assignment mechanism, the estimand, and the geometry of the covariates.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (19)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Regression Adjustment.