Papers
Topics
Authors
Recent
Search
2000 character limit reached

Theoretical Properties of Covariate-Adaptive Randomization with a Diverging Number of Covariates

Published 13 Aug 2026 in stat.ME and math.ST | (2608.13442v1)

Abstract: Covariate-adaptive randomization procedures are widely used in clinical trials to improve covariate balance. In modern applications, experimenters often have access to many covariates, motivating the need for a theory of covariate-adaptive randomization procedures with a diverging number of covariates. In this paper, we study the theoretical properties of two unified families of covariate-adaptive randomization procedures under high-dimensional settings. For the two procedures, we establish the convergence rate of the imbalance measure corresponding to the specified covariates. In addition, for one of them, we study the asymptotic properties of the imbalance of unspecified covariates under and apply these results to derive the asymptotic properties of the difference-in-means estimator for the average treatment effect and construct asymptotic 95% confidence intervals. Furthermore, we provide extensive numerical and empirical studies to illustrate the practical relevance of our theoretical results.

Authors (2)

Summary

  • The paper develops asymptotic theory for IE-CAR and IR-CAR when the feature dimension diverges, showing that both outperform complete randomization only when q grows slower than n.
  • IR-CAR avoids the shift problem under unequal allocation, improves balance for unspecified covariates, and supports asymptotically valid average treatment effect inference without outcome-model assumptions.
  • The results establish a trade-off controlled by γ: larger γ supports faster covariate-dimension growth and robust normality, while slowing specified-covariate balance, with no CAR advantage when q is of order n.

This paper develops asymptotic theory for covariate-adaptive randomization (CAR) procedures when the number of covariates, or more precisely the dimension qq of the balancing feature map, diverges with the sample size nn. The authors analyze two unified families of procedures—imbalance-efficient CAR (IE-CAR), based on the framework of Ma et al., and imbalance-robust CAR (IR-CAR), proposed by Zhang et al.—and establish convergence rates for the imbalance measure, laws of large numbers and central limit theorems for covariates not included in the feature map, and asymptotically valid inference for the average treatment effect (ATE) under IR-CAR. To date, theoretical guarantees for CAR procedures have been almost exclusively low-dimensional; this work extends them to regimes where both pp and qq grow with nn.

Framework

The setting is a sequentially randomized experiment with allocation proportion π(0,1)\pi \in (0,1). Each unit has covariate vector XiRpX_i \in \mathbb{R}^p, mapped through a feature map ϕ(Xi):RpRq\phi(X_i): \mathbb{R}^p \to \mathbb{R}^q. The imbalance vector is Λn=i=1n(Tiπ)ϕ(Xi)\Lambda_n = \sum_{i=1}^n (T_i - \pi)\phi(X_i), and the imbalance measure is Imbn=Λn2\mathrm{Imb}_n = \|\Lambda_n\|^2. Under complete randomization (CR), nn0, which serves as the benchmark throughout.

The two procedure families differ in how the allocation probability responds to imbalance. IE-CAR uses a constant biased-coin rule: unit nn1 is assigned to treatment with probability nn2 if nn3 and probability nn4 if the inner product is positive, with nn5. The influence of imbalance on assignment remains at a constant level across steps. IR-CAR instead scales the imbalance by nn6 with nn7, assigning treatment with probability nn8 for a non-increasing allocation function nn9 with pp0; the influence of imbalance on assignment decays over time. A recommended parametric form is pp1.

Three assumptions drive the analysis: i.i.d. covariates; moment conditions requiring pp2 and bounded moments of all one-dimensional projections (satisfied under sub-exponential tails); and bounded nonzero eigenvalues of pp3. These are mild by high-dimensional standards. Notably, if pp4 satisfies a small-ball condition, the moment requirements can be weakened to pp5 and pp6 without changing any conclusion.

Convergence rates under IE-CAR

The first main result establishes that under IE-CAR, if pp7 then pp8 with lower bound pp9; if qq0, then qq1. When Assumption 2 holds for all qq2, the upper bound sharpens to qq3 for any qq4.

The implication is a phase transition: CAR retains an advantage over CR only while the feature dimension grows sublinearly in qq5. Once qq6 is of order qq7 or larger, no benefit remains—the imbalance measure is of the same order as under CR. This is a direct quantitative answer to how many covariates can be usefully balanced.

A significant limitation is acknowledged here: the analysis of additional covariates not included in the feature map, which in low dimensions relies on Markov chain tools (drift conditions, invariant measures, Poisson equations), does not extend to high dimensions. Consequently, no asymptotic properties of additional covariates or of treatment effect estimators are available under IE-CAR in this regime.

Convergence rates and additional-covariate behavior under IR-CAR

For IR-CAR, the rate results are three-regime: if qq8, qq9 lies between nn0 and an upper bound of order nn1; if nn2 and nn3, the bounds become nn4 and nn5; and if nn6, again nn7. Thus IR-CAR exhibits the same qualitative phase transition as IE-CAR, but with a slower rate when nn8—the price paid for robustness to unspecified covariates.

The central contribution for IR-CAR concerns additional covariates nn9 (possibly unobservable) not used in randomization. Under dimension conditions involving π(0,1)\pi \in (0,1)0 and π(0,1)\pi \in (0,1)1, the paper proves:

  • No shift problem: π(0,1)\pi \in (0,1)2, even under unequal allocation (π(0,1)\pi \in (0,1)3). This contrasts sharply with IE-CAR under continuous covariates and unequal allocation, where the shift problem—imbalance centered away from zero—is known to arise.
  • Asymptotic normality: the pair π(0,1)\pi \in (0,1)4 converges to a bivariate standard normal, where π(0,1)\pi \in (0,1)5.

Because π(0,1)\pi \in (0,1)6, the asymptotic variance of the additional-covariate imbalance is bounded above by its CR counterpart. IR-CAR therefore dominates CR for unspecified covariates—a guarantee unavailable for IE-CAR even in low dimensions. The parameter π(0,1)\pi \in (0,1)7 governs a trade-off: larger π(0,1)\pi \in (0,1)8 permits faster covariate-dimension growth in the normality result but slows the convergence of the specified-covariate imbalance.

Inference for the average treatment effect

Under IR-CAR, the difference-in-means estimator π(0,1)\pi \in (0,1)9 satisfies

XiRpX_i \in \mathbb{R}^p0

where XiRpX_i \in \mathbb{R}^p1 aggregates residual variances after projecting potential outcomes onto XiRpX_i \in \mathbb{R}^p2, and XiRpX_i \in \mathbb{R}^p3 captures variance reduction from balancing. Consistent plug-in estimators are constructed via group-wise OLS regressions of centered outcomes on XiRpX_i \in \mathbb{R}^p4, yielding asymptotically valid Wald confidence intervals. Importantly, these results require no linear model relating outcomes to covariates—only finite second moments of potential outcomes—addressing a common criticism of inference under stratified CAR.

Numerical evidence

Simulations with XiRpX_i \in \mathbb{R}^p5 for XiRpX_i \in \mathbb{R}^p6, under both Gaussian continuous and binary discrete covariates and equal or unequal allocation, closely match the theory. On log–log plots, the estimated growth exponent of XiRpX_i \in \mathbb{R}^p7 is approximately XiRpX_i \in \mathbb{R}^p8 under IE-CAR when XiRpX_i \in \mathbb{R}^p9, approximately ϕ(Xi):RpRq\phi(X_i): \mathbb{R}^p \to \mathbb{R}^q0 under IR-CAR when ϕ(Xi):RpRq\phi(X_i): \mathbb{R}^p \to \mathbb{R}^q1 and ϕ(Xi):RpRq\phi(X_i): \mathbb{R}^p \to \mathbb{R}^q2 when ϕ(Xi):RpRq\phi(X_i): \mathbb{R}^p \to \mathbb{R}^q3, and ϕ(Xi):RpRq\phi(X_i): \mathbb{R}^p \to \mathbb{R}^q4 under CR throughout—all consistent with the predicted rates. For ϕ(Xi):RpRq\phi(X_i): \mathbb{R}^p \to \mathbb{R}^q5, all procedures converge to the CR exponent, confirming empirically that CAR loses its advantage once ϕ(Xi):RpRq\phi(X_i): \mathbb{R}^p \to \mathbb{R}^q6.

For treatment effect estimation with ϕ(Xi):RpRq\phi(X_i): \mathbb{R}^p \to \mathbb{R}^q7 and ϕ(Xi):RpRq\phi(X_i): \mathbb{R}^p \to \mathbb{R}^q8 up to 56, the difference-in-means estimator shows negligible bias, standard errors track Monte Carlo standard deviations (with mild upward bias at larger ϕ(Xi):RpRq\phi(X_i): \mathbb{R}^p \to \mathbb{R}^q9), and empirical coverage of nominal 95% intervals stays between roughly 0.94 and 0.96 across all configurations, including unequal allocation.

Real-data illustration

Applying the procedures to the Mayo Clinic primary biliary cirrhosis trial data (312 randomized patients, 16 baseline covariates), redesigned assignments over 5,000 replications show the expected ordering: CR yields the largest mean imbalance (308.53 under Λn=i=1n(Tiπ)ϕ(Xi)\Lambda_n = \sum_{i=1}^n (T_i - \pi)\phi(X_i)0), IE-CAR the smallest (12.66), and IR-CAR intermediate values increasing with Λn=i=1n(Tiπ)ϕ(Xi)\Lambda_n = \sum_{i=1}^n (T_i - \pi)\phi(X_i)1 (29.53 to 75.86). Using synthetic potential outcomes imputed by a causal forest, the 95% confidence intervals under IR-CAR with Λn=i=1n(Tiπ)ϕ(Xi)\Lambda_n = \sum_{i=1}^n (T_i - \pi)\phi(X_i)2 are substantially shorter than under CR (e.g., width 4.75 versus 8.57 under equal allocation), while all intervals cover zero—consistent with the original finding that D-penicillamine provides no clear survival benefit. Confidence intervals are not reported for IE-CAR or IR-CAR with Λn=i=1n(Tiπ)ϕ(Xi)\Lambda_n = \sum_{i=1}^n (T_i - \pi)\phi(X_i)3 because the required dimension conditions fail, a candid reflection of the theory's applicability boundaries.

Limitations and open questions

Several restrictions deserve emphasis. First, component-wise convergence rates of Λn=i=1n(Tiπ)ϕ(Xi)\Lambda_n = \sum_{i=1}^n (T_i - \pi)\phi(X_i)4 are established only under the strong assumption of exchangeable feature components; generalizing beyond this is left open. Second, the asymptotic analysis of additional covariates and ATE inference is carried out only for IR-CAR; extending these results to IE-CAR in high dimensions appears difficult because classical Markov chain machinery breaks down. Third, when Λn=i=1n(Tiπ)ϕ(Xi)\Lambda_n = \sum_{i=1}^n (T_i - \pi)\phi(X_i)5, neither procedure improves on CR, and whether any CAR procedure can perform satisfactorily in that regime remains unresolved. Fourth, multi-armed experiments and non-Markovian CAR frameworks designed to address the shift problem under unequal allocation have not yet been analyzed in high-dimensional settings. Finally, the practical choice of Λn=i=1n(Tiπ)ϕ(Xi)\Lambda_n = \sum_{i=1}^n (T_i - \pi)\phi(X_i)6 involves a genuine trade-off between allowable covariate dimension and balance speed, for which the paper provides guidance but no automatic selection rule.

Conclusion

This paper supplies the first systematic high-dimensional asymptotic theory for two unified families of CAR procedures. Its principal findings are a sharp dimensional phase transition at Λn=i=1n(Tiπ)ϕ(Xi)\Lambda_n = \sum_{i=1}^n (T_i - \pi)\phi(X_i)7 for the specified-covariate imbalance, dominance of IR-CAR over CR for unspecified covariates together with absence of the shift problem under unequal allocation, and distribution-free asymptotic inference for the ATE with valid variance estimation. The simulation and PBC analyses confirm that the predicted rates and coverage properties hold in finite samples, supporting the use of CAR in trials collecting large numbers of baseline covariates, provided the feature-map dimension grows strictly slower than Λn=i=1n(Tiπ)ϕ(Xi)\Lambda_n = \sum_{i=1}^n (T_i - \pi)\phi(X_i)8.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.