High-Dimensional Empirical Likelihood Weighting
- The paper introduces doubly penalized empirical likelihood methods that simultaneously select variables and estimating equations for sparse, efficient high-dimensional inference.
- It demonstrates how adjustments like pseudo-observations and soft calibration overcome feasibility issues such as convex-hull failures and unstable inverse probability weighting.
- The approach is applied across randomized trials, causal inference, and dependent data analyses, offering robust, likelihood-free estimation under diverging dimensions.
High-dimensional empirical likelihood weighting denotes a family of nonparametric or semi-parametric methods that assign probability weights to observations so that estimating equations, calibration equations, or auxiliary moment restrictions are satisfied while the dimension of covariates, parameters, or moment conditions may diverge with sample size. In its classical form, empirical likelihood maximizes a product of weights under normalization and moment constraints; in high-dimensional regimes, this basic construction is modified through double penalization, adjusted pseudo-observations, rescaling, sample splitting, side-information constraints, or soft balancing to remain feasible, stable, and statistically valid (Chang et al., 2017, Liu et al., 2010, Wang et al., 2023, Liang et al., 2020). Across longitudinal models, randomized trials, treatment-effect estimation, big-data subsampling, voluntary samples, structural equation models, and dependent time series, the common objective is to retain the likelihood-free flexibility of empirical likelihood while overcoming convex-hull failure, unstable inverse probability weighting, and nonstandard asymptotics that arise when , , or both are large relative to (Li et al., 2021, Fan et al., 2022, Lahiri et al., 2013, Chang et al., 26 Feb 2025).
1. Foundational formulation
Empirical likelihood is built on probability weights over observed data. A standard formulation maximizes
where is an estimating function or moment restriction (Liu et al., 2010). In treatment-control settings, the construction is often group-specific: two-sample empirical likelihood weighted estimation uses separate weights and for treated and control samples, constrained so that weighted covariate moments equal pooled moments (Tan et al., 2020). In semiparametric weighting with side information, the resulting empirical likelihood weights take the explicit form
with determined by the constraint equations (Wang et al., 2023).
The high-dimensional weighting problem appears when the number of covariates, parameters, or moment restrictions is large and may diverge with . In this regime, standard empirical likelihood encounters several difficulties. The convex hull condition can fail, the chi-square approximation can have low precision, and exact balancing constraints can become infeasible (Liu et al., 2010). In randomized trials and causal inference, additional instability arises when inverse probability weighting uses very small inclusion or propensity probabilities (Fan et al., 2022). In high-dimensional testing, classical likelihood ratio approximations based on fixed-0 chi-square limits no longer apply (Qi et al., 2022, Ciuperca et al., 2015).
A central feature of the literature is that the weights are not merely numerical devices for reweighting means. They encode identification, select among estimating equations, absorb auxiliary information, and define calibrated empirical distributions. This suggests that “weighting” in the high-dimensional empirical likelihood literature is best understood as a constrained optimization mechanism for estimation and inference rather than as a narrow alternative to inverse probability weighting.
2. Penalization, adjustment, and feasibility in high dimensions
A major line of development introduces penalties on both the target parameter and the Lagrange multipliers. The doubly penalized empirical likelihood criterion is written as
1
or in equivalent notational variants, with 2 and 3 chosen from classes including SCAD, MCP, and 4 penalties (Chang et al., 2017, Li et al., 2021, Chang et al., 26 Feb 2025). The penalty on 5 promotes sparse parameter estimation, while the penalty on 6 promotes sparsity in the active estimating equations, so that moment selection and variable selection are carried out simultaneously (Chang et al., 2017). Several papers state explicitly that both the dimensionalities of model parameters and estimating equations can grow exponentially with the sample size under sparsity and regularity conditions (Chang et al., 2017, Li et al., 2021).
A distinct route to feasibility is adjusted empirical likelihood. When sample size is small or the dimension of the estimating function is high, the empirical likelihood equations may have no solution. Adjusted empirical likelihood adds pseudo-observations such as
7
in the scalar case, or two pseudo-observations in the multivariate case, so that the solution always exists and the statistic remains well defined (Liu et al., 2010). With a specific level of adjustment, the adjusted empirical likelihood attains the high-order precision of the Bartlett correction and preserves the guarantee of existence (Liu et al., 2010).
For high-dimensional mean inference, another strategy is penalized empirical likelihood in which the likelihood is modified by a penalty on deviations of the weighted sample mean from the hypothesized mean: 8 with component-wise scaling 9 and recommended 0 (Lahiri et al., 2013). This formulation is well-defined for all 1 and avoids inversion of the sample covariance matrix (Lahiri et al., 2013).
These constructions address a common technical fact: in high dimensions, empirical likelihood must often relax the classical exact finite-dimensional geometry without abandoning its moment-based logic.
3. Weight construction, calibration, and auxiliary information
The explicit form of the weights varies with the inferential task, but several recurring patterns appear. In big-data M-estimation under capture-recapture sampling, empirical likelihood weighting replaces unstable inverse-probability factors by weights
2
obtained from normalization, inclusion-probability constraints, and optional auxiliary-moment constraints (Fan et al., 2022). The resulting estimator
3
circumvents the use of inverse probabilities and utilizes auxiliary information including the size and certain sample moments of big data (Fan et al., 2022).
In voluntary samples under a nonignorable sample selection model, the final empirical likelihood weights are constructed by maximizing 4 subject to normalization, a bias calibration constraint, and a benchmarking constraint: 5 This weighting scheme is designed to reduce selection bias when sample inclusion depends on both 6 and 7 (Kim et al., 2022).
In randomized trials, empirical likelihood weighted estimation imposes covariate moment constraints motivated by randomization. The treated-arm and control-arm weights have the forms
8
and the average treatment effect estimator is the difference of weighted means (Tan et al., 2020). In the machine-learning and data-splitting empirical likelihood approach for high-dimensional covariates, the weights in each arm become
9
where the nuisance regressions are fit on data not used for estimation (Liang et al., 2020).
Other weighting schemes incorporate side information or robustness. In structural equation models with side information, empirical likelihood weights again take the rational form 0 and are used to construct EL-weighted covariance estimators (Wang et al., 2023). In high-dimensional ATE inference with multiple working propensity score models, the weights are defined under calibration constraints and soft covariate balancing constraints, with
1
and the inequality-based soft calibration relaxes exact balancing, which is infeasible in high dimensions (Xia et al., 30 Aug 2025). In robust longitudinal analysis, leverage points are downweighted through a diagonal weight matrix 2 in robust estimating functions, and in weighted jackknife empirical likelihood, depth-based weights
3
tilt the likelihood so that outliers receive smaller weights (Li et al., 2021, Sang et al., 2019).
4. Asymptotic theory and nonstandard limiting behavior
The asymptotic theory of high-dimensional empirical likelihood weighting departs sharply from the fixed-dimensional Wilks paradigm. In high-dimensional sparse estimation with doubly penalized empirical likelihood, the main results are consistency, sparsity, and asymptotically normal nonzero components, with selection consistency for both model parameters and estimating equations (Chang et al., 2017). In robust penalized empirical likelihood for longitudinal data, oracle properties and bounded influence functions are established, and the nonzero components have the same distribution as if the true support set were known under appropriate regularity conditions (Li et al., 2021).
Generalized empirical likelihood for high-dimensional moment restrictions with weakly dependent data yields consistency with rates, asymptotic normality, and a high-dimensional Wilks phenomenon. Under appropriate restrictions on the growth rates of the dimensions and the dependence strength,
4
and a consistent test for over-identification is available (Chang et al., 2013). For dependent time series with double-penalty empirical likelihood, asymptotic normality is established for the nonzero coefficients but includes an explicit bias term; projected penalized empirical likelihood is then introduced for low-dimensional unbiased inference after orthogonalizing nuisance effects (Chang et al., 26 Feb 2025).
In testing problems, the limiting law is often no longer chi-square. For the complete independence test in high-dimensional data, the one-sided empirical likelihood statistic satisfies
5
and the rescaled statistic has the same limit while improving power in simulations (Qi et al., 2022). In the high-dimensional two-sample change-point linear model, the empirical likelihood ratio statistic, after normalization, is asymptotically standard normal rather than 6 (Ciuperca et al., 2015). For penalized empirical likelihood for the population mean when 7 may grow faster than 8, the limit distribution of the ratio statistic varies with the component-wise dependence structure and can be non-Gaussian, normal after centering and scaling, or a stochastic integral limit, depending on whether dependence is non-Ergodic, long-range, or short-range (Lahiri et al., 2013).
A common misconception is that empirical likelihood ratio statistics retain the usual chi-squared limit once weights are constructed. The high-dimensional literature repeatedly shows that this is false: the limiting law can be truncated chi-square-type, standard normal after normalization, dependence-specific, or Bartlett-corrected only after explicit adjustment (Qi et al., 2022, Ciuperca et al., 2015, Lahiri et al., 2013, Liu et al., 2010).
5. Major application domains
High-dimensional empirical likelihood weighting has been developed in several distinct but connected domains. The applications differ in what is being weighted—observations, sampled units, treatment groups, jackknife pseudo-values, or blockwise moment averages—but they share the use of moment restrictions and data-adaptive weights.
| Setting | Weighting device | Representative papers |
|---|---|---|
| High-dimensional longitudinal data | Double penalty on 9 and 0; robust estimating functions | (Li et al., 2021) |
| Big-data subsampling and M-estimation | Capture-recapture sampling with ELW weights instead of IPW | (Fan et al., 2022) |
| Voluntary samples | Bias calibration and benchmarking constraints | (Kim et al., 2022) |
| Randomized trials and ATE estimation | Group-specific EL weights; data splitting; multiple working models | (Tan et al., 2020, Liang et al., 2020, Xia et al., 30 Aug 2025) |
| Structural equation models | EL-weighted estimators with growing side-information constraints | (Wang et al., 2023) |
| Dependent time series and high-dimensional moments | GEL or doubly penalized EL under weak dependence or 1-mixing | (Chang et al., 2013, Chang et al., 26 Feb 2025) |
In big-data subsampling, the primary emphasis is computational cost and stability: empirical likelihood weighting is proposed because inverse probability weighting can become unstable when the probability weights are close to zero and cannot incorporate auxiliary information (Fan et al., 2022). In voluntary samples, the emphasis is nonignorable selection bias and the use of population auxiliary totals (Kim et al., 2022). In randomized experiments and observational treatment-effect problems, the emphasis is semiparametric efficiency, double robustness or multiple robustness, and valid inference with high-dimensional nuisance estimation (Tan et al., 2020, Liang et al., 2020, Xia et al., 30 Aug 2025). In structural equation models, the emphasis is efficiency gain from side information and from allowing the number of constraints to grow with the sample size (Wang et al., 2023). In dependent time series, the emphasis is sparse high-dimensional estimation under many moment restrictions without relying solely on classical blockwise approaches (Chang et al., 2013, Chang et al., 26 Feb 2025).
This distribution of applications suggests that high-dimensional empirical likelihood weighting is less a single algorithm than a methodological template adaptable to sampling design, causal inference, semiparametric efficiency, robustness, and dependent data.
6. Recurring difficulties, methodological debates, and current directions
Several recurring difficulties organize the field. One is feasibility. Standard empirical likelihood may be undefined because the convex hull does not contain the target moment vector, particularly when the dimension of the estimating function is high (Liu et al., 2010). Penalized empirical likelihood, adjusted empirical likelihood, and soft calibration are three different responses: penalization relaxes exact moment fitting through regularization, adjustment adds pseudo-observations to guarantee existence, and soft balancing replaces exact equality constraints by inequalities (Lahiri et al., 2013, Liu et al., 2010, Xia et al., 30 Aug 2025).
A second issue is whether empirical likelihood weighting should be viewed as a direct substitute for inverse probability weighting. In capture-recapture subsampling, the answer is explicitly negative: empirical likelihood weighting overcomes the instability of IPW by circumventing the use of inverse probabilities and by using auxiliary information (Fan et al., 2022). In treatment-effect estimation with missing outcomes, the comparison is more nuanced: empirical likelihood weighting can attain the semiparametric efficiency bound and can be doubly robust or multiply robust when constraints are built from working models, but this depends on correct specification conditions stated by the relevant papers (Tan et al., 2020, Xia et al., 30 Aug 2025).
A third debate concerns dependence. Earlier high-dimensional generalized empirical likelihood for dependent data uses a blocking technique to preserve dependence (Chang et al., 2013). More recent work on high-dimensional moment restrictions with dependent data emphasizes a marginal empirical likelihood approach despite temporal dependence in the data and derives theory under 2-mixing conditions (Chang et al., 26 Feb 2025). This suggests an active distinction between methods that encode dependence directly in the criterion and methods that retain an i.i.d.-looking criterion while moving dependence handling into the asymptotic analysis.
Finally, robustness remains a continuing direction. Robust estimating equations with bounded score functions, diagonal leverage-point downweighting, and depth-based weighting all aim to limit the impact of outliers or heavy tails (Li et al., 2021, Sang et al., 2019). At the same time, sample splitting, cross-fitting, and regularized nuisance estimation are used to stabilize inference when the number of covariates is high and machine-learning models are inserted into the constraints (Liang et al., 2020, Xia et al., 30 Aug 2025). A plausible implication is that contemporary high-dimensional empirical likelihood weighting is increasingly defined by the joint management of sparsity, calibration, and nuisance complexity rather than by likelihood maximization alone.