Penalized Empirical Likelihood
- Penalized empirical likelihood is a class of methods that integrates regularization with empirical likelihood to address high-dimensional challenges, outliers, and invalid moment conditions.
- It applies various penalties on model parameters, Lagrange multipliers, and moment deviations to induce sparsity and enhance moment selection with consistent asymptotics.
- Applications span causal inference, longitudinal analysis, econometrics, and capture–recapture, offering improved numerical stability and robust estimation in complex settings.
Penalized empirical likelihood (PEL) denotes a family of methods that augments empirical likelihood with explicit regularization on model parameters, Lagrange multipliers, weighted moment deviations, or application-specific targets. Across the current literature, PEL is used when classical empirical likelihood encounters high-dimensional estimating equations, convex-hull failure, invalid or weak moments, outliers and heavy tails, low-overlap causal designs, decentralized computation, or unstable abundance estimation. In that sense, PEL is not a single estimator but a class of regularized empirical-likelihood constructions built on estimating equations and nonparametric likelihood weights (Chang et al., 2017, Li et al., 2021, Chang et al., 2024).
1. Empirical-likelihood basis
Empirical likelihood (EL) starts from estimating equations . In the standard formulation, one assigns nonnegative weights to observations and maximizes the nonparametric likelihood under simplex and moment constraints: Using Lagrange multipliers, the maximizing weights take the familiar form
and the negative log empirical likelihood ratio can be written as
The multiplier solves
This likelihood-like construction is attractive because it avoids specifying a full parametric distribution and works directly with moment restrictions. At fixed dimension, the classical EL ratio has Wilks-type chi-squared behavior under regularity, but the underlying positivity constraint and the convex-hull geometry become problematic as the number of moments grows. One formulation in the recent literature states that classical EL can degenerate when grows with , and cites the case in which, when 0, EL may be identically zero near 1 (Chang et al., 2024). That difficulty is a primary motivation for penalization.
2. Penalization architectures
The literature uses several distinct penalization mechanisms rather than a single canonical PEL objective. A central high-dimensional formulation penalizes both the model parameter and the dual multiplier: 2 Here 3 and 4 are sparsity-inducing penalties such as 5, SCAD, or MCP. This construction is designed to encourage sparsity in 6 and to select a sparse subset of informative estimating equations through 7 (Chang et al., 2017).
A robust longitudinal variant keeps the dual-penalty structure but replaces the raw estimating equations with robustified versions based on bounded score functions and leverage downweights. In that setting, the criterion is
8
or, after positivity stabilization, the pseudo-log version 9 with 0 replacing 1 (Li et al., 2021).
Other papers regularize different objects. For inference on a high-dimensional mean 2, PEL is defined by
3
with 4 and 5. This formulation does not enforce 6; it penalizes its violation directly (Lahiri et al., 2013).
A bias-reduction construction penalizes EL by an empirical-likelihood prior rather than by a sparsity penalty. In that case,
7
where 8 is a Jeffreys-type EL prior derived by maximizing an EL-based mutual information criterion (Vexler et al., 2018).
Application-specific PEL objectives also appear. For multiple-treatment causal inference, the primal EL is penalized by an 9 term on 0: 1 subject to empirical-likelihood moment restrictions that encode CBPS and outcome-regression conditions (Lee et al., 19 Sep 2025). For closed-population capture–recapture, the profile log EL is stabilized by adding 2, with
3
to penalize large abundance values and overcome flat EL tails when the capture probability is low (Liu et al., 2022).
3. Sparsity, moment selection, and asymptotic structure
A defining feature of many PEL formulations is that the penalty on 4 induces sparsity in the active estimating equations. In the dual-penalty high-dimensional framework, this produces “drastic dimension reduction in the number of estimating equations” and makes moment selection an explicit part of estimation. The resulting estimator is proved to be sparse and consistent, with asymptotically normally distributed nonzero components after bias correction, even when the dimensions of both the parameter vector and the estimating equations grow exponentially with the sample size (Chang et al., 2017).
The robust longitudinal formulation sharpens this theme. Its KKT conditions take the form
5
for multiplier coordinates and an analogous condition for 6. Because the robust scores, leverage weights, and pseudo-log derivatives are bounded, the influence functions for both 7 and 8 are bounded. Under the stated regularity conditions, the paper establishes a rate of convergence, selection consistency, and oracle properties for the nonzero coefficients (Li et al., 2021).
In econometric moment-selection problems, PEL is used to separate valid and invalid moments. One construction introduces an auxiliary parameter 9 for the candidate moment block and penalizes both the Lagrange multipliers and 0. Valid moments are detected through zeros of 1, with
2
and the paper proves 3. It further develops a projected PEL procedure that removes the asymptotic bias induced by high-dimensional moment selection and yields bias-free asymptotic normality for low-dimensional targets (Chang et al., 2021).
PEL does not, however, imply a universal Wilks-type limit. In the high-dimensional mean problem, the limit law of 4 depends on the component-wise dependence structure. Under non-ergodic dependence, the limit is a stochastic integral involving a Gaussian process; under short-range dependence, 5 is asymptotically normal; under strong long-range dependence with 6, the limit is non-Normal and is expressed as a bivariate Wiener–Itô integral (Lahiri et al., 2013). A common misconception is therefore that penalization merely restores standard EL asymptotics; in some high-dimensional regimes it changes the effective asymptotic geometry altogether.
4. Computation and numerical methods
Most PEL procedures are solved by alternating or nested optimization over the primal parameter and the dual multiplier. In the multiple-treatment causal framework, a standard implementation alternates between solving for the dual multiplier by minimizing the EL dual and updating the structural parameter through the penalized dual objective,
7
with coordinate descent for the 8 penalty, second-order methods for smooth components, convergence checks on the KKT conditions, and tuning by cross-validation or information criteria. Warm starts are recommended for regularization paths (Lee et al., 19 Sep 2025).
High-dimensional sparse PEL and robust longitudinal PEL both use nested coordinate descent together with a pseudo-log transformation to preserve numerical stability near the EL feasibility boundary. A representative choice is
9
with 0. This replacement yields a twice differentiable objective even when 1 becomes small, and supports Newton or coordinate-wise updates with thresholding for exact sparsity (Li et al., 2021, Chang et al., 2017).
Decentralized network settings use a different computational strategy. There, the penalty is a fused lasso term on node-specific multipliers,
2
and inference is carried out by two ADMM-based algorithms: the Pairwise Copy Method and the Modified Approximation Objective Method. The second algorithm has linear convergence on spanning-tree network structures and closed-form node-level updates, while still recovering the global EL multiplier asymptotically (Du et al., 2024).
Capture–recapture PEL employs EM rather than direct dual optimization. The E-step computes expected counts for unobserved individuals, the M-step updates 3 in closed form and updates 4 through a weighted binomial logistic regression, and a separate step updates 5 by maximizing
6
The penalized log EL is nondecreasing after each EM iteration, and the algorithm converges to a local maximum PEL estimator (Liu et al., 2022).
5. Major application domains
A recent causal-inference application integrates the generalized propensity score (GPS), covariate balancing propensity score (CBPS), and outcome regression into a PEL formulation for continuous outcomes under multiple treatment levels. The stacked moments encode CBPS balance and outcome-regression conditions, while the resulting doubly robust estimator
7
is consistent if either the propensity model or the outcome model is correctly specified. In simulations with 8 covariates, sample sizes 9, and contamination ratios 0, the proposed method had lower bias and error than high-dimensional A-learning, deep Q-learning, and generalized survival forests; for example, at 1 and 2, bias/MSE were 3 versus 4 for A-learning, and on the sepsis dataset with 5 and 6, the reported performance was 7 (Lee et al., 19 Sep 2025).
In high-dimensional longitudinal analysis, robust PEL is used with QIF-style moment blocks, bounded score functions such as Huber, exponential, and Tukey’s biweight, and leverage downweights based on robust Mahalanobis distance. Simulations with continuous and count outcomes show that robust methods such as ERPEL, HRPEL, and TRPEL yield smaller average estimation error and MSE, and higher correct-model frequencies, than non-robust NPEL and PEL under heavy tails and contamination. In the yeast cell-cycle data, robust RPEL methods selected transcription factors such as SWI4, SWI6, and MBP1, which the paper identifies as known regulators of G1 (Li et al., 2021).
In structural econometrics, PEL addresses many potentially invalid moment conditions by penalizing an auxiliary misspecification vector 8 and the multiplier block attached to the candidate moments. In simulations for high-dimensional linear IV and nonlinear dynamic-panel models, the projected PEL procedure produced coverage matching nominal levels well, while the ordinary de-biased PEL retained finite-sample bias and 2SLS intervals were substantially wider (Chang et al., 2021).
Closed-population capture–recapture offers a different application logic. There, PEL penalizes large abundance values because the EL ratio in 9 can be flat when the overall capture probability 0 is small. In simulations, the PEL estimator consistently had the smallest RMSE among the EL, CL, and VGAM competitors; in one scenario with 1 and 2, the reported RMSEs for 3 were 4, 5, 6, and 7, respectively. In the Fort Drum black bear data, PEL yielded materially more stable upper confidence bounds than EL and CL under the 8 model (Liu et al., 2022).
6. Relations, distinctions, and current directions
PEL sits inside a broader EL/GEL/GMM landscape. One recent causal paper states the relationship directly: EL is a likelihood-based alternative to GMM, GEL generalizes EL through alternative discrepancy measures such as exponential tilting and continuous updating, and PEL augments EL/GEL with penalties—most often on 9—to enable variable selection and stabilize estimation in HDLSS settings (Lee et al., 19 Sep 2025).
A useful distinction is between PEL proper and EL-based model-selection criteria. The empirical-likelihood consistent information criterion
0
uses EL as a data-driven likelihood component in a BIC-type criterion, but the paper is explicit that it “does not penalize the empirical likelihood objective itself.” Instead, penalization enters through the complexity term 1 and through externally computed plug-in estimators such as SCAD-penalized GEE fits (Chen et al., 2020). This distinction matters because model selection with EL is not automatically equivalent to penalized EL estimation.
Bayesian PEL introduces another layer. In that framework, the profile penalized EL defines a posterior surrogate
2
and inference is performed by Metropolis–Hastings or Modified Adaptive Multiple Importance Sampling rather than by direct optimization. The paper proves a Bernstein–von Mises result in total variation around the penalized estimator 3, with covariance 4, and emphasizes that the posterior concentrates around 5, not 6, because of the penalty-induced bias (Chang et al., 2024).
Decentralized network EL shows that penalties need not target sparsity in coefficients at all. There, a fused penalty on local multipliers enforces consensus across graph edges, and the resulting distributed empirical log-likelihood ratio remains asymptotically 7 under connectedness and suitable growth of 8, even when the number of machines diverges (Du et al., 2024). This suggests that PEL is best understood as a design pattern for regularized moment-based likelihood inference rather than as a single estimator, penalty location, or asymptotic regime.
Open directions stated in the current literature include adaptive selection of correlation structures via 9-sparsity, extension to ultrahigh-dimensional screening and robust composite EL, non-i.i.d. and privacy-constrained decentralized settings, and data-driven tuning of network or penalty parameters (Li et al., 2021, Du et al., 2024). The common theme is that penalization is being used not merely to shrink estimates, but to reshape the feasible moment structure of empirical likelihood itself.