Papers
Topics
Authors
Recent
Search
2000 character limit reached

Penalized Empirical Likelihood

Updated 10 July 2026
  • Penalized empirical likelihood is a class of methods that integrates regularization with empirical likelihood to address high-dimensional challenges, outliers, and invalid moment conditions.
  • It applies various penalties on model parameters, Lagrange multipliers, and moment deviations to induce sparsity and enhance moment selection with consistent asymptotics.
  • Applications span causal inference, longitudinal analysis, econometrics, and capture–recapture, offering improved numerical stability and robust estimation in complex settings.

Penalized empirical likelihood (PEL) denotes a family of methods that augments empirical likelihood with explicit regularization on model parameters, Lagrange multipliers, weighted moment deviations, or application-specific targets. Across the current literature, PEL is used when classical empirical likelihood encounters high-dimensional estimating equations, convex-hull failure, invalid or weak moments, outliers and heavy tails, low-overlap causal designs, decentralized computation, or unstable abundance estimation. In that sense, PEL is not a single estimator but a class of regularized empirical-likelihood constructions built on estimating equations and nonparametric likelihood weights (Chang et al., 2017, Li et al., 2021, Chang et al., 2024).

1. Empirical-likelihood basis

Empirical likelihood (EL) starts from estimating equations E{g(Zi,θ)}=0E\{g(Z_i,\theta)\}=0. In the standard formulation, one assigns nonnegative weights pip_i to observations and maximizes the nonparametric likelihood under simplex and moment constraints: RF=supθ,p1,,pn{i=1nnpi; pi0, i=1npi=1, i=1npig(Zi,θ)=0}.R^F=\sup_{\theta,p_1,\ldots,p_n}\left\{\prod_{i=1}^{n} np_i;\ p_i\ge 0,\ \sum_{i=1}^{n}p_i=1,\ \sum_{i=1}^{n}p_i\,g(Z_i,\theta)=0\right\}. Using Lagrange multipliers, the maximizing weights take the familiar form

pi(θ,λ)=1n11+λg(Zi,θ),p_i(\theta,\lambda)=\frac{1}{n}\frac{1}{1+\lambda^\top g(Z_i,\theta)},

and the negative log empirical likelihood ratio can be written as

l=logRF(θ^,λ^)=i=1nlog{1+λ^g(Zi,θ^)}.l=-\log R^F(\hat\theta,\hat\lambda)=\sum_{i=1}^{n}\log\{1+\hat\lambda^\top g(Z_i,\hat\theta)\}.

The multiplier λ^\hat\lambda solves

1ni=1ng(Zi,θ^)1+λg(Zi,θ^)=0\frac{1}{n}\sum_{i=1}^{n}\frac{g(Z_i,\hat\theta)}{1+\lambda^\top g(Z_i,\hat\theta)}=0

(Chen et al., 2020).

This likelihood-like construction is attractive because it avoids specifying a full parametric distribution and works directly with moment restrictions. At fixed dimension, the classical EL ratio has Wilks-type chi-squared behavior under regularity, but the underlying positivity constraint 1+λgi(θ)>01+\lambda^\top g_i(\theta)>0 and the convex-hull geometry become problematic as the number of moments grows. One formulation in the recent literature states that classical EL can degenerate when qq grows with nn, and cites the case in which, when pip_i0, EL may be identically zero near pip_i1 (Chang et al., 2024). That difficulty is a primary motivation for penalization.

2. Penalization architectures

The literature uses several distinct penalization mechanisms rather than a single canonical PEL objective. A central high-dimensional formulation penalizes both the model parameter and the dual multiplier: pip_i2 Here pip_i3 and pip_i4 are sparsity-inducing penalties such as pip_i5, SCAD, or MCP. This construction is designed to encourage sparsity in pip_i6 and to select a sparse subset of informative estimating equations through pip_i7 (Chang et al., 2017).

A robust longitudinal variant keeps the dual-penalty structure but replaces the raw estimating equations with robustified versions based on bounded score functions and leverage downweights. In that setting, the criterion is

pip_i8

or, after positivity stabilization, the pseudo-log version pip_i9 with RF=supθ,p1,,pn{i=1nnpi; pi0, i=1npi=1, i=1npig(Zi,θ)=0}.R^F=\sup_{\theta,p_1,\ldots,p_n}\left\{\prod_{i=1}^{n} np_i;\ p_i\ge 0,\ \sum_{i=1}^{n}p_i=1,\ \sum_{i=1}^{n}p_i\,g(Z_i,\theta)=0\right\}.0 replacing RF=supθ,p1,,pn{i=1nnpi; pi0, i=1npi=1, i=1npig(Zi,θ)=0}.R^F=\sup_{\theta,p_1,\ldots,p_n}\left\{\prod_{i=1}^{n} np_i;\ p_i\ge 0,\ \sum_{i=1}^{n}p_i=1,\ \sum_{i=1}^{n}p_i\,g(Z_i,\theta)=0\right\}.1 (Li et al., 2021).

Other papers regularize different objects. For inference on a high-dimensional mean RF=supθ,p1,,pn{i=1nnpi; pi0, i=1npi=1, i=1npig(Zi,θ)=0}.R^F=\sup_{\theta,p_1,\ldots,p_n}\left\{\prod_{i=1}^{n} np_i;\ p_i\ge 0,\ \sum_{i=1}^{n}p_i=1,\ \sum_{i=1}^{n}p_i\,g(Z_i,\theta)=0\right\}.2, PEL is defined by

RF=supθ,p1,,pn{i=1nnpi; pi0, i=1npi=1, i=1npig(Zi,θ)=0}.R^F=\sup_{\theta,p_1,\ldots,p_n}\left\{\prod_{i=1}^{n} np_i;\ p_i\ge 0,\ \sum_{i=1}^{n}p_i=1,\ \sum_{i=1}^{n}p_i\,g(Z_i,\theta)=0\right\}.3

with RF=supθ,p1,,pn{i=1nnpi; pi0, i=1npi=1, i=1npig(Zi,θ)=0}.R^F=\sup_{\theta,p_1,\ldots,p_n}\left\{\prod_{i=1}^{n} np_i;\ p_i\ge 0,\ \sum_{i=1}^{n}p_i=1,\ \sum_{i=1}^{n}p_i\,g(Z_i,\theta)=0\right\}.4 and RF=supθ,p1,,pn{i=1nnpi; pi0, i=1npi=1, i=1npig(Zi,θ)=0}.R^F=\sup_{\theta,p_1,\ldots,p_n}\left\{\prod_{i=1}^{n} np_i;\ p_i\ge 0,\ \sum_{i=1}^{n}p_i=1,\ \sum_{i=1}^{n}p_i\,g(Z_i,\theta)=0\right\}.5. This formulation does not enforce RF=supθ,p1,,pn{i=1nnpi; pi0, i=1npi=1, i=1npig(Zi,θ)=0}.R^F=\sup_{\theta,p_1,\ldots,p_n}\left\{\prod_{i=1}^{n} np_i;\ p_i\ge 0,\ \sum_{i=1}^{n}p_i=1,\ \sum_{i=1}^{n}p_i\,g(Z_i,\theta)=0\right\}.6; it penalizes its violation directly (Lahiri et al., 2013).

A bias-reduction construction penalizes EL by an empirical-likelihood prior rather than by a sparsity penalty. In that case,

RF=supθ,p1,,pn{i=1nnpi; pi0, i=1npi=1, i=1npig(Zi,θ)=0}.R^F=\sup_{\theta,p_1,\ldots,p_n}\left\{\prod_{i=1}^{n} np_i;\ p_i\ge 0,\ \sum_{i=1}^{n}p_i=1,\ \sum_{i=1}^{n}p_i\,g(Z_i,\theta)=0\right\}.7

where RF=supθ,p1,,pn{i=1nnpi; pi0, i=1npi=1, i=1npig(Zi,θ)=0}.R^F=\sup_{\theta,p_1,\ldots,p_n}\left\{\prod_{i=1}^{n} np_i;\ p_i\ge 0,\ \sum_{i=1}^{n}p_i=1,\ \sum_{i=1}^{n}p_i\,g(Z_i,\theta)=0\right\}.8 is a Jeffreys-type EL prior derived by maximizing an EL-based mutual information criterion (Vexler et al., 2018).

Application-specific PEL objectives also appear. For multiple-treatment causal inference, the primal EL is penalized by an RF=supθ,p1,,pn{i=1nnpi; pi0, i=1npi=1, i=1npig(Zi,θ)=0}.R^F=\sup_{\theta,p_1,\ldots,p_n}\left\{\prod_{i=1}^{n} np_i;\ p_i\ge 0,\ \sum_{i=1}^{n}p_i=1,\ \sum_{i=1}^{n}p_i\,g(Z_i,\theta)=0\right\}.9 term on pi(θ,λ)=1n11+λg(Zi,θ),p_i(\theta,\lambda)=\frac{1}{n}\frac{1}{1+\lambda^\top g(Z_i,\theta)},0: pi(θ,λ)=1n11+λg(Zi,θ),p_i(\theta,\lambda)=\frac{1}{n}\frac{1}{1+\lambda^\top g(Z_i,\theta)},1 subject to empirical-likelihood moment restrictions that encode CBPS and outcome-regression conditions (Lee et al., 19 Sep 2025). For closed-population capture–recapture, the profile log EL is stabilized by adding pi(θ,λ)=1n11+λg(Zi,θ),p_i(\theta,\lambda)=\frac{1}{n}\frac{1}{1+\lambda^\top g(Z_i,\theta)},2, with

pi(θ,λ)=1n11+λg(Zi,θ),p_i(\theta,\lambda)=\frac{1}{n}\frac{1}{1+\lambda^\top g(Z_i,\theta)},3

to penalize large abundance values and overcome flat EL tails when the capture probability is low (Liu et al., 2022).

3. Sparsity, moment selection, and asymptotic structure

A defining feature of many PEL formulations is that the penalty on pi(θ,λ)=1n11+λg(Zi,θ),p_i(\theta,\lambda)=\frac{1}{n}\frac{1}{1+\lambda^\top g(Z_i,\theta)},4 induces sparsity in the active estimating equations. In the dual-penalty high-dimensional framework, this produces “drastic dimension reduction in the number of estimating equations” and makes moment selection an explicit part of estimation. The resulting estimator is proved to be sparse and consistent, with asymptotically normally distributed nonzero components after bias correction, even when the dimensions of both the parameter vector and the estimating equations grow exponentially with the sample size (Chang et al., 2017).

The robust longitudinal formulation sharpens this theme. Its KKT conditions take the form

pi(θ,λ)=1n11+λg(Zi,θ),p_i(\theta,\lambda)=\frac{1}{n}\frac{1}{1+\lambda^\top g(Z_i,\theta)},5

for multiplier coordinates and an analogous condition for pi(θ,λ)=1n11+λg(Zi,θ),p_i(\theta,\lambda)=\frac{1}{n}\frac{1}{1+\lambda^\top g(Z_i,\theta)},6. Because the robust scores, leverage weights, and pseudo-log derivatives are bounded, the influence functions for both pi(θ,λ)=1n11+λg(Zi,θ),p_i(\theta,\lambda)=\frac{1}{n}\frac{1}{1+\lambda^\top g(Z_i,\theta)},7 and pi(θ,λ)=1n11+λg(Zi,θ),p_i(\theta,\lambda)=\frac{1}{n}\frac{1}{1+\lambda^\top g(Z_i,\theta)},8 are bounded. Under the stated regularity conditions, the paper establishes a rate of convergence, selection consistency, and oracle properties for the nonzero coefficients (Li et al., 2021).

In econometric moment-selection problems, PEL is used to separate valid and invalid moments. One construction introduces an auxiliary parameter pi(θ,λ)=1n11+λg(Zi,θ),p_i(\theta,\lambda)=\frac{1}{n}\frac{1}{1+\lambda^\top g(Z_i,\theta)},9 for the candidate moment block and penalizes both the Lagrange multipliers and l=logRF(θ^,λ^)=i=1nlog{1+λ^g(Zi,θ^)}.l=-\log R^F(\hat\theta,\hat\lambda)=\sum_{i=1}^{n}\log\{1+\hat\lambda^\top g(Z_i,\hat\theta)\}.0. Valid moments are detected through zeros of l=logRF(θ^,λ^)=i=1nlog{1+λ^g(Zi,θ^)}.l=-\log R^F(\hat\theta,\hat\lambda)=\sum_{i=1}^{n}\log\{1+\hat\lambda^\top g(Z_i,\hat\theta)\}.1, with

l=logRF(θ^,λ^)=i=1nlog{1+λ^g(Zi,θ^)}.l=-\log R^F(\hat\theta,\hat\lambda)=\sum_{i=1}^{n}\log\{1+\hat\lambda^\top g(Z_i,\hat\theta)\}.2

and the paper proves l=logRF(θ^,λ^)=i=1nlog{1+λ^g(Zi,θ^)}.l=-\log R^F(\hat\theta,\hat\lambda)=\sum_{i=1}^{n}\log\{1+\hat\lambda^\top g(Z_i,\hat\theta)\}.3. It further develops a projected PEL procedure that removes the asymptotic bias induced by high-dimensional moment selection and yields bias-free asymptotic normality for low-dimensional targets (Chang et al., 2021).

PEL does not, however, imply a universal Wilks-type limit. In the high-dimensional mean problem, the limit law of l=logRF(θ^,λ^)=i=1nlog{1+λ^g(Zi,θ^)}.l=-\log R^F(\hat\theta,\hat\lambda)=\sum_{i=1}^{n}\log\{1+\hat\lambda^\top g(Z_i,\hat\theta)\}.4 depends on the component-wise dependence structure. Under non-ergodic dependence, the limit is a stochastic integral involving a Gaussian process; under short-range dependence, l=logRF(θ^,λ^)=i=1nlog{1+λ^g(Zi,θ^)}.l=-\log R^F(\hat\theta,\hat\lambda)=\sum_{i=1}^{n}\log\{1+\hat\lambda^\top g(Z_i,\hat\theta)\}.5 is asymptotically normal; under strong long-range dependence with l=logRF(θ^,λ^)=i=1nlog{1+λ^g(Zi,θ^)}.l=-\log R^F(\hat\theta,\hat\lambda)=\sum_{i=1}^{n}\log\{1+\hat\lambda^\top g(Z_i,\hat\theta)\}.6, the limit is non-Normal and is expressed as a bivariate Wiener–Itô integral (Lahiri et al., 2013). A common misconception is therefore that penalization merely restores standard EL asymptotics; in some high-dimensional regimes it changes the effective asymptotic geometry altogether.

4. Computation and numerical methods

Most PEL procedures are solved by alternating or nested optimization over the primal parameter and the dual multiplier. In the multiple-treatment causal framework, a standard implementation alternates between solving for the dual multiplier by minimizing the EL dual and updating the structural parameter through the penalized dual objective,

l=logRF(θ^,λ^)=i=1nlog{1+λ^g(Zi,θ^)}.l=-\log R^F(\hat\theta,\hat\lambda)=\sum_{i=1}^{n}\log\{1+\hat\lambda^\top g(Z_i,\hat\theta)\}.7

with coordinate descent for the l=logRF(θ^,λ^)=i=1nlog{1+λ^g(Zi,θ^)}.l=-\log R^F(\hat\theta,\hat\lambda)=\sum_{i=1}^{n}\log\{1+\hat\lambda^\top g(Z_i,\hat\theta)\}.8 penalty, second-order methods for smooth components, convergence checks on the KKT conditions, and tuning by cross-validation or information criteria. Warm starts are recommended for regularization paths (Lee et al., 19 Sep 2025).

High-dimensional sparse PEL and robust longitudinal PEL both use nested coordinate descent together with a pseudo-log transformation to preserve numerical stability near the EL feasibility boundary. A representative choice is

l=logRF(θ^,λ^)=i=1nlog{1+λ^g(Zi,θ^)}.l=-\log R^F(\hat\theta,\hat\lambda)=\sum_{i=1}^{n}\log\{1+\hat\lambda^\top g(Z_i,\hat\theta)\}.9

with λ^\hat\lambda0. This replacement yields a twice differentiable objective even when λ^\hat\lambda1 becomes small, and supports Newton or coordinate-wise updates with thresholding for exact sparsity (Li et al., 2021, Chang et al., 2017).

Decentralized network settings use a different computational strategy. There, the penalty is a fused lasso term on node-specific multipliers,

λ^\hat\lambda2

and inference is carried out by two ADMM-based algorithms: the Pairwise Copy Method and the Modified Approximation Objective Method. The second algorithm has linear convergence on spanning-tree network structures and closed-form node-level updates, while still recovering the global EL multiplier asymptotically (Du et al., 2024).

Capture–recapture PEL employs EM rather than direct dual optimization. The E-step computes expected counts for unobserved individuals, the M-step updates λ^\hat\lambda3 in closed form and updates λ^\hat\lambda4 through a weighted binomial logistic regression, and a separate step updates λ^\hat\lambda5 by maximizing

λ^\hat\lambda6

The penalized log EL is nondecreasing after each EM iteration, and the algorithm converges to a local maximum PEL estimator (Liu et al., 2022).

5. Major application domains

A recent causal-inference application integrates the generalized propensity score (GPS), covariate balancing propensity score (CBPS), and outcome regression into a PEL formulation for continuous outcomes under multiple treatment levels. The stacked moments encode CBPS balance and outcome-regression conditions, while the resulting doubly robust estimator

λ^\hat\lambda7

is consistent if either the propensity model or the outcome model is correctly specified. In simulations with λ^\hat\lambda8 covariates, sample sizes λ^\hat\lambda9, and contamination ratios 1ni=1ng(Zi,θ^)1+λg(Zi,θ^)=0\frac{1}{n}\sum_{i=1}^{n}\frac{g(Z_i,\hat\theta)}{1+\lambda^\top g(Z_i,\hat\theta)}=00, the proposed method had lower bias and error than high-dimensional A-learning, deep Q-learning, and generalized survival forests; for example, at 1ni=1ng(Zi,θ^)1+λg(Zi,θ^)=0\frac{1}{n}\sum_{i=1}^{n}\frac{g(Z_i,\hat\theta)}{1+\lambda^\top g(Z_i,\hat\theta)}=01 and 1ni=1ng(Zi,θ^)1+λg(Zi,θ^)=0\frac{1}{n}\sum_{i=1}^{n}\frac{g(Z_i,\hat\theta)}{1+\lambda^\top g(Z_i,\hat\theta)}=02, bias/MSE were 1ni=1ng(Zi,θ^)1+λg(Zi,θ^)=0\frac{1}{n}\sum_{i=1}^{n}\frac{g(Z_i,\hat\theta)}{1+\lambda^\top g(Z_i,\hat\theta)}=03 versus 1ni=1ng(Zi,θ^)1+λg(Zi,θ^)=0\frac{1}{n}\sum_{i=1}^{n}\frac{g(Z_i,\hat\theta)}{1+\lambda^\top g(Z_i,\hat\theta)}=04 for A-learning, and on the sepsis dataset with 1ni=1ng(Zi,θ^)1+λg(Zi,θ^)=0\frac{1}{n}\sum_{i=1}^{n}\frac{g(Z_i,\hat\theta)}{1+\lambda^\top g(Z_i,\hat\theta)}=05 and 1ni=1ng(Zi,θ^)1+λg(Zi,θ^)=0\frac{1}{n}\sum_{i=1}^{n}\frac{g(Z_i,\hat\theta)}{1+\lambda^\top g(Z_i,\hat\theta)}=06, the reported performance was 1ni=1ng(Zi,θ^)1+λg(Zi,θ^)=0\frac{1}{n}\sum_{i=1}^{n}\frac{g(Z_i,\hat\theta)}{1+\lambda^\top g(Z_i,\hat\theta)}=07 (Lee et al., 19 Sep 2025).

In high-dimensional longitudinal analysis, robust PEL is used with QIF-style moment blocks, bounded score functions such as Huber, exponential, and Tukey’s biweight, and leverage downweights based on robust Mahalanobis distance. Simulations with continuous and count outcomes show that robust methods such as ERPEL, HRPEL, and TRPEL yield smaller average estimation error and MSE, and higher correct-model frequencies, than non-robust NPEL and PEL under heavy tails and contamination. In the yeast cell-cycle data, robust RPEL methods selected transcription factors such as SWI4, SWI6, and MBP1, which the paper identifies as known regulators of G1 (Li et al., 2021).

In structural econometrics, PEL addresses many potentially invalid moment conditions by penalizing an auxiliary misspecification vector 1ni=1ng(Zi,θ^)1+λg(Zi,θ^)=0\frac{1}{n}\sum_{i=1}^{n}\frac{g(Z_i,\hat\theta)}{1+\lambda^\top g(Z_i,\hat\theta)}=08 and the multiplier block attached to the candidate moments. In simulations for high-dimensional linear IV and nonlinear dynamic-panel models, the projected PEL procedure produced coverage matching nominal levels well, while the ordinary de-biased PEL retained finite-sample bias and 2SLS intervals were substantially wider (Chang et al., 2021).

Closed-population capture–recapture offers a different application logic. There, PEL penalizes large abundance values because the EL ratio in 1ni=1ng(Zi,θ^)1+λg(Zi,θ^)=0\frac{1}{n}\sum_{i=1}^{n}\frac{g(Z_i,\hat\theta)}{1+\lambda^\top g(Z_i,\hat\theta)}=09 can be flat when the overall capture probability 1+λgi(θ)>01+\lambda^\top g_i(\theta)>00 is small. In simulations, the PEL estimator consistently had the smallest RMSE among the EL, CL, and VGAM competitors; in one scenario with 1+λgi(θ)>01+\lambda^\top g_i(\theta)>01 and 1+λgi(θ)>01+\lambda^\top g_i(\theta)>02, the reported RMSEs for 1+λgi(θ)>01+\lambda^\top g_i(\theta)>03 were 1+λgi(θ)>01+\lambda^\top g_i(\theta)>04, 1+λgi(θ)>01+\lambda^\top g_i(\theta)>05, 1+λgi(θ)>01+\lambda^\top g_i(\theta)>06, and 1+λgi(θ)>01+\lambda^\top g_i(\theta)>07, respectively. In the Fort Drum black bear data, PEL yielded materially more stable upper confidence bounds than EL and CL under the 1+λgi(θ)>01+\lambda^\top g_i(\theta)>08 model (Liu et al., 2022).

6. Relations, distinctions, and current directions

PEL sits inside a broader EL/GEL/GMM landscape. One recent causal paper states the relationship directly: EL is a likelihood-based alternative to GMM, GEL generalizes EL through alternative discrepancy measures such as exponential tilting and continuous updating, and PEL augments EL/GEL with penalties—most often on 1+λgi(θ)>01+\lambda^\top g_i(\theta)>09—to enable variable selection and stabilize estimation in HDLSS settings (Lee et al., 19 Sep 2025).

A useful distinction is between PEL proper and EL-based model-selection criteria. The empirical-likelihood consistent information criterion

qq0

uses EL as a data-driven likelihood component in a BIC-type criterion, but the paper is explicit that it “does not penalize the empirical likelihood objective itself.” Instead, penalization enters through the complexity term qq1 and through externally computed plug-in estimators such as SCAD-penalized GEE fits (Chen et al., 2020). This distinction matters because model selection with EL is not automatically equivalent to penalized EL estimation.

Bayesian PEL introduces another layer. In that framework, the profile penalized EL defines a posterior surrogate

qq2

and inference is performed by Metropolis–Hastings or Modified Adaptive Multiple Importance Sampling rather than by direct optimization. The paper proves a Bernstein–von Mises result in total variation around the penalized estimator qq3, with covariance qq4, and emphasizes that the posterior concentrates around qq5, not qq6, because of the penalty-induced bias (Chang et al., 2024).

Decentralized network EL shows that penalties need not target sparsity in coefficients at all. There, a fused penalty on local multipliers enforces consensus across graph edges, and the resulting distributed empirical log-likelihood ratio remains asymptotically qq7 under connectedness and suitable growth of qq8, even when the number of machines diverges (Du et al., 2024). This suggests that PEL is best understood as a design pattern for regularized moment-based likelihood inference rather than as a single estimator, penalty location, or asymptotic regime.

Open directions stated in the current literature include adaptive selection of correlation structures via qq9-sparsity, extension to ultrahigh-dimensional screening and robust composite EL, non-i.i.d. and privacy-constrained decentralized settings, and data-driven tuning of network or penalty parameters (Li et al., 2021, Du et al., 2024). The common theme is that penalization is being used not merely to shrink estimates, but to reshape the feasible moment structure of empirical likelihood itself.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Penalized Empirical Likelihood.