Penalized Empirical Likelihood (PEL)
- Penalized Empirical Likelihood (PEL) is a method that enhances classical empirical likelihood by incorporating explicit penalties for regularization, bias reduction, and moment selection.
- PEL methods use diverse penalty placements—on model parameters, Lagrange multipliers, or moment discrepancies—to address high-dimensional challenges and ensure the selection of informative moments.
- These techniques offer practical benefits in robust causal inference, capture–recapture abundance stabilization, Bayesian posterior computation, and bias correction across various applications.
Penalized empirical likelihood (PEL) denotes a family of methods that modifies empirical likelihood by adding an explicit penalty or, in one influential high-dimensional construction, by replacing exact moment feasibility with a quadratic penalty on weighted moment discrepancies. Across this literature, the common empirical-likelihood core is optimization over probability weights subject to estimating-equation structure, while the penalty is used for bias reduction, sparsity, moment selection, computational stabilization, or problem-specific regularization. The term does not refer to a single canonical estimator: different papers penalize the model parameter, the Lagrange multipliers, the discrepancy in empirical moments, or a target parameter such as abundance in capture–recapture models (Chang et al., 2017, Vexler et al., 2018, Lahiri et al., 2013, Liu et al., 2022).
1. Empirical-likelihood foundation and the scope of penalization
A standard empirical-likelihood formulation for estimating-equation models is
with a dual representation involving Lagrange multipliers through terms of the form (Chang et al., 2017, Chang et al., 2024). Penalized empirical likelihood retains this structure but augments it with a regularizing term. In some papers the penalty is added directly to the empirical-likelihood objective; in others it appears in the dual problem through the multiplier vector; in still others it regularizes a particular scientific target such as the population size (Vexler et al., 2018, Chang et al., 2024, Liu et al., 2022).
This breadth matters because the same label covers substantively different methodological aims. In high-dimensional estimating-equation problems, PEL is primarily a sparse estimation and moment-selection device (Chang et al., 2017, Chang et al., 2021). In bias-reduction work, it is the empirical-likelihood analogue of Jeffreys-prior penalization (Vexler et al., 2018). In capture–recapture, it is a one-sided stabilizer against implausibly large abundance estimates (Liu et al., 2022). In recent causal work, PEL is used on the treatment-weighting side to impose covariate-balancing propensity score restrictions with an penalty (Lee et al., 19 Sep 2025).
A further terminological complication is that the acronym “PEL” is not stable across subfields. One causal-inference paper explicitly uses “PEL” to mean pseudo-empirical likelihood, not penalized empirical likelihood, and contains no regularization penalty in the high-dimensional or shrinkage sense (Huang et al., 2024). By contrast, the empirical-likelihood-based consistent information criterion ELCIC is adjacent to PEL because it evaluates candidate penalized fits, but it is not a penalized empirical-likelihood estimator (Chen et al., 2020).
2. Main penalization architectures
The literature contains several recurring penalty placements. Their differences are structural rather than cosmetic.
| Penalization target | Representative criterion | Main role |
|---|---|---|
| Moment discrepancy | High-dimensional mean inference | |
| Model parameter | Bias reduction | |
| Parameters and multipliers | Variable and moment selection | |
| Multipliers only | Many-moment dimension reduction | |
| Target parameter 0 | 1 | Stabilization in capture–recapture |
| Propensity parameter 2 | 3 subject to EL constraints | CBPS-based causal weighting |
The 2013 high-dimensional mean paper is distinctive because it does not impose the exact empirical-likelihood moment constraint at all. Instead it replaces the hard feasibility condition by a variance-normalized quadratic penalty on the weighted discrepancy of the sample mean from 4. This makes the criterion well-defined even when 5, precisely the regime in which the convex-hull condition for ordinary empirical likelihood can fail with positive probability (Lahiri et al., 2013).
By contrast, the double-penalty estimating-equation literature keeps the usual EL dual form and adds penalties to 6, to 7, or to both. The representative formulation is
8
with 9 inducing sparsity in 0 and 1 inducing sparsity in the active estimating equations through 2 (Chang et al., 2017).
The bias-reduction paper uses a different logic. It derives an empirical-likelihood prior
3
and then defines the penalized estimator
4
Here the penalty is not a sparsity device; it is a Jeffreys-type information correction aimed at reducing the first-order asymptotic bias of the maximum empirical-likelihood estimator (Vexler et al., 2018).
3. High-dimensional sparse estimation and moment selection
The most developed PEL strand is the high-dimensional estimating-equation literature. Its central claim is that the main obstacle in empirical likelihood is often not only the parameter dimension 5, but also the moment dimension 6, because the dual variable 7 is 8-dimensional (Chang et al., 2017). This perspective led to the “new scope” in which a sparse 9 is estimated using a sparse subset of informative moments, rather than treating all 0 moments as equally active.
A key insight is that sparsity in 1 is equivalent to sparsity in the set of active estimating equations. In the 2017 double-penalty framework, for fixed 2,
3
and the inner optimizer 4 selects only a thresholded subset of moments. The support of 5 is shown, with high probability, to be contained in a set of coordinates with sufficiently large empirical discrepancies 6 (Chang et al., 2017). This provides a direct moment-selection interpretation of the multiplier penalty.
The theoretical payoff is substantial. Under suitable conditions, including 7, the estimator is sparse and consistent, with
8
and after a bias correction 9, the nonzero coordinates satisfy
0
for any unit vector 1 (Chang et al., 2017). The same paper emphasizes that both 2 and 3 may grow exponentially with 4 under the stated rates.
The 2021 paper “Culling the herd of moments with penalized empirical likelihood” recasts possible moment invalidity as a sparse nuisance-parameter problem. It introduces an auxiliary vector 5 with components
6
so that valid moments correspond to 7 and invalid moments to 8. The proposed PEL criterion penalizes both 9 and the dual variables 0 for candidate moments,
1
This yields consistent detection of invalid moments through
2
and an oracle-type asymptotic theory after bias correction (Chang et al., 2021).
The 2013 mean-inference paper occupies a different place within high-dimensional PEL. Its target is the 3-dimensional population mean, and the emphasis is not sparse selection but asymptotic behavior under component-wise dependence structures. The paper derives different limit laws for the PEL ratio statistic under non-Ergodic, long-range dependence, and short-range dependence, showing that the limit is typically not the classical fixed-dimensional 4 law. It also establishes a unified subsampling calibration valid across those dependence regimes (Lahiri et al., 2013).
4. Bias reduction, robustness, and Bayesian computation
Bias reduction enters PEL through the empirical-likelihood prior of Vexler, Zou, and Hutson. Starting from an integrated Kullback–Leibler or Shannon mutual information criterion,
5
the paper shows that the asymptotically optimal prior is
6
The resulting penalized empirical likelihood removes the 7-component of the 8 bias of the ordinary maximum empirical-likelihood estimator, and when 9, the first-order bias vanishes entirely (Vexler et al., 2018). This PEL variant is therefore closest in spirit to Firth-type penalization, not to sparse regularization.
Robustness is the defining feature of the high-dimensional longitudinal PEL developed for marginal models with repeated measurements. That paper replaces the ordinary QIF-style estimating equations by bounded robust estimating functions
0
and combines them with penalties on both 1 and 2: 3 The concrete variants ERPEL, HRPEL, and TRPEL correspond to exponential, Huber, and Tukey robust score functions. The paper argues that robustness in the estimating equations improves estimating-equation selection and therefore improves variable selection under heavy tails and contamination. Its formal robustness measure is a bounded influence function (Li et al., 2021).
Bayesian penalized empirical likelihood, or BPEL, shifts the focus from point estimation to posterior computation. The defining estimator penalizes the dual variables only,
4
and the corresponding posterior is
5
The penalty creates sparse multiplier vectors and hence automatic moment selection, while the Bayesian layer replaces difficult global optimization over 6 by sampling. The paper develops both Metropolis–Hastings and a modified adaptive multiple importance sampler (MAMIS), proves convergence of both schemes, and establishes a Bernstein–von Mises theorem: 7 The posterior is asymptotically Gaussian around the penalized estimator, not directly around 8, reflecting the bias induced by the penalty (Chang et al., 2024).
5. Domain-specific formulations and applications
One of the most problem-specific PEL constructions arises in closed-population capture–recapture. In that setting the instability concerns the abundance parameter 9, especially when the capture probability is low or the empirical-likelihood profile in 0 has a flat right tail. The proposed penalized log empirical likelihood is
1
with the recommended one-sided quadratic penalty
2
This shrinks overly large EL abundance estimates toward Chao’s lower bound without penalizing smaller values. The paper proves asymptotic normality of the maximum PEL estimator and Wilks-type limits for the PEL ratio statistic, and it develops standard EM algorithms for PEL, EL, and CL. In a difficult scenario B with 3 and 4, the reported RMSEs were
5
which the paper presents as direct evidence of stabilization (Liu et al., 2022).
A recent causal paper integrates generalized propensity score, covariate balancing propensity score, and outcome regression into a PEL formulation for continuous outcomes under multiple treatment levels. Its defining treatment-side optimization is
6
subject to
7
The paper states that 8 encodes CBPS and outcome-regression moment conditions, and it gives the explicit CBPS balance restriction
9
Its stated goals are covariate balance, high-dimensional regularization, and support for a doubly robust estimator (Lee et al., 19 Sep 2025).
That causal paper also illustrates a recurring limitation in applied PEL work: the conceptual framework can be broader than the displayed formulas. It claims accommodation of multi-valued treatments, yet the explicit propensity-score and CBPS equations are written in binary-treatment form. It also claims double robustness, finite-sample validity, and outlier resistance, but the supplied manuscript excerpt does not provide theorem statements for consistency of the PEL treatment-model estimator, oracle properties, or asymptotic normality (Lee et al., 19 Sep 2025).
6. Adjacent literature, misconceptions, and unresolved issues
A persistent misconception is that any empirical-likelihood procedure attached to penalized or sparse fits is automatically PEL. The 2020 empirical-likelihood-based consistent information criterion is an instructive counterexample. ELCIC is
0
and is designed for model selection among candidate models, including models generated by penalized procedures such as PGEE with SCAD. Its relevance to PEL is indirect: it uses empirical likelihood to score candidate penalized fits, but it does not formulate or analyze a penalized empirical-likelihood estimator (Chen et al., 2020).
A second misconception concerns acronym reuse. The 2024 causal-inference paper on average treatment effects is explicitly about pseudo-empirical likelihood. Its PEL objective uses estimated propensity-score weights and calibration constraints, and the resulting maximum PEL estimator is equivalent to a Hájek-type IPW estimator. There is no regularization penalty on parameters, multipliers, or moments, so it belongs to a different methodological lineage despite the shared acronym (Huang et al., 2024).
Several open issues recur across the genuine PEL literature. First, the field lacks a single canonical penalty geometry. Penalties have been placed on 1, on 2, on moment discrepancies, on nuisance misspecification vectors such as 3, and on scientific targets such as 4. This suggests that “PEL” is better understood as a design principle than as a single estimator class (Chang et al., 2017, Chang et al., 2021, Liu et al., 2022). Second, multiplier penalization typically introduces asymptotic bias, which then requires explicit correction, projected inference, or posterior centering around the penalized estimator rather than the truth (Chang et al., 2017, Chang et al., 2021, Chang et al., 2024). Third, tuning remains problem-dependent: the literature uses BIC-type criteria, theoretical rate conditions, or ad hoc simulation-scale choices, but no universal rule emerges from the cited papers (Chang et al., 2017, Chang et al., 2021, Chang et al., 2024).
The resulting picture is technically coherent but heterogeneous. Penalized empirical likelihood has evolved into a broad toolkit for semiparametric and estimating-equation problems: sparse estimation in high dimensions, selection among many moments, invalid-moment screening, bias reduction, robust longitudinal analysis, Bayesian sampling, abundance stabilization, and treatment-model regularization. Its unifying idea is not a particular penalty form, but the use of regularization to make empirical-likelihood methods workable outside the classical low-dimensional, well-specified regime (Lahiri et al., 2013, Chang et al., 2017, Vexler et al., 2018, Chang et al., 2024).