---
title: Adaptive Empirical Likelihood Procedures
url: https://www.emergentmind.com/topics/adaptive-empirical-likelihood-procedure
type: topic
---

# Adaptive Empirical Likelihood Procedures

Searching arXiv for recent papers on adaptive/adjusted/penalized empirical likelihood and related variants.
Adaptive empirical likelihood procedure is not a single universally standardized method, but a family of empirical-likelihood-based constructions in which the likelihood surrogate, the estimating equations, the calibration device, or the computational representation is made data-driven to accommodate difficult inferential regimes. In the literature represented here, the phrase is most naturally associated with procedures that modify ordinary empirical likelihood to restore feasibility, improve finite-sample calibration, select informative estimating equations, handle heavy tails or dependence, or adapt inference to unknown asymptotic regimes [1603.04093], [1604.06170], [1704.00566], [2103.10613], [1205.6055]. A crucial distinction runs through this literature: some methods are explicitly **adjusted empirical likelihood** procedures, some are **penalized** or **split-sample** EL procedures, and some are only *adaptive in an informal sense* because they are data-driven or automatically calibrated rather than “adaptive” in a formal decision-theoretic sense [1904.08420], [2112.09206], [2403.05080].

## 1. Conceptual scope and terminological boundaries

Empirical likelihood is a nonparametric likelihood construction based on estimating equations rather than a fully specified parametric density. In the formulations summarized here, the common starting point is a moment or estimating-equation restriction such as \(E\{g(X,\theta_0)\}=0\), followed by maximization over probability weights subject to those constraints [1704.00566], [2303.07259], [1808.06222]. The attraction of this framework is that it often preserves likelihood-ratio-style inference and Wilks-type chi-square calibration without requiring direct covariance estimation [1604.06170], [2303.07259].

Within this broad framework, “adaptive empirical likelihood procedure” is used most coherently as an umbrella description for methods that alter the empirical likelihood construction in response to the statistical or computational structure of the problem. The data block supports several distinct meanings.

One meaning is **finite-sample or geometric adaptation**. Adjusted empirical likelihood adds a pseudo-observation so that the empirical likelihood ratio is always well-defined and the convex-hull feasibility problem is repaired [1904.08420], [1603.04093], [1604.06170], [1602.09128]. Another is **estimating-equation adaptation**, where penalization of the empirical-likelihood Lagrange multiplier induces selection among many candidate estimating equations, resulting in a data-driven subset of active constraints [1704.00566], [2103.10613]. A third is **computational adaptation**, in which the empirical likelihood is rebuilt on blockwise summaries rather than the full sample to address massive data while preserving first-order efficiency and Wilks’ theorem [1703.03312], [2303.07259]. A fourth is **automatic calibration**, where empirical-likelihood statistics are calibrated by Monte Carlo or bootstrap rather than by fixed parametric covariance formulas, especially for simultaneous inference [2112.09206].

The literature also marks clear boundaries. "Likelihood based inference for current status data on a grid: A boundary phenomenon and an adaptive inference procedure" is explicitly adaptive and likelihood-based, but it is **not an empirical likelihood method in the standard Owen sense** [1205.6055]. Conversely, several papers are not named “adaptive empirical likelihood,” yet are adaptive in an informal but technically substantive sense because they are nonparametric, data-driven, or automatically calibrated [2112.09206], [2403.05080], [2011.07721].

## 2. Feasibility repair and finite-sample adjustment

A large part of the adaptive-EL literature is driven by a basic geometric problem: empirical likelihood may be undefined when the origin is not in the convex hull of the estimating functions. In ordinary EL or JEL, the constraints are feasible only if the estimating equation can be solved with nonnegative weights summing to one [1603.04093], [1604.06170], [1602.09128]. This difficulty is amplified in small samples, in tail problems where only a small number of extremes are used, and in time-series settings where frequency-domain estimating functions can be poorly behaved [1904.08420], [1604.06170], [1602.09128].

The standard adjustment is to add one **pseudo-observation** proportional to the negative empirical mean of the estimating functions. In the heavy-tail tail-index paper, this takes the form
\[
y_{k_n+1}(\gamma) = -\frac{a_n}{k_n}\sum_{i=1}^{k_n} y_i(\gamma) = -a_n(\hat{\gamma}_n-\gamma),
\]
with \(a_n>0\), leading to the adjusted empirical likelihood ratio
\[
l_{\mathrm{AEL}(\gamma)} = -2\log\Bigg( \sup\Big\{ \prod_{i=1}^{k_n+1}((k_n+1)p_i): p_i\ge 0,\ \sum_{i=1}^{k_n+1}p_i=1,\ \sum_{i=1}^{k_n+1}p_i y_i(\gamma)=0 \Big\} \Bigg)
\]
and
\[
l_{\mathrm{AEL}(\gamma)} = 2\sum_{i=1}^{k_n+1}\log\bigl(1+\lambda y_i(\gamma)\bigr)
\]
with \(\lambda\) solving the corresponding constraint equation [1904.08420]. The paper states explicitly that one can verify there always exists a probability vector satisfying the constraint, so \(l_{\mathrm{AEL}(\gamma)}\) is always well-defined [1904.08420].

The same repair appears in jackknife EL. "Adjusted Jackknife Empirical Likelihood" constructs jackknife pseudo-values \(\widehat V_i\), then adds
\[
\widehat V_{n+1} = -a_n U_n = -\frac{a_n}{n}\sum_{i=1}^n \widehat V_i
\]
and defines the adjusted likelihood ratio on \(n+1\) pseudo-observations [1603.04093]. The paper’s main theorem is that, for one-sample and two-sample U-statistics,
\[
-2\log R(\theta)\xrightarrow{d}\chi_1^2,
\]
so the adjustment preserves the Wilks-type first-order limit while guaranteeing that the statistic is well-defined for all parameter values [1603.04093].

Time-series AEL papers apply the same idea in the frequency domain. For short-memory models, "Adjusted Empirical Likelihood for Time Series Models" augments Monti’s Whittle-based empirical likelihood by
\[
\psi_{n+1}(\beta) = -\frac{a_n}{n}\sum_{j=1}^n \psi_j(\beta) = -a_n\bar\psi_n(\beta),
\]
with
\[
a_n=\max\left(1,\frac{\log n}{2}\right),
\]
and proves that the adjusted statistic still has asymptotic \(\chi_k^2\) calibration [1602.09128]. "Adjusted Empirical Likelihood for Long-memory Time Series Models" applies the same construction to ARFIMA models and proves
\[
W^*(\beta)\xrightarrow{d}\chi_k^2
\]
under the Fox and Taqqu conditions [1604.06170].

A common misconception is that these procedures are “adaptive” because they change the target parameter or estimating equations. The papers do not support that interpretation. Their contribution is more specific: they adapt the empirical likelihood geometry so that the optimization problem exists and finite-sample coverage improves, while leaving first-order asymptotics unchanged [1904.08420], [1603.04093], [1604.06170], [1602.09128].

## 3. Adaptive selection of estimating equations and variables

The most explicit adaptive empirical-likelihood framework in the supplied material is the doubly penalized high-dimensional EL literature. "A new scope of penalized empirical likelihood with high-dimensional estimating equations" studies i.i.d. data \(X_1,\ldots,X_n\), a \(p\)-dimensional parameter \(\theta\), and an \(r\)-dimensional estimating function \(g(X;\theta)\) satisfying
\[
\mathbb{E}\{g(X_i;\theta_0)\}=0. \tag{2.1}
\]
The paper’s main methodological proposal is
\[
\hat\theta_n = \arg\min_{\theta\in\Theta}\max_{\lambda\in\widehat\Lambda_n(\theta)} \left[ \sum_{i=1}^n \log\{1+\lambda^\top g(X_i;\theta)\} - n\sum_{j=1}^r P_{2,\nu}(|\lambda_j|) + n\sum_{k=1}^p P_{1,\pi}(|\theta_k|) \right], \tag{3.1}
\]
where \(P_{1,\pi}\) penalizes \(\theta\) and \(P_{2,\nu}\) penalizes \(\lambda\) [1704.00566].

The adaptive feature is not incidental. Penalizing \(\lambda\) induces sparsity in the EL Lagrange multiplier, which the paper identifies as equivalent to selection among high-dimensional estimating equations [1704.00566]. For fixed \(\theta\), if
\[
\bar g_j(\theta)=\frac1n\sum_{i=1}^n g_j(X_i;\theta),
\]
then the active set of estimating equations is controlled by thresholding the sample moment magnitudes relative to \(\nu p_2'(0+)\) [1704.00566]. This is a genuine data-driven estimating-equation selection mechanism inside EL.

The asymptotic theory is correspondingly adaptive to sparsity. With active parameter set
\[
\mathcal S=\{1\le k\le p:\theta_k^0\neq 0\},\qquad s=|\mathcal S|,
\]
Theorem 1 establishes a sparse local minimizer \(\hat\theta_n\) such that
\[
|\hat\theta_{n,\mathcal S}-\theta_{0,\mathcal S}|_\infty = O_p\{b_n^{1/(2\beta)}\}, \qquad P(\hat\theta_{n,\mathcal S^c}=0)\to 1,
\]
and Theorem 2 gives asymptotic normality for the nonzero coordinates after bias correction [1704.00566]. The paper emphasizes that this permits both parameter dimension and number of estimating equations to grow exponentially with sample size under sparsity and effective-selection conditions [1704.00566].

"Robust penalized empirical likelihood in high dimensional longitudinal data analysis" extends the same architecture to longitudinal marginal models, but replaces the standard QIF-type estimating functions by robust estimating equations
\[
g(X_i;\beta)= \begin{pmatrix} D_i^T A_i^{-1/2} M_1 h_i(\mu_i(\beta))\ \vdots\ D_i^T A_i^{-1/2} M_\ell h_i(\mu_i(\beta)) \end{pmatrix},
\]
with
\[
h_i(\mu_i)=W_i\,[\psi(y_i-\mu_i(\beta))-C_i(\mu_i(\beta))].
\]
Its penalized criterion is
\[
S_n(\beta) = \sum_{i=1}^n \log\{1+\lambda^T g(X_i;\beta)\} - n\sum_{j=1}^r P_{1,\nu}(|\lambda_j|) + n\sum_{k=1}^p P_{2,\omega}(|\beta_k|). \tag{1}
\]
This again performs simultaneous selection of variables and estimating equations, but now with bounded influence functions for both \(\lambda\) and \(\beta\), due to bounded \(\psi\) and leverage weights \(W_i\) [2103.10613].

This suggests that in the high-dimensional literature, “adaptive empirical likelihood procedure” is best understood not as a general slogan but as a concrete min-max mechanism: sparsity in \(\beta\) adapts to the active model, and sparsity in \(\lambda\) adapts to the informative subset of estimating equations [1704.00566], [2103.10613].

## 4. Computational adaptation: split-sample, composite, and massive-data EL

A separate line of work adapts empirical likelihood to computational scale rather than to sparsity or robustness. "Split Sample Empirical Likelihood" proposes “a new approach that combines multiple non-parametric likelihood-type components to build a data-driven approximation of the true likelihood function” and shows that “the asymptotic behaviors of our approach are identical to those seen in empirical likelihood” while “significantly decreasing computational time” [1703.03312]. The supplied details do not provide the exact paper text, so no more specific construction can be attributed reliably [1703.03312].

A more concrete massive-data construction is given in "A novel approach of empirical likelihood with massive data," which introduces **split sample mean empirical likelihood (SSMEL)** [2303.07259]. The method partitions the sample into \(K\) subsets of size \(m=n/K\), computes one mean estimating function per block,
\[
\bar g^{(k)}(\theta)=\frac{1}{m}\sum_{i=1}^{m} g(x_i^{(k)},\theta),
\]
and then applies EL to the \(K\) block means:
\[
R_S(\theta)=\sup\left\{\prod_{k=1}^{K} Kp_k:\; p_k\ge 0,\ \sum_{k=1}^{K} p_k=1,\ \sum_{k=1}^{K} p_k \bar g^{(k)}(\theta)=0\right\}. \tag{4}
\]
The key identity is
\[
\frac{1}{K}\sum_{k=1}^{K}\bar g^{(k)}(\theta)=\frac{1}{n}\sum_{i=1}^{n}g(x_i,\theta),
\]
which preserves the full-sample mean estimating equation exactly [2303.07259]. The paper proves
\[
\sqrt n (\hat\theta_S-\theta_0)\overset d\longrightarrow N(0,\Sigma), \tag{7}
\]
and
\[
\mathcal{W}(\theta_0)=2\big[\ell_S(\theta_0)-\ell_S(\hat\theta_S)\big]\overset d\longrightarrow \chi_p^2, \tag{11}
\]
so SSMEL has the same first-order efficiency as centralized EL and retains Wilks’ theorem [2303.07259].

This is adaptive in a computational sense. The user chooses \(K\) to trade optimization cost against information retention. The paper recommends \(K>p\) and notes that if \(K\) is too small, the method over-compresses the information; if \(K\) is too close to \(p\), the convex-hull geometry becomes weak [2303.07259]. This suggests a notion of adaptive EL in which the empirical support is changed from individual observations to block means while preserving the inferential target.

The idea of combining multiple components also appears in "Composite Empirical Likelihood," whose abstract states that the method “combines multiple non-parametric likelihood-type components to build a data-driven approximation of the true function” by borrowing “the ability to avoid a parametric specification” from empirical likelihood and “multiple likelihood components” from composite likelihood [1511.04635]. The supplied details explicitly say that the paper text was unavailable and that no verified formulas or theorems could be extracted, so only this abstract-level description can be used responsibly [1511.04635]. Even at that level, it supports the view that adaptive EL can also mean modular combination of multiple likelihood-type pieces [1511.04635].

## 5. Automatic calibration, multiple testing, and adaptive likelihood-based inference

The experimental-design paper develops an empirical-likelihood framework that is not called adaptive in a formal sense, but it is strongly data-driven in its calibration [2112.09206]. In blocked designs, the estimating function is
\[
g(X_i, \theta) = (X_i - \theta) \circ c_i,
\]
and the EL statistic is built from
\[
l_n(\theta) = 2\sum_{i = 1}^n \log \left(1 + \lambda^\top g(X_i, \theta)\right),
\]
with hypotheses expressed as smooth constraints \(H_j=\{\theta:h_j(\theta)=0\}\) and test statistics
\[
T_{nj} = \frac{a_n^2}{n} \inf_{\theta \in H_j \cap \overline{K}_n} l_n(\theta)
\]
[2112.09206].

Its main adaptive feature is calibration. The paper proposes two single-step multiple testing procedures: **asymptotic Monte Carlo (AMC)** and **nonparametric bootstrap (NB)** [2112.09206]. AMC estimates the joint covariance structure from the data and simulates the limiting quadratic forms. NB resamples from null-transformed data and recomputes the entire vector of EL statistics [2112.09206]. Both procedures asymptotically control the generalized family-wise error rate and construct simultaneous confidence intervals “without explicitly considering the underlying covariance structure” [2112.09206]. The paper explicitly interprets this as data-driven and robust to violations of standard mixed-model assumptions [2112.09206].

A related but conceptually distinct example is "Likelihood based inference for current status data on a grid: A boundary phenomenon and an adaptive inference procedure" [1205.6055]. This paper studies the current status NPMLE on a grid with spacing \(cn^{-\gamma}\), where asymptotics change depending on whether \(\gamma<1/3\), \(\gamma=1/3\), or \(\gamma>1/3\) [1205.6055]. Its adaptive procedure works by pretending \(\gamma=1/3\), computing a surrogate \(\hat c\), and using the boundary family of limit laws to construct confidence intervals that are asymptotically valid without knowing \(\gamma\) [1205.6055]. The paper is explicit that this is **not an empirical likelihood method in the standard Owen sense** [1205.6055]. Its relevance is conceptual: it shows that adaptation can also mean avoiding regime estimation by exploiting a boundary distribution that interpolates between competing asymptotic laws [1205.6055].

This suggests a broader perspective. In empirical-likelihood-related work, “adaptive procedure” often refers less to a particular primal-dual optimization change than to an inference scheme whose calibration is learned from the data or from a unifying asymptotic boundary law [2112.09206], [1205.6055].

## 6. Simulation-based and Bayesian empirical likelihood as adaptive surrogates

Several papers use empirical likelihood as a likelihood surrogate in Bayesian or likelihood-free inference. Their adaptivity is mostly simulation-driven and summary-driven rather than formal in the penalized-EL sense.

"Bayesian computation via empirical likelihood" defines
\[
L_{el}(\theta\mid y) = \max_{\mathbf p}\; \prod_{i=1}^n p_i,
\]
subject to
\[
p_i\in[0,1],\qquad \sum_{i=1}^n p_i=1,\qquad \sum_{i=1}^n p_i\,h(y_i,\theta)=0,
\]
and then treats
\[
\pi_{BC}(\theta\mid y)\propto \pi(\theta)\,L_{el}(\theta\mid y)
\]
as an approximate posterior [1205.5658]. The basic sampler draws \(\theta_i\sim \pi\) and weights by \(\omega_i=L_{el}(\theta_i\mid y)\), while BC-AMIS adaptively updates the proposal distribution using weighted means and covariances from earlier particles [1205.5658]. This is an explicitly adaptive Monte Carlo mechanism built around EL, even though the EL constraints themselves are not adaptively chosen [1205.5658].

The ABC-EL papers move further toward simulation-based adaptation. "An easy-to-use empirical likelihood ABC method" and "On a Variational Approximation based Empirical Likelihood ABC Method" define, for observed summary \(g(X_o)\) and simulated summaries \(g(X_i(\theta))\),
\[
h_i(\theta)=g(X_i(\theta))-g(X_o),
\]
and the feasible set
\[
\mathcal{W}_{\theta} = \left\{w:\sum_{i=1}^m w_i h_i(\theta)=0\right\}\cap \Delta_{m-1},
\]
with empirical-likelihood weights
\[
\hat w(\theta)=\argmax_{w\in\mathcal W_\theta}\prod_{i=1}^m m w_i
\]
[1810.01675], [2011.07721]. The resulting pseudo-posterior is
\[
\Pi(\theta\mid X_o)\propto \exp\left(\frac{1}{m}\sum_{i=1}^m \log \hat w_i(\theta)\right)\pi(\theta)
\]
in the 2018 paper [1810.01675], and, in the 2020 variational version,
\[
\hat{\Pi}(\theta\mid g(X_o)) \propto \exp\left( \frac{1}{m}\sum_{i=1}^m \log \hat w_i(\theta) + \hat H^0_{g\mid\theta}(\theta) \right)\pi(\theta)
\]
after adding an estimated entropy term [2011.07721].

"On an Empirical Likelihood based Solution to the Approximate Bayesian Computation Problem" formalizes the same idea as ABCel, defining
\[
h_i(\theta)=s\!\left(X_i(\theta)\right)-s\!\left(X_o(\theta_o)\right), \tag{1}
\]
\[
\mathcal{W}_{\theta} =\left\{w:\sum_{i=1}^m w_i\left[s\!\left(X_i(\theta)\right)-s\!\left(X_o(\theta_o)\right)\right]=0\right\}\cap \Delta_{m-1}, \tag{2}
\]
\[
\hat w(\theta)=\argmax_{w\in\mathcal{W}_{\theta}} \prod_{i=1}^m m w_i, \tag{3}
\]
and posterior approximation
\[
\hat{\Pi}(\theta\mid s(X_o)) = \frac{ \left[e^{\left(\frac{1}{m}\sum_{i=1}^m \log \hat w_i(\theta)+\hat H^0_{s\mid\theta}(\theta)\right)}\right]\pi(\theta) }{ \int_{t\in\Theta} \left[e^{\left(\frac{1}{m}\sum_{i=1}^m \log \hat w_i(t)+\hat H^0_{s\mid t}(t)\right)}\right]\pi(t)\,dt }. \tag{4}
\]
The paper explicitly says it does **not** call the method “adaptive empirical likelihood,” but it is data-dependent through the observed-summary constraints, parameter-specific reweighting, and entropy estimation [2403.05080].

A plausible implication is that in the Bayesian and ABC literature, adaptive EL means the likelihood surrogate is rebuilt at each \(\theta\) from simulation output and observed summaries, rather than from a fixed analytic estimating equation. The papers do not claim more than that, and they also emphasize limitations: summary choice remains problem-dependent, feasibility can fail, and posterior support may be non-convex [1810.01675], [2011.07721], [2403.05080].

## 7. Scope, controversies, and recurring limitations

Several recurring limitations appear across the literature.

First, **adaptivity is not uniform across papers**. Some methods are explicitly adjusted EL procedures and should be described that way, not as general adaptive EL frameworks [1904.08420], [1603.04093], [1604.06170], [1602.09128]. Some are explicitly penalized EL procedures with adaptive moment selection [1704.00566], [2103.10613]. Others are adaptive only in an informal sense because they are data-driven, automatically calibrated, or simulation-based [2112.09206], [2403.05080], [2011.07721].

Second, **constraint choice remains central**. In Bayesian computation via EL, the success of BCel depends heavily on the identifying quality of the estimating equations, and the normal example in the paper shows that adding more constraints can worsen posterior approximation [1205.5658]. The easy-to-use EL-ABC paper likewise emphasizes that the method cannot automatically down-weight uninformative summaries, so poor summaries can cause infeasibility or underestimation of uncertainty [1810.01675].

Third, **convex-hull feasibility remains a structural issue**. Adjusted EL papers exist precisely because unadjusted EL may have no solution when the zero vector is outside the convex hull of the estimating functions [1904.08420], [1603.04093], [1604.06170], [1602.09128]. ABC-EL methods inherit related support irregularities because the observed summary may lie outside the convex hull of simulated summaries at a candidate parameter value [1810.01675], [2403.05080].

Fourth, **different adaptive mechanisms target different problems**. Robust doubly penalized EL targets outliers, heavy tails, and high-dimensional moment sets [2103.10613]. SSMEL targets massive-data computation [2303.07259]. Adjusted EL targets existence and small-sample calibration [1904.08420], [1603.04093]. Bootstrap- or Monte-Carlo-calibrated EL targets simultaneous inference under complex covariance structures [2112.09206]. These should not be conflated.

Finally, **not every adaptive likelihood-based procedure is empirical likelihood**. The current-status-grid paper provides a strong example of adaptive inference through a boundary family of asymptotic distributions, but it belongs to nonparametric maximum likelihood under shape constraints, not to empirical likelihood in the Owen/Qin–Lawless sense [1205.6055].

Taken together, the literature suggests that an adaptive empirical likelihood procedure is best defined operationally rather than doctrinally. It is an EL-based method that changes the empirical likelihood construction, the active estimating equations, the support geometry, the calibration rule, or the computational representation in a data-driven way to maintain inference under nonstandard conditions. The most explicit realizations in the supplied material are pseudo-observation-based adjusted EL [1904.08420], [1603.04093], [1604.06170], [1602.09128], doubly penalized EL with estimating-equation selection [1704.00566], [2103.10613], and split-sample mean EL for massive data [2303.07259].

Source: https://www.emergentmind.com/topics/adaptive-empirical-likelihood-procedure