---
title: Maximum Empirical Likelihood Estimation
url: https://www.emergentmind.com/topics/maximum-empirical-likelihood-estimation
type: topic
---

# Maximum Empirical Likelihood Estimation

Searching arXiv for recent and foundational papers on maximum empirical likelihood estimation and closely related empirical likelihood variants.
Maximum empirical likelihood estimation (MELE) is the empirical-likelihood analogue of maximum likelihood estimation for models defined by estimating equations or moment restrictions rather than a fully specified parametric density. In its standard form, empirical likelihood assigns nonnegative weights to observed data and maximizes a nonparametric likelihood subject to constraints such as \(E[g(X,\theta_0)]=0\) or \(\int \Phi(\theta_0,x)\,d\mu_0(x)=0\). MELE then selects the parameter value that maximizes the resulting profile empirical likelihood, equivalently minimizes the empirical log-likelihood ratio. Across the literature, the method appears in i.i.d. estimating-equation models, generalized empirical likelihood and maximum-entropy formulations, dependent-data problems in the frequency domain, high-dimensional inference, semiparametric efficiency constructions with many constraints, Bayesian computation via empirical likelihood, and specialized rare-event Monte Carlo schemes built around empirical likelihood maximization [1306.1493], [1202.6469], [2303.16410], [1010.0313], [2302.14768], [0708.0197], [1805.10742], [1205.5658], [1312.3027].

## 1. Definition and basic optimization structure

In the estimating-equation framework, the parameter of interest \(\theta_0\in\mathbb R^p\) is defined by
\[
E[g(X,\theta_0)] = 0,
\]
with \(g(X,\theta)\in\mathbb R^q\), and the data \(X_1,\dots,X_n\) are i.i.d. copies of \(X\). The literature explicitly treats both just-determined problems, where \(q=p\), and over-determined problems, where \(q>p\) [1306.1493]. A closely related moment-condition formulation writes the target as the unique \(\theta_0\in\Theta\subset\mathbb R^d\) satisfying
\[
\int_{\mathcal X} \Phi(\theta_0,x)\,d\mu_0(x)=0,
\]
with \(\Phi:\Theta\times\mathcal X\to\mathbb R^k\) and \(k\ge d\) [1202.6469].

For fixed \(\theta\), the original empirical likelihood maximizes a multinomial-type likelihood over weights \(w_i\) or \(p_i\). One standard formulation is
\[
R(\theta)=\sup\left\{\prod_{i=1}^n n w_i: \sum_{i=1}^n w_i g(X_i,\theta)=0,\; w_i\ge 0,\; \sum_{i=1}^n w_i=1\right\},
\]
with empirical log-likelihood ratio
\[
l(\theta)=-2\log R(\theta).
\]
Equivalent notation also appears as
\[
L_n(\theta) = \sup\left\{ \prod_{i=1}^n p_i:\; p_i\ge 0,\; \sum_{i=1}^n p_i=1,\; \sum_{i=1}^n p_i g(x_i;\theta)=0 \right\},
\]
or, in profile form,
\[
e_n(\theta) = \sup \left\{ \sum_{i=1}^n \log p_i: \sum_{i=1}^n p_i=1,\; \sum_{i=1}^n p_i g(x_i;\theta)=0 \right\}
\]
[1306.1493], [1010.0313], [2303.16410].

When the constraints are feasible, Lagrange multiplier arguments yield the familiar weights
\[
p_i(\theta)=\frac{1}{n\{1+\lambda^\top h(y_i,\theta)\}}
\]
or
\[
\hat p_i=\frac{1}{n[1+\hat\lambda_n^\tau g(x_i;\theta)]},
\]
where the multiplier solves
\[
\sum_{i=1}^n \frac{g(X_i,\theta)}{1+\lambda^T g(X_i,\theta)}=0
\]
or its analogous form under the chosen notation [1306.1493], [1205.5658], [2303.16410], [1010.0313].

MELE is defined by maximizing the empirical likelihood over \(\theta\), equivalently minimizing the empirical log-likelihood ratio. Representative definitions are
\[
\hat\theta = \arg\max_\theta L_n(\theta)\quad\text{equivalently}\quad \hat\theta = \arg\min_\theta R_n(\theta),
\]
\[
\hat\theta_n=\arg\max_\theta e_n(\theta)=\arg\min_\theta W_n(\theta),
\]
and, in fixed-dimensional classical notation,
\[
\check\theta_n=\arg\max_{\theta\in\Theta}L(\theta)
\]
[1010.0313], [2303.16410], [1805.10742].

This optimization-based definition makes empirical likelihood a nonparametric likelihood method under moment constraints. The method is repeatedly described as “maximum likelihood under moment constraints” or as an estimator that “plays the role of a likelihood-based estimator, but without an explicit parametric likelihood” [1010.0313], [1205.5658]. A plausible implication is that the central object in MELE is not a parametric sampling density but the compatibility between observed data and imposed estimating equations.

## 2. Geometric constraints, existence, and domain corrections

A fundamental limitation of ordinary empirical likelihood is that it is only defined when the origin belongs to the convex hull of the estimating-function vectors. In the i.i.d. estimating-equation setting, the domain \(\Theta_n\) of the original empirical likelihood is generally a subset of \(\mathbb R^p\), often bounded, because feasibility requires
\[
0 \in \mathrm{conv}\{g(X_i,\theta)\}_{i=1}^n.
\]
If this convex-hull condition fails, the weights do not exist and the empirical likelihood is undefined [1306.1493]. The same difficulty is emphasized in adjusted empirical likelihood, where the ratio is defined only if \(0\) lies in the convex hull of \(\{g(x_i;\theta): i=1,\dots,n\}\) [1010.0313].

This domain mismatch matters statistically. The original empirical likelihood domain may fail to coincide with the true parameter space \(\mathbb R^p\), and this is identified as a principal reason for undercoverage and poor finite-sample performance, especially in over-determined problems [1306.1493]. In small samples or with high-dimensional estimating functions, nonexistence of solutions can be frequent enough to hinder practical use [1010.0313].

Two distinct remedies appear in the supplied literature.

The first is **extended empirical likelihood**. Rather than altering the empirical likelihood objective itself, the construction expands the restricted domain through a geometric mapping centered at the maximum empirical likelihood estimator \(\tilde\theta\):
\[
h_n^C(\theta)=\tilde\theta+\left(1+\frac{l(\theta)}{2n}\right)(\theta-\tilde\theta), \qquad \theta\in\Theta_n.
\]
This “composite similarity mapping” is a similarity transformation on each empirical-likelihood contour and is surjective from \(\Theta_n\) onto \(\mathbb R^p\) under Conditions 1–3; with an additional nested-contour condition it becomes bijective [1306.1493]. Because the mapping need not be one-to-one, a generalized inverse is defined by
\[
h_n^{-C}(\theta)=\arg\min_{\theta'\in s(\theta)}\|\theta'-\theta\|,
\quad s(\theta)=\{\theta':h_n^C(\theta')=\theta\},
\]
and the extended empirical log-likelihood ratio is
\[
l^*(\theta)=l(h_n^{-C}(\theta)), \qquad \theta\in\mathbb R^p.
\]
The result is an empirical likelihood defined on the full parameter space while retaining the original contour shape [1306.1493].

The second is **adjusted empirical likelihood**, which forces feasibility by adding pseudo-observations. With \(g_i=g(x_i;\theta)\),
\[
g_{n+1}=-a_n\bar g_n, \qquad \bar g_n=\frac{1}{n}\sum_{i=1}^n g_i, \qquad a_n>0.
\]
The adjusted empirical likelihood becomes
\[
L_n(\theta;a_n) = \sup\left\{ \prod_{i=1}^{n+1}p_i:\; p_i\ge0,\; \sum_{i=1}^{n+1}p_i=1,\; \sum_{i=1}^{n+1}p_i g_i=0 \right\},
\]
with adjusted ratio
\[
R_n(\theta;a_n) = -2\log\{(n+1)^{n+1}L_n(\theta;a_n)\}.
\]
Because \(g_{n+1}=-a_n\bar g_n\) lies on the opposite side of the origin from \(\bar g_n\), the augmented vectors can always be arranged so that the origin lies in their convex hull, and the adjusted optimization problem is always feasible [1010.0313].

These constructions do not redefine the estimator in a wholly different way. In the extended version, the minimum of \(l^*(\theta)\) remains at the same center \(\tilde\theta\), and maximizing the extended empirical likelihood does not produce a new estimator different from the MELE [1306.1493]. In the adjusted version, the core profile-likelihood logic remains the same, but feasibility is guaranteed [1010.0313].

## 3. Asymptotic theory, Wilks-type limits, and higher-order accuracy

The empirical-likelihood literature represented here repeatedly emphasizes that MELE inherits many asymptotic features of parametric likelihood under regularity conditions. For standard empirical likelihood, the empirical log-likelihood ratio at the true parameter has a chi-square approximation, and maximum empirical likelihood estimators are consistent and asymptotically normal under regularity [1205.5658]. In the extended empirical likelihood setting, the main Wilks-type result is
\[
l^*(\theta_0)\xrightarrow{d}\chi_q^2,
\]
so the extension preserves the first-order chi-square limit of the original empirical likelihood [1306.1493].

The higher-order theory is particularly explicit in the literature on Bartlett correction and adjusted or extended variants. For original empirical likelihood,
\[
l(\theta_0)=nR^TR+O_p(n^{-3/2}),
\]
and
\[
l(\theta_0)\bigl[1-bn^{-1}+O_p(n^{-3/2})\bigr]
\]
has a \(\chi_q^2\) approximation with error \(O(n^{-2})\), where \(b\) is the Bartlett correction factor [1306.1493]. In the scalar mean case, the factor is
\[
b=\frac{1}{2}\alpha_4-\frac{1}{3}\alpha_3^2,
\]
with \(\alpha_r=E\{g(X;\theta)\}^r\) after standardization so that \(\alpha_2=1\) [1010.0313].

Adjusted empirical likelihood shows that a specific adjustment level reproduces this higher-order behavior. If
\[
a_n = a + O_p(n^{-1/2}),
\]
then, under the stated regularity assumptions, the adjusted ratio has an expansion of the same order, and when
\[
a=\frac{b}{2},
\]
the adjusted empirical likelihood achieves the same second-order accuracy as Bartlett-corrected empirical likelihood [1010.0313]. The mechanism appears in the asymptotic relation
\[
\lambda_a = \lambda - n^{-1}a\,\bar g + O_p(n^{-2}),
\]
which leads to
\[
R_n(\theta_0;a_n) = R_n(\theta_0)-2a\,R_1^TR_1+O_p(n^{-3/2})
\]
[1010.0313].

Extended empirical likelihood also admits a second-order construction in the just-determined case:
\[
\gamma_2(n,l(\theta))=1+\frac{b}{2n}[l(\theta)]^{\delta(n)},
\qquad \delta(n)=O(n^{-1/2}),
\]
yielding
\[
l_2^*(\theta_0)=l(\theta_0)\bigl[1-bn^{-1}+O_p(n^{-3/2})\bigr].
\]
The paper further gives the first-order expansion
\[
l^*(\theta_0)=l(\theta_0)\bigl[1-l(\theta_0)n^{-1}+O_p(n^{-3/2})\bigr],
\]
and states that this helps explain the strong finite-sample performance of the first-order extended empirical likelihood even without explicit Bartlett correction [1306.1493].

The regularity conditions required for these expansions are not minimal. Extended empirical likelihood uses Conditions 1–3, including positive definiteness of \(V[g(X,\theta_0)]\), smoothness of \(g\), a Cramér-type condition, and the moment condition
\[
E\|g(X,\theta_0)\|^{15}<\infty
\]
[1306.1493]. Adjusted empirical likelihood assumes Cramér’s condition, finite \(18\)th moments,
\[
E\|g(X;\theta)\|^{18}<\infty,
\]
and nonsingular covariance, with additional smoothness in the over-identified case [1010.0313].

A common misconception is that empirical-likelihood refinements merely repair existence problems. The supplied papers show a broader point: domain correction and feasibility adjustment are tied to higher-order coverage behavior as well as definition on the parameter space [1306.1493], [1010.0313].

## 4. Global maximization, identifiability, and consistency of the empirical-likelihood maximizer

The distinction between a local empirical-likelihood maximizer and the global maximizer is a central issue in the modern theory of MELE. In the profile empirical-likelihood framework,
\[
W_n(\theta) = -\sum_{i=1}^n \log(n\hat p_i) = \sum_{i=1}^n \log\left[1+\hat\lambda_n^\tau g(x_i;\theta)\right],
\]
and the MELE is defined as the global minimizer of \(W_n(\theta)\) [2303.16410]. The significance of global optimization is that, like parametric likelihood, the empirical-likelihood criterion can have multiple local or global extrema. The local asymptotic theory is not sufficient unless one knows that the extremum under consideration is the global one [2303.16410].

A major recent contribution is a strong global consistency theorem for the empirical-likelihood maximizer under Conditions C1–C5. These conditions require: uniqueness of the population root
\[
E\{g(X;\theta)\}=0 \quad \text{has a unique solution } \theta^*,
\]
finite moments and positive definiteness of
\[
E\{g(X;\theta)g(X;\theta)^\tau\},
\]
local Lipschitz continuity on compact sets, a closed parameter space, and a scaling vector \(b(\theta)\) such that \(h(X;\theta)=b^\tau(\theta)g(X;\theta)\) behaves properly at infinity [2303.16410]. Under these conditions,
\[
\hat\theta_n \to \theta^* \quad \text{almost surely as } n\to\infty.
\]
This is explicitly a theorem on strong global consistency, not merely local consistency near \(\theta^*\) [2303.16410].

The same work also introduces a **global maximum test**. Given a numerically found local maximizer \(\tilde\theta_n\), one rejects the claim that it is global when
\[
p\text{-value} = P\{\chi^2_{q-m}>2W_n(\tilde\theta_n)\} <\alpha.
\]
This is presented as a diagnostic for global optimality rather than a model-validity test [2303.16410].

A more structural remedy is to enlarge the estimating-function vector. If an initial \(g_1(x;\theta)\) leads to multiple global maxima or fails C1–C5, the proposed remedy is to add unbiased estimating functions \(g_2\) and use
\[
g(x;\theta)= \begin{pmatrix} g_1(x;\theta)\\ g_2(x;\theta) \end{pmatrix}
\]
so that the combined system satisfies C1–C5 [2303.16410]. The paper illustrates this strategy in several examples. For the Cauchy model, the score-like function
\[
g_1(x;\theta)=\frac{x-\theta}{1+(x-\theta)^2}
\]
admits multiple roots and fails C5, but adding \((x-\theta)^{1/3}\) yields a globally consistent MELE [2303.16410]. In nonlinear regression with
\[
y_i=\theta+\theta^2 x_i+\epsilon_i,
\]
a one-dimensional estimating function may have three roots, whereas the overidentified system
\[
g(x,y;\theta)= \begin{pmatrix} y-\theta-\theta^2 x\\ x(y-\theta-\theta^2 x) \end{pmatrix}
\]
restores global consistency [2303.16410].

This body of results modifies an often implicit assumption in empirical-likelihood practice: solving the estimating equations or finding a numerically acceptable local optimum is not itself enough to justify the MELE. The paper’s position is that global identifiability and behavior at infinity must be checked, and, when necessary, additional estimating equations should be introduced [2303.16410].

## 5. Generalized empirical likelihood, maximum entropy, and Bayesian interpretation

MELE sits inside the broader class of generalized empirical likelihood (GEL) estimators. In moment-condition models, GEL is defined through divergence minimization. For a convex \(f\) with \(f(1)=f'(1)=0\), the \(f\)-divergence is
\[
\mathcal D_f(\nu\mid \mu)= \begin{cases} \displaystyle \int_{\mathcal X} f\!\left(\frac{d\nu}{d\mu}\right)\,d\mu, & \nu\ll\mu,\\[1em] +\infty,& \text{otherwise.} \end{cases}
\]
The GEL estimator is
\[
\hat\theta=\arg\min_{\theta\in\Theta}\mathcal D_f(\mathcal M_\theta,\mathbb P_n),
\]
where \(\mathcal M_\theta=\{\mu:\mu[\Phi(\theta,\cdot)]=0\}\) and \(\mathbb P_n=n^{-1}\sum_{i=1}^n\delta_{X_i}\) [1202.6469].

The paper on Bayesian interpretation of GEL shows that a large class of GEL estimators can be represented as **maximum entropy on the mean (MEM)** solutions, with priors on the weight vector \(W\). In the i.i.d. prior case \(\nu_0=\nu^{\otimes n}\), the estimator has saddle-point representation
\[
\hat \theta = \arg \min_{\theta \in \Theta} \ \sup_{(\gamma,\lambda)\in\mathbb R \times \mathbb R^k} \left\{ \gamma-\mathbb P_n\!\left[\Lambda_\nu\big(\gamma+\lambda^t\Phi(\theta,\cdot)\big)\right] \right\},
\]
where \(\Lambda_\nu\) is the log-Laplace transform of the prior [1202.6469].

Within this framework, empirical likelihood itself corresponds to an exponential prior
\[
d\nu(x)=e^{-x}dx \quad \text{on } x>0,
\]
for which
\[
\Lambda_\nu(s)=-\log(1-s),\qquad s<1.
\]
Exponential tilting corresponds to a Poisson prior, and the continuous updating estimator corresponds to a Gaussian prior \(\mathcal N(1,1)\) [1202.6469]. The paper therefore treats empirical likelihood not as an isolated construction but as one member of a maximum-entropy family of discrepancy-based estimators.

This viewpoint also supports robustness to approximate moment conditions. If \(\Phi\) is approximated by \(\Phi_m\), the approximate GEL estimator becomes
\[
\hat \theta_m = \arg \min_{\theta \in \Theta} \ \sup_{(\gamma,\lambda) \in \mathbb R \times \mathbb R^{k} } \left\{ \gamma-\mathbb P_n\!\left[\Lambda(\gamma+\lambda^t\Phi_m(\theta,\cdot))\right] \right\},
\]
and if A.1–A.9 hold,
\[
n\|\hat\theta_m-\hat\theta\|^2 = O_P\!\left(n\varphi_m^{-2}\right)+o_P(1).
\]
If
\[
n\varphi_m^{-2}\to 0,
\]
then \(\hat\theta_m\) is \(\sqrt n\)-consistent and asymptotically equivalent to the exact GEL estimator [1202.6469].

The Bayesian strand continues in empirical-likelihood-based posterior approximation. In Bayesian computation via empirical likelihood, the intractable likelihood is replaced by an empirical likelihood \(L_{\text{el}}(\theta\mid y)\), and the target posterior approximation is
\[
\pi_{\text{BC}}(\theta\mid y) \propto \pi(\theta)\, L_{\text{el}}(\theta\mid y).
\]
The basic BC algorithm samples \(\theta_i\sim\pi(\theta)\) and assigns weights
\[
\omega_i = L_{\text{el}}(\theta_i\mid y),
\]
with approximation quality monitored by the effective sample size
\[
\text{ESS} = \frac{1}{\sum_{i=1}^M \left(\omega_i / \sum_{j=1}^M \omega_j\right)^2 }.
\]
An adaptive BC-AMIS variant updates multivariate Student proposals \(q_t=t_3(\cdot\mid \mathbf m_t,\Sigma_t)\) and uses mixture-proposal weights [1205.5658].

These Bayesian and maximum-entropy formulations do not replace MELE in frequentist inference, but they show that empirical-likelihood maximization admits a probabilistic interpretation through priors on weights and can serve as a surrogate likelihood inside Bayesian samplers [1202.6469], [1205.5658].

## 6. Extensions to dependent data, high dimension, and many constraints

MELE has been extended well beyond the classical fixed-dimensional i.i.d. setting.

For dependent time series, **frequency domain empirical likelihood (FDEL)** replaces time-domain observations by periodogram ordinates at Fourier frequencies. With spectral estimating functions \(G_\theta(\lambda)\), the profile FDEL is
\[
L_n(\theta) = \sup \left\{ \prod_{j=1}^N w_j: w_j\ge 0,\ \sum_{j=1}^N w_j=\pi,\ 
 \sum_{j=1}^N w_j\, G_\theta(\lambda_j) I_n(\lambda_j)=M \right\},
\]
where
\[
I_n(\lambda)=\frac{1}{2\pi n}\left|\sum_{t=1}^n X_t e^{-it\lambda}\right|^2
\]
is the periodogram and \(\lambda_j=2\pi j/n\) [0708.0197]. The corresponding MELE is
\[
\hat\theta_n = \arg\max_{\theta\in\Theta} R_n(\theta),
\]
and Theorem 2 gives a local maximizer \(\hat\theta_n\) near \(\theta_0\) such that
\[
\hat\theta_n \xrightarrow{p} \theta_0, \qquad
\sqrt{n}(\hat\theta_n-\theta_0)\xrightarrow{d}N(0,V_{\theta_0}).
\]
Theorem 3 provides conditions under which the global maximizer exists with probability tending to \(1\) and is consistent [0708.0197]. This extension covers short-range and long-range dependence and applies to autocorrelations, normalized spectral distribution quantities, Whittle estimation, and spectral goodness-of-fit [0708.0197].

For high-dimensional estimating-equation models, the challenge is that the parameter dimension \(p\) and the number of estimating equations \(r\) may grow exponentially. The cited work uses a penalized empirical likelihood estimator
\[
\hat{\theta}_{PEL}=\arg\min_{\theta\in\Theta}\max_{\lambda\in\hat{\Lambda}_n(\theta)} \Bigg[ \sum_{i=1}^n\log\{1+\lambda^\top g_i(\theta)\} +n\sum_{k=1}^pP_{1,\pi}(|\theta_k|) -n\sum_{j=1}^rP_{2,\nu}(|\lambda_j|) \Bigg],
\]
with penalties on both \(\theta\) and \(\lambda\) [1805.10742]. For valid inference on a low-dimensional target subvector \(\theta_{\mathcal M}\), the paper constructs transformed estimating equations
\[
f^{A_n}(X;\theta)=A_n g(X;\theta),
\]
where \(A_n\) is chosen so that nuisance-gradient effects are asymptotically negligible. The resulting transformed empirical-likelihood ratio
\[
\ell_{A_n}^*(\theta_{\mathcal M})=-2\log\{n^nL_{A_n}^*(\theta_{\mathcal M};\theta_{\mathcal M^c}^*)\}
\]
has a Wilks-type limit under stated rate conditions, either \(\chi_m^2\) for fixed \(m\) or a normal approximation when \(m\to\infty\) [1805.10742].

A distinct extension studies estimation of a linear functional
\[
\theta=\int \varphi(z)\,dQ(z)=E\{\varphi(Z)\}
\]
under side information characterized by infinitely many constraints. The “easy EL” estimator uses weights
\[
\hat\pi_j=\frac{1}{n}\frac{1}{1+\hat u(Z_j)^\top \hat\lambda_n},
\]
producing
\[
\hat\theta_n=\sum_{j=1}^n \hat\pi_j\,\varphi(Z_j)
=\frac1n\sum_{j=1}^n \frac{\varphi(Z_j)}{1+\hat u(Z_j)^\top \hat\lambda_n}.
\]
With estimated constraint functions and a growing number \(m_n\) of constraints, the paper establishes semiparametric efficiency in settings including known marginals, unknown but identical marginals, and distributional symmetry [2302.14768]. The key expansion is
\[
\hat\theta_n=\bar\varphi-\overline{\varphi_0}+o_p(n^{-1/2}),
\]
where \(\varphi_0\) is the projection of \(\varphi\) onto the closed linear span of the constraints [2302.14768]. This suggests that empirical-likelihood weighting can be understood as projection-based variance reduction in semiparametric models.

Taken together, these extensions show that MELE is not confined to low-dimensional i.i.d. moment problems. It has been adapted to spectral inference for dependent data, sparse and overidentified high-dimensional models, and semiparametric estimation with infinitely many constraints [0708.0197], [1805.10742], [2302.14768].

## 7. Applications, computational roles, and practical interpretation

MELE and empirical-likelihood maximization appear in several distinct applied and computational roles.

In **Bayesian computation via empirical likelihood**, empirical likelihood replaces an intractable model likelihood. The method is positioned as an alternative to ABC because it does not simulate pseudo-data in the likelihood-approximation step, does not require a tolerance \(\epsilon\), and does not require ad hoc summary-statistic matching in the ABC sense [1205.5658]. The paper illustrates this with standard distributions, time series, the \(g\)-and-\(k\) family, ARCH(1), GARCH(1,1), and population-genetics models based on pairwise composite likelihood score constraints [1205.5658].

In **rare-event probability estimation**, the phrase **Empirical Likelihood Maximization (ELM)** is used for a Monte Carlo method that estimates unknown normalizing constants in a sequence of densities
\[
f_t(x) = \frac{w_t(x)}{\ell_t} = \frac{f(x)H_t(x)}{\ell_t}, \qquad t=1,\dots,s,
\]
one of which embeds the rare-event probability
\[
\ell = \mathbb{P}_f(S(X)\ge \gamma)
\]
as a normalizing constant [1312.3027]. With pooled samples from the \(f_t\), the method reparameterizes via
\[
\mathbf{z}=\left(-\log\left(\frac{\ell_1}{\lambda_1}\right),\dots,-\log\left(\frac{\ell_s}{\lambda_s}\right)\right),
\]
and minimizes the convex objective
\[
\widehat{D}(\mathbf{z}) = \sum_{j=1}^n \log \left (\sum_{k=1}^s w_k(X_j)e^{z_k} \right )  - \sum_{k=1}^s n_k z_k.
\]
The target estimator is then recovered from
\[
\widehat{\ell}_s = \lambda_s e^{-\widehat z_s},
\]
up to the chosen normalization [1312.3027]. Although this usage differs from classical statistical MELE, the paper explicitly frames ELM as paralleling MLE logic: maximize an empirical likelihood to estimate unknown rare-event probabilities [1312.3027].

Practical recommendations and limitations recur across the literature. One recommendation in Bayesian computation via empirical likelihood is to keep the number of constraints equal to the dimension of \(\theta\), because overconstraining can degrade the fit [1205.5658]. In classical EL inference, poor finite-sample behavior is linked to domain mismatch and infeasible constraints [1306.1493], [1010.0313]. In the high-dimensional setting, too large a testing index set \(\mathcal J\) can inflate the critical value and reduce power [1805.10742]. In the semiparametric many-constraint setting, efficiency depends on growth-rate conditions such as
\[
m_n^4/n\to 0
\quad\text{or}\quad
m_n^6/n\to 0
\]
depending on whether constraints are known or estimated [2302.14768].

A common misconception is that maximum empirical likelihood estimation is a single fixed estimator with one canonical form. The supplied papers instead present a family of closely related constructions. The classical MELE is the maximizer of a profile empirical likelihood under exact estimating constraints [2303.16410], [1010.0313]. Extended empirical likelihood leaves the estimator unchanged while expanding the domain [1306.1493]. Adjusted empirical likelihood guarantees feasibility and can match Bartlett-corrected accuracy [1010.0313]. Generalized empirical likelihood embeds empirical likelihood inside a broader divergence and maximum-entropy class [1202.6469]. Frequency-domain, high-dimensional, and many-constraint formulations change the underlying empirical-likelihood problem to fit dependence, nuisance structure, or side information [0708.0197], [1805.10742], [2302.14768].

In that broader sense, MELE is best understood as a nonparametric likelihood principle for moment-defined parameters: choose the parameter value whose associated constraints yield the largest feasible empirical likelihood. The continuing research themes in the cited work are geometric correction of the likelihood domain, higher-order calibration, global rather than merely local maximization, robustness to approximate moments, exploitation of auxiliary constraints for efficiency, and adaptation to computational settings where ordinary likelihood is unavailable or inconvenient [1306.1493], [1010.0313], [2303.16410], [1202.6469], [1205.5658].

Source: https://www.emergentmind.com/topics/maximum-empirical-likelihood-estimation