---
title: Semiparametric Empirical Likelihood
url: https://www.emergentmind.com/topics/semiparametric-empirical-likelihood-el
type: topic
---

# Semiparametric Empirical Likelihood

to=arxiv_search.search  天天中彩票能json
{"query":"semiparametric empirical likelihood arXiv empirical likelihood estimating equations semiparametric", "max_results": 10, "sort_by":"relevance"}
to=arxiv_search.search ിച്ചിട്ടുണ്ട്ഞ്ജसन
{"query":"1801.08812 empirical likelihood robust estimation linear regression", "max_results": 5, "sort_by":"relevance"}
Semiparametric empirical likelihood (EL) is a likelihood-based inferential framework in which a finite-dimensional parameter is specified through estimating equations or structural moment restrictions, while the underlying distribution is left unspecified and represented by unknown probability masses on the observed sample points [1010.0313][1801.08812][2301.09704]. Rather than postulating a full parametric density, semiparametric EL maximizes a nonparametric likelihood over weights subject to simplex and estimating-equation constraints, then profiles out those weights with Lagrange multipliers to obtain estimators, likelihood-ratio statistics, confidence regions, and hypothesis tests [1010.0313][2102.13232]. Across recent work, the same core construction has been adapted to robust regression, missing data, density ratio models, structural equation models, weakly dependent time series, stratified metric spaces, capture–recapture models, spatial hierarchies, and ensemble learners.

## 1. Core construction and semiparametric viewpoint

In the standard estimating-equation formulation, the parameter \(\theta\) is defined by
\[
E\{g(X;\theta)\}=0,
\]
and empirical likelihood maximizes
\[
L_n(\theta)= \sup \Biggl\{ \prod_{i=1}^n p_i : p_i\ge 0,\; \sum_{i=1}^n p_i=1,\; \sum_{i=1}^n p_i g(x_i;\theta)=0 \Biggr\}.
\]
The empirical log-likelihood ratio is
\[
R_n(\theta)=-2\log\bigl(n^n L_n(\theta)\bigr),
\]
and, when the constraints are feasible, the maximizing weights satisfy the usual dual form
\[
p_i=\frac{1}{n}\frac{1}{1+\lambda^\top g(x_i;\theta)},
\]
with \(\lambda\) determined by the multiplier equation
\[
\sum_{i=1}^{n}\frac{g(x_i;\theta)}{1+\lambda^\top g(x_i;\theta)}=0
\]
[1010.0313].

The semiparametric character is explicit in regression formulations. For the linear model
\[
Y_i=\mathbf X_i^{T}\boldsymbol\beta+\varepsilon_i,
\]
the regression structure is specified parametrically through \(\boldsymbol\beta\), while the distribution of the observations is left unspecified and represented nonparametrically by unknown masses \(p_i\). The EL criterion is
\[
L(\boldsymbol\beta)=\prod_{i=1}^{n}p_i,
\]
under
\[
\sum_{i=1}^{n}p_i=1,\qquad
\sum_{i}^{n}p_{i}\left(Y_{i}-\mathbf X_{i}^{T}\beta\right)X_{i}=0,
\]
which is the EL analogue of the OLS normal equations, with equal weights \(1/n\) replaced by unknown \(p_i\) [1801.08812].

The same logic extends beyond single-sample moment models. In the two-sample density ratio model, the baseline distribution \(F_0\) is represented nonparametrically on the pooled observed support by masses \(p_{ij}\), while the second sample is linked through
\[
dF_1(x)=\exp\{\theta^\top Q(x)\}\,dF_0(x),
\]
and auxiliary information is imposed by
\[
\sum_{i,j}p_{ij} g(X_{ij};\eta)=0
\]
[2102.13232]. This pattern—parametric low-dimensional structure plus nonparametric masses under unbiased constraints—is the defining architecture of semiparametric EL.

## 2. Likelihood-ratio inference, Wilks phenomena, and high-order refinement

Under standard regularity conditions, empirical likelihood has the central likelihood-ratio property
\[
P\{R_n(\theta_0)\le x\}=P\{\chi_q^2\le x\}+O(n^{-1}),
\]
which yields asymptotically valid confidence regions of the form
\[
\{\theta: R_n(\theta)\le c_{1-\alpha}(q)\}
\]
[1010.0313]. This Wilks-type behavior is one reason EL is routinely described as likelihood-style inference without a fully specified parametric model.

A practical obstruction is the convex-hull condition: ordinary EL is undefined if \(0\) does not lie in the convex hull of the estimating functions. Adjusted empirical likelihood (AEL) resolves this by adding a pseudo-observation
\[
g_{n+1}=-a_n \bar g_n,\qquad \bar g_n = n^{-1}\sum_{i=1}^n g_i,
\]
and maximizing the augmented likelihood over \(n+1\) points. With a specific adjustment level,
\[
a_n = \frac{b}{2} + O_p(n^{-1/2}),
\]
where \(b\) is the Bartlett correction factor for ordinary EL, AEL achieves the same \(O(n^{-2})\) chi-square accuracy as Bartlett-corrected EL while also guaranteeing existence of the adjusted estimating equations [1010.0313]. In the scalar case \(q=1\), the Bartlett factor is
\[
b=\frac{1}{2}\alpha_4-\frac{1}{3}\alpha_3^2
\]
under \(E[g^2]=1\), and the paper also discusses the two-pseudo-observation construction for \(q>1\).

Jackknife empirical likelihood addresses nonlinear estimands, especially U-statistics, by replacing the original problem with a mean problem on pseudo-values
\[
V_i = nU_n - (n-1)U_{n-1}^{(i)}.
\]
Adjusted jackknife empirical likelihood (AJEL) adds
\[
V_{n+1} = -a_n U_n,
\]
with \(a_n=\log n/2\) suggested and \(a_n=o_p(n^{2/3})\) sufficient, so that the empirical likelihood ratio is well-defined for all \(\theta\). The resulting statistic still satisfies
\[
-2\log R(\theta_0)\xrightarrow{d}\chi^2_1
\]
for one-sample and two-sample U-statistics [1603.04093].

Higher-order accuracy also appears in non-Euclidean EL. For Fréchet means on open books, bootstrap calibration improves the coverage error of EL confidence regions from
\[
O(n^{-1}) \longrightarrow O(n^{-2}),
\]
and the paper explicitly notes that EL is Bartlett correctable asymptotically [2412.18818]. This places semiparametric EL among the few nonparametric likelihood methods that can be pushed beyond first-order chi-square calibration.

## 3. Robust, local, and globally consistent variants

Classical regression EL inherits the sensitivity of OLS-like estimating equations. In the linear model, standard EL uses the weighted normal equations, so merely replacing equal weights by unknown \(p_i\) does not sufficiently protect against outliers. A direct robustification replaces the residual moment condition by the M-estimation score equation
\[
\frac{1}{n}\sum_{i=1}^{n}\psi(Y_{i}-\mathbf{X_{i}^{T}\beta})\mathbf X_{i}=0,
\]
or, in EL form,
\[
\sum_{i=1}^{n}p_{i}\psi\left(Y_{i}-X_{i}^{T}\beta\right)X_{i}=0.
\]
Here \(\psi=\rho'\), and bounded or redescending \(\psi\) functions such as Huber’s and Tukey’s bisquare downweight large residuals [1801.08812]. In the contaminated normal simulation \((0.90)N(0,1)+(0.10)N(20,1)\), with \(k=2\) and \(n=50\), the reported MSEs were \(1.1896\) for EL, \(0.6350\) for EL-Huber, \(0.2023\) for EL-Tukey, and \(1.3495\) for OLS, with Tukey’s bisquare usually the best under contamination [1801.08812].

A different line of development treats EL as a local likelihood experiment. A local representation paper shows that for local perturbations \(\theta=\theta_0+\delta_n\tau_n\), the implied EL measures admit the approximation
\[
\sum_{i=1}^n \log \frac{d\tilde P_{\theta+\delta_n\tau_n}}{d\tilde P_{\theta}}(X_i)
= \tau_n^T S_{\theta,n} -\frac12 \tau_n^T K_{\theta,n}\tau_n +o_p(1),
\]
and derives the local estimator
\[
T_n = \theta_n^* + \delta_n K_n^{-1} S_n.
\]
The paper proves consistency, local asymptotic normality, and asymptotic optimality, and emphasizes that the consistency theorem does not require differentiability of \(m(X,\theta)\) [1403.6782].

Global behavior is a separate issue. Standard EL theory is local: it shows that a local maximizer near the truth is consistent, but it does not imply that the global maximizer is the correct one when the EL surface has multiple local or global maxima. A global-consistency treatment establishes strong consistency of the global empirical likelihood maximizer under conditions C1–C5, including identifiability, finite moments, local Lipschitz continuity, a closed parameter space, and a growth/nondegeneracy condition at infinity [2303.16410]. The same paper proposes a global maximum test based on the value of the profile EL ratio at a candidate maximizer and a remedy that expands the estimating-function vector so that the EL criterion becomes globally identifying.

## 4. Efficiency gains, side information, and doubly robust estimation

A major attraction of semiparametric EL is that valid side information can be encoded as constraints on the weights and converted directly into efficiency gains. For a linear functional
\[
\theta = \int \psi(z)\, dQ(z),
\]
with side information \(E\{g(Z)\}=0\), the EL weights satisfy
\[
\pi_j=\frac{1}{n}\frac{1}{1+\lambda^\top g(Z_j)},
\]
and the EL-weighted estimator becomes
\[
\hat\theta_{\text{EL}} = \frac{1}{n}\sum_{j=1}^n \frac{\psi(Z_j)}{1+\lambda^\top g(Z_j)}.
\]
Under fixed finite constraints,
\[
\hat\theta_{\text{EL}} = \bar\psi - \bar\psi_0 + o_p(n^{-1/2}),\qquad
\sqrt{n}(\hat\theta_{\text{EL}}-\theta) \Rightarrow N(0,\Sigma_0),
\]
with
\[
\Sigma_0 = \operatorname{Var}\!\big(\psi(Z)-\Pi(\psi\mid[g])\big),
\]
so the asymptotic variance is the variance of the residual after projection onto the span of the constraint functions [2301.09704]. The same paper allows estimated constraints \(\hat g\) and a growing number of constraints \(m_n\to\infty\), provided growth and moment conditions hold.

In density ratio models, auxiliary information is incorporated by unbiased estimating equations
\[
E_0\{g(X;\eta)\}=0
\]
on top of the semiparametric relation
\[
dF_1(x)=\exp\{\theta^\top Q(x)\}\,dF_0(x).
\]
The maximum empirical likelihood estimators \((\hat\theta,\hat\eta)\) are asymptotically normal, the ELR statistic for hypotheses \(H(\theta,\eta)=0\) converges to \(\chi_q^2\), and the paper proves a monotonicity statement: if \(r>p\), dropping one valid estimating equation cannot reduce the asymptotic variance; if \(r=p\), including the just-identified valid estimating equations does not change the asymptotic variance relative to the DRM-only estimator [2102.13232].

Missing-data problems make the semiparametric efficiency role of EL especially explicit. For parameters defined by general estimating equations under missing at random,
\[
P(\delta=1\mid y,x)=P(\delta=1\mid x)=:\omega(x),
\]
an efficient doubly robust EL approach constructs stacked estimating functions combining a propensity model and a working model for \(E\{s(z,\beta)\mid x\}\). The resulting estimator is consistent if either the propensity model is correct or the regression model is correct, and when both are correctly specified it achieves the full semiparametric efficiency bound [1612.00922]. In randomized trials, two-sample EL weighting for the average treatment effect,
\[
\hat\theta = \sum_{i=1}^m \hat p_i Y_{1i}-\sum_{j=1}^n \hat q_j Y_{0j},
\]
uses arm-specific empirical likelihood weights with covariate-balancing constraints, is semiparametric efficient when the working outcome regressions are correct, and extends under MAR missingness to estimators with double robustness and multiple robustness [2008.12989].

## 5. Dependence, nuisance functions, and singular geometry

Semiparametric EL is not confined to i.i.d. settings. For stationary, strongly mixing time series, a conditional heteroscedastic partially linear single-index model is rewritten as a finite set of unconditional moments,
\[
\mathbb E\!\left[\Psi(Z_i,Z_i^{\{r\};\theta,\eta_\gamma})\right]=0,
\]
where the nuisance vector \(\eta_\gamma\) contains nonparametric objects such as \(m_\gamma\), \(m_\gamma'\), conditional expectations, and the index density. These components are estimated by kernel smoothing, yet the empirical log-likelihood ratio
\[
\ell_n(\theta,\eta_\gamma) = \sum_{i=1}^n \log\!\left(1+\lambda(\theta,\eta_\gamma)^\top \Psi(Z_i,Z_i^{\{r\};\theta,\eta_\gamma})\right)
\]
has the same first-order limit when \(\eta_\gamma\) is replaced by \(\widehat\eta_\gamma\). The resulting Wilks theorem is
\[
2\ell_n(\theta_0,\widehat\eta_{\gamma_0}) \xrightarrow{d}\chi^2_{d_\theta-1}
\]
[2006.06350].

Recent work on Fréchet means shows that the asymptotic law of the EL ratio can depend on the local topology of the parameter space. On an open book \(M\), EL is defined off the spine by a folding map \(F_k\) and on the spine by projection \(P_s\) plus page-direction inequalities. If the population Fréchet mean lies in the interior of a page,
\[
-2\log(\mathcal R_n(x_0)) \xrightarrow{d} \chi_p^2.
\]
If the mean lies on the spine and is sticky,
\[
-2\log(\mathcal R_n(x_0)) \xrightarrow{d} \chi_{p-1}^2,
\]
while in the half-sticky case
\[
-2\log(\mathcal R_n(x_0)) \xrightarrow{d} \frac12\left\{\chi_p^2+\chi_{p-1}^2\right\}.
\]
The paper’s central point is that the EL limit is governed not just by ambient dimension but by the stratified geometry near the mean [2412.18818]. This suggests that semiparametric EL can remain viable even when the parameter of interest lives on singular spaces, but the Wilks law must then be read geometrically.

## 6. Specialized implementations and contemporary extensions

The flexibility of semiparametric EL is visible in the range of model classes to which the same constrained-weighting mechanism has been adapted.

| Setting | Distinctive EL construction | Paper |
|---|---|---|
| Linear regression with AR(\(p\)) errors | Transformed regression moments jointly constrain \(\boldsymbol\beta\), \(\boldsymbol\phi\), and \(\sigma^2\); simulation shows smaller MSE and bias than CML in almost all configurations | [2008.03282] |
| Linear SEMs with dependent, non-Gaussian errors | Profiles out \(\Omega\), adds AEL and EEL, and reports runtime up to 40 times faster for profiled EL | [1710.02588] |
| Bayesian spatial hierarchical models | Places EL at the data stage and a spatial prior at the process stage in SHEL models; all three data examples report lower MSPE than standard parametric analyses | [1405.3880] |
| Capture–recapture abundance | EL ratio for abundance has limiting \(\chi_1^2\); extensions handle one-inflation and MAR covariates by semiparametric profiling and score testing | [2507.10388], [2507.10404] |
| High-dimensional moment systems | Penalizes the Lagrange multipliers in Bayesian penalized EL and uses Metropolis–Hastings or MAMIS rather than direct optimization | [2412.17354] |
| Random forests and ensembles | Recasts predictions as incomplete generalized \(U\)-statistics and uses modified EL to restore \(\chi_1^2\) pivotality under sparse subsampling | [2511.13934] |

These implementations share the same statistical grammar. Unknown masses or weights are optimized under model-implied constraints; nuisance structure is either profiled out, estimated in a first stage, or absorbed into auxiliary estimating equations; and the inferential target is typically a profile EL ratio with asymptotic chi-square calibration or a modified version thereof. What changes across domains is the constraint geometry: autoregressive residual structure in time series regression, structural zeros in mixed-graph SEMs, spatial latent processes in SHEL models, one-inflated count mechanisms in capture–recapture, sparse multiplier support in penalized EL, or jackknife pseudo-values for incomplete \(U\)-statistics.

Taken together, these developments portray semiparametric empirical likelihood as a general inferential technology for models that are too structured to be purely nonparametric and too distribution-sensitive to be handled comfortably by full parametric likelihood. The recurring advantages are likelihood-ratio inference from moment restrictions, compatibility with nuisance estimation and side information, and a capacity for refinement through adjustment, robustification, bootstrap calibration, or penalization when the basic EL geometry becomes fragile.

Source: https://www.emergentmind.com/topics/semiparametric-empirical-likelihood-el