---
title: Empirical Likelihood Ratio CI
url: https://www.emergentmind.com/topics/empirical-likelihood-ratio-confidence-interval
type: topic
---

# Empirical Likelihood Ratio CI

An empirical likelihood ratio confidence interval is a nonparametric, likelihood-driven confidence set obtained by inverting an empirical likelihood ratio statistic constructed from estimating equations rather than a fully specified parametric model. In its standard form, the method assigns probabilities \(p_1,\dots,p_n\) to observed data and maximizes the nonparametric likelihood subject to normalization and moment constraints; when a Wilks-type theorem holds, the resulting set is typically \(\{\theta:-2\log R(\theta)\le \chi^2_{q,1-\alpha}\}\) [1604.02573, 2112.09206]. The modern literature extends this construction well beyond i.i.d. mean problems to blocked experiments, time series, high-dimensional inference, censored data, complex surveys, nonparametric regression, and policy evaluation, and it includes profile, simultaneous, adjusted, jackknife, and robust variants designed to address nuisance parameters, multiplicity, convex-hull nonexistence, dependence, and bias correction [1805.10742, 2504.01535].

## 1. Core construction

For data \(X_1,\dots,X_n\) and estimating equations \(g(X_i,\theta)\), the empirical likelihood ratio is defined by
\[
R(\theta)
=
\max_{p_1,\dots,p_n}\;\prod_{i=1}^n p_i,
\quad
\text{subject to}
\quad
p_i\ge 0,\;
\sum_{i=1}^n p_i=1,\;
\sum_{i=1}^n p_i\,g(X_i,\theta)=0.
\]
Introducing a Lagrange multiplier \(\lambda(\theta)\), the maximizing weights take the form
\[
p_i(\theta)=\frac{1}{n\{1+\lambda(\theta)^\top g(X_i,\theta)\}},
\]
where \(\lambda(\theta)\) solves
\[
\sum_{i=1}^n \frac{g(X_i,\theta)}{1+\lambda(\theta)^\top g(X_i,\theta)}=0.
\]
The empirical log-likelihood ratio statistic is
\[
-2\log R(\theta)
=
2\sum_{i=1}^n \log\bigl(1+\lambda(\theta)^\top g(X_i,\theta)\bigr)
\equiv l_n(\theta).
\]
On finite support, the unconstrained nonparametric likelihood is maximized at the uniform weights \(u=(1/n,\dots,1/n)\), so the empirical likelihood ratio is naturally interpreted relative to the empirical distribution itself [2112.09206, 1604.02573].

This construction is likelihood-like but distribution-free in the sense that it uses only the moment restrictions encoded by \(g\). In optimization with expected-value objectives and constraints, the same profile ratio appears over empirical weights \(p\in\Delta_n\), and the acceptance set \(\{-2\log R(\theta)\le \chi^2_{k,1-\alpha}\}\) is equivalent to a Burg-entropy divergence ball around the empirical distribution, with radius \(\delta=\chi^2_{k,1-\alpha}/(2n)\) [1604.02573]. In blocked experiments, the canonical choice
\[
g(X_i,\theta)=(X_i-\theta)\circ c_i
\]
encodes missingness and incompleteness directly through the block-incidence vector \(c_i\), so empirical likelihood targets treatment means without explicit covariance modeling [2112.09206].

## 2. Wilks-type calibration and profile inversion

The central inferential step is calibration of \(l_n(\theta)\). Under mild conditions, \(l_n(\theta_0)\overset{d}{\to}\chi^2_q\), where \(q\) is the dimension of the parameter subset or the number of constraints [2112.09206]. In the blocked-design framework, the stated regularity conditions include: \(P\{0\in\mathrm{Conv}_n(\theta_0)\}\to1\); continuous differentiability and uniform consistency of \(g\) and \(G(\theta)=E[g(X_i,\theta)]\); covariance consistency of \(S_n(\theta)=n^{-1}\sum gg^\top\); asymptotic normality of \(a_n G_n(\theta_0)\); maximum bounds on \(g\) and \(\partial_\theta g\); and continuously differentiable hypothesis functions with full-rank Jacobians [2112.09206]. Under these conditions, Wilks-type results extend from a single constraint to profiled and multivariate settings.

For a single contrast \(\psi_j=u_j^\top\theta\), the profile empirical likelihood ratio confidence interval is obtained by inversion:
\[
\mathrm{CI}_j
=
\Bigl\{\psi_j=r:\ \inf_{\theta:\,u_j^\top\theta=r} l_n(\theta)\le c_\alpha\Bigr\}.
\]
This form recurs across applications. In heavy-tail inference for the tail index \(\gamma\), the EL and adjusted EL intervals are
\[
C^{\mathrm{EL}}_{1-\alpha}=\{\gamma:\ l_{\mathrm{EL}}(\gamma)\le \chi^2_{1,1-\alpha}\},
\qquad
C^{\mathrm{AEL}}_{1-\alpha}=\{\gamma:\ l_{\mathrm{AEL}}(\gamma)\le \chi^2_{1,1-\alpha}\},
\]
with \(g_i(\gamma)=y_i-\gamma\) based on top-order log-spacings [1904.08420]. For GPD exceedances, joint confidence regions for \((\gamma,\sigma_n)\) and profile intervals for \(\gamma\) follow from
\[
C_{1-\alpha}=\{(\gamma,\sigma): -2\log R(\gamma,\sigma)\le \chi^2_{2,1-\alpha}\},
\]
and
\[
CI_{1-\alpha}(\gamma)=\{\gamma:\ell_p(\gamma)\le \chi^2_{1,1-\alpha}\},
\]
where \(\ell_p(\gamma)\) profiles out \(\sigma\) [1008.3229].

Wilks-type calibration is not universal in its simplest form. For generalized Lorenz ordinates, the modified empirical likelihood ratio statistics converge to scaled chi-squared distributions with one degree of freedom, with scale \(c(t)=\sigma_p^2(t)/\sigma_v^2(t)\), so the confidence interval uses \(\hat c(t)\ell(\theta)\le \chi^2_{1,1-\alpha}\) rather than the unscaled statistic [2304.04124]. By contrast, for right-censored lifetime data, a specific influence-function construction yields
\[
-2\log R(\theta_0)\Rightarrow \chi^2_1
\]
without estimating any scale parameter, restoring the standard Wilks phenomenon under censoring [1203.5955].

## 3. Multivariate and simultaneous confidence intervals

Empirical likelihood ratio confidence intervals can be embedded in simultaneous inference by exploiting the joint limit of multiple profiled EL statistics. For \(m\) hypotheses \(H_j:h_j(\theta)=0\), the blocked-design theory defines
\[
T_{nj}
=
\frac{a_n^2}{n}\inf_{\theta\in H_j\cap\overline{K}_n} l_n(\theta),
\qquad
\overline{K}_n=\{\theta:\|\theta-\theta_0\|\le K/a_n\},
\]
and shows that
\[
T_n=(T_{n1},\dots,T_{nm})\Rightarrow \chi^2(m,q,R),
\]
a multivariate chi-square law with marginal degrees \(q=(q_1,\dots,q_m)\) and correlation matrix \(R\) induced by the transformed Gaussian vector \(Z_j=(J_jMJ_j^\top)^{-1/2}J_jWU\) [2112.09206]. Asymptotically, subset pivotality holds, so any subset of the \(T_{nj}\) has the same joint limiting distribution under the intersection null as under the complete null.

This permits single-step simultaneous confidence intervals. For contrasts \(\psi_1,\dots,\psi_m\),
\[
\mathcal C=\prod_{j=1}^m \mathrm{CI}_j,
\]
with a common cutoff \(c_\alpha\) chosen to control the generalized family-wise error rate or FWER asymptotically [2112.09206]. Two procedures are developed for \(c_\alpha\). The asymptotic Monte Carlo procedure approximates the limiting \(\chi^2(m,q,R)\) law using the plug-in covariance \(S_n(\hat\theta)\), while the nonparametric bootstrap applies a null transformation, resamples the transformed data, and calibrates the statistic by the empirical quantile of the \(v\)-th largest bootstrap component [2112.09206].

The simulation evidence is specific. In balanced incomplete block designs with \(p=5\), block size \(b_i=2\), and scenarios including non-normality, heteroscedasticity, and violations of compound symmetry, the nonparametric bootstrap achieves \(\mathrm{FWER}\approx 0.05\) and \(\mathrm{CP}\approx 0.95\) even at \(n=50\), whereas AMC converges slower and is more sensitive to skew and thick tails; bootstrap intervals are typically slightly wider because the bootstrap cutoff is larger [2112.09206]. This establishes empirical likelihood ratio intervals not only as marginal devices but also as simultaneous confidence procedures compatible with multiplicity control.

## 4. Convex-hull failure, adjusted EL, and robust weighting

A recurrent obstacle is the convex-hull condition. In time-series EL based on Whittle estimating functions, the maximization fails if \(0\) does not lie in the convex hull of \(\{g_i(\theta)\}\); then no Lagrange multiplier satisfies the constraints and the empirical likelihood is undefined or \(-\infty\) in Owen’s convention [1602.09128]. The same issue appears in heavy-tail tail-index inference when \(k\) is small, in jackknife EL for probability weighted moments, and in blockwise EL for weakly dependent data [1904.08420, 1807.04450, 1912.07699].

The standard remedy is adjusted empirical likelihood. In stationary time-series models, the adjustment augments the estimating equations with
\[
g_{n+1}(\theta)=-a_n\bar g_n(\theta),
\qquad
a_n=\max\{1,(\log n)/2\},
\]
forcing the origin into the convex hull of the augmented set while preserving the \(\chi^2_p\) limit for \(-2\log R_{\mathrm{adj}}(\theta_0)\) [1602.09128]. The same idea is carried to long-memory ARFIMA models, where
\[
\psi_{n+1}(\beta)=-a_n\bar\psi_n(\beta)
\]
guarantees existence of the adjusted EL solution and maintains the \(\chi^2_k\) limit for the adjusted statistic [1604.06170]. For the tail index of a heavy-tailed distribution, adjusted EL uses the pseudo-observation
\[
y_{k_n+1}(\gamma)=-a_n(\hat\gamma_n-\gamma),
\]
with \(a_n=o(k_n^{2/3})\), and the paper recommends \(a_n=19/12\) because the top-order log-spacings are approximately \(\mathrm{Exp}(\gamma)\) and this choice yields a Bartlett-like correction [1904.08420].

Jackknife variants attack the same problem through pseudo-values. For probability weighted moments, the jackknife pseudo-values are
\[
\widehat V_k=n\widehat\beta_r-(n-1)\widehat\beta_{r,k},
\]
and adjusted JEL adds
\[
\widehat V_{n+1}=-\frac{a_n}{n}\sum_{k=1}^n \widehat V_k,
\]
after which \(-2l_1(\beta_r)\to\chi^2(1)\) under \(a_n=o_p(n^{2/3})\) [1807.04450]. Generalized Lorenz inference develops three modified EL approaches—adjusted EL, transformed EL, and transformed adjusted EL—and shows that the modified ratio statistics follow scaled chi-squared limits; in simulations, TAEL consistently delivers the highest coverage probability, whereas AEL gives the shortest intervals on average [2304.04124].

A distinct line of development replaces ad hoc bias subtraction by robust weights that incorporate both bias correction and the additional variability from estimated bias. In nonparametric regression and regression discontinuity designs, robust empirical likelihood defines weights such as
\[
W_{i,h,b}^\star(x_0)
=
W_{i,h}(x_0)
-
W_{i,2,2,b}(x_0)\Bigl\{n^{-1}b^{-2}\sum_{k=1}^n W_{k,h}(x_0)(X_k-x_0)^2\Bigr\},
\]
or the difference-based analogue \(W_{i,h,b}^{\diamond}(x_0)\), and then uses \(g_i(\theta)=W_{i,h,b}^{\mathrm{rob}}(x_0)[Y_i-\theta]\) in the EL constraint [2504.01535]. The resulting robust EL ratio satisfies Wilks’ theorem under \(h/b\to\kappa\in[0,1]\), avoiding the undersmoothing otherwise required by conventional local-linear EL [2504.01535].

## 5. Extensions beyond i.i.d. low-dimensional models

The empirical likelihood ratio confidence interval has been extended to dependent, censored, survey-weighted, and high-dimensional settings by altering the estimating equations rather than abandoning EL calibration. In the frequency domain, empirical likelihood based on periodogram ordinates uses
\[
\psi_j(\theta)=2\pi G_\theta(\lambda_j)I_n(\lambda_j)-M
\]
at Fourier frequencies and yields
\[
I_n(\theta)=-4\log R(\theta)\Rightarrow \chi^2_q
\]
for both short- and long-range dependence, provided the spectral estimating functions satisfy the stated growth and smoothness conditions near zero [0708.0197]. In stationary and long-memory time-series models, Whittle score-like equations and their adjusted variants similarly lead to asymptotic \(\chi^2\) confidence regions for ARMA and ARFIMA parameters [1602.09128, 1604.06170].

For weakly dependent multivariate data, blockwise empirical likelihood replaces raw estimating functions by block averages \(T_i(\theta)\). Adjusted blockwise EL adds the pseudo-block
\[
T_{Q+1}(\theta)=-a\bar T(\theta),
\]
which removes the finite-sample convex-hull upper bound on coverage and preserves the \(\chi^2_q\) limit of the calibrated statistic \(-2\,SF\,\mathrm{ABELRQ}(\theta)\); with a suitable high-order choice of \(a\), the coverage error improves from \(O(n^{-2/3})\) to \(O(n^{-5/6})\) [1912.07699].

High-dimensional inference requires a different modification. When \(\theta=(\theta_M,\theta_{M^c})\) includes a high-dimensional nuisance component, transformed moments
\[
f^{A_n}(X;\theta)=A_n g(X;\theta)
\]
are constructed so that the impact of estimating \(\theta_{M^c}\) becomes asymptotically negligible [1805.10742]. The resulting EL ratio \(\ell^*_{A_n}(\theta_M)\) satisfies
\[
\ell^*_{A_n}(\theta_{0,M})\to_d \chi^2_m
\]
when \(m\) is fixed, and
\[
\frac{\ell^*_{A_n}(\theta_{0,M})-m}{\sqrt{2m}}\to_d N(0,1)
\]
when \(m\to\infty\) under the stated rate conditions [1805.10742]. This preserves ELR-based confidence regions in regimes where classical profile EL would fail because nuisance estimation is slower than \(n^{-1/2}\) [1805.10742].

Right-censored lifetime data and complex surveys illustrate two other strategies. Under right censoring, the EL is built from influence-function estimating equations \(W_{ni}(\theta)\) derived from Kaplan–Meier integrals, and the resulting log-EL ratio converges to \(\chi^2_1\) without any unknown scale parameter [1203.5955]. Under unequal-probability complex survey sampling, jackknife pseudo-values \(V_i=nT_n-(n-1)T_{(i)}\) are inserted into a weighted pseudo-EL objective with effective sample-size scaling \(n^*=n/\mathrm{deff}\), giving \(-2\log R(\theta_0)\Rightarrow\chi^2_q\) under design-based conditions, with or without auxiliary information [2303.09737].

## 6. Applications and empirical behavior

Empirical likelihood ratio confidence intervals have been used for treatment contrasts in blocked experiments, optimal values and optimality gaps in sample average approximation, off-policy values in contextual bandits, regression coefficients under spatial dependence, tail parameters of heavy-tailed laws, and inequality or poverty functionals [2112.09206, 1604.02573, 1906.03323, 1808.08793, 1008.3229, 1707.04998, 2603.17327]. In blocked experiments with highly unbalanced or incomplete block designs, EL constructs simultaneous confidence intervals for pairwise treatment differences without explicit covariance specification; in the clothianidin application, NB and AMC agreed on significant differences except Fungicide vs Low, whereas HBW differed because the parametric assumptions were violated [2112.09206]. In contextual bandits, the EL confidence interval for policy value is formulated as a low-dimensional convex optimization problem using the moments \(E[w]=1\) and \(E[wr]=v\), and the reported empirical behavior is that the EL interval is much tighter than the binomial interval and has near-nominal coverage, whereas the asymptotic Gaussian interval undercovers [1906.03323].

In heavy-tail problems, empirical likelihood regions for \((\gamma,\sigma_n)\) based on Zhang’s estimating equations are reported to perform better than Wald-type regions, especially those derived from the asymptotic normality of the maximum likelihood estimators [1008.3229]. For the tail index \(\gamma\), adjusted EL outperforms normal approximation in terms of coverage probability and interval length, and the paper recommends defaulting to AEL with \(a_n=19/12\) [1904.08420]. For spatial error linear models, ELR statistics are constructed from linear and quadratic estimating equations and shown to converge to chi-squared limits, which are then used to form confidence regions for \((\beta,\rho,\sigma^2)\) and profiled intervals for individual coefficients [1808.08793].

Inequality and poverty applications show the same pattern. For generalized Lorenz ordinates, TAEL provides the highest coverage probability, AEL the shortest intervals, and all interval lengths increase with \(t\) while decreasing with \(n\) [2304.04124]. For S-Gini indices, JEL intervals have better or comparable coverage and shorter lengths than BCEL and bootstrap-\(t\) intervals across exponential, Pareto, and lognormal designs [1707.04998]. For the Sen and Sen–Shorrocks–Thon indices, EL and JEL-based intervals achieve substantially better coverage than normal or Wald intervals, especially in small samples or heavy-tailed settings; the empirical illustrations based on PSID and CPHS data use JEL confidence intervals to compare poverty intensity across years and states [2603.17327].

Taken together, these developments show that the empirical likelihood ratio confidence interval is not a single formula but a family of inversion procedures built around the same core geometry: a nonparametric likelihood over empirical weights, constrained by moment equations that encode the target parameter. The principal methodological differences across the literature concern how those constraints are chosen, profiled, adjusted, jackknifed, blocked, or robustified so that the resulting ELR statistic retains a usable chi-squared calibration in the presence of nuisance parameters, dependence, bias, incomplete designs, or nonstandard sampling schemes [2112.09206, 1912.07699, 2504.01535].

Source: https://www.emergentmind.com/topics/empirical-likelihood-ratio-confidence-interval