---
title: 'Likelihood Ratio Testing: Theory & Extensions'
url: https://www.emergentmind.com/topics/likelihood-ratio-testing-lrt
type: topic
---

# Likelihood Ratio Testing: Theory & Extensions

The likelihood ratio test (LRT) is a central statistical tool for hypothesis testing in parametric frameworks. It formalizes the comparison between nested models via the maximized likelihood under the null and alternative, and underpins the theory and practice of model selection, goodness-of-fit, high-dimensional inference, latent variable modeling, irregular settings, missing-data analysis, and the construction of universally valid tests. This article documents the theoretical foundations of the LRT, its asymptotic and finite-sample properties, its generalizations, its modern high-dimensional variants, and key practical and conceptual issues.

## 1. Definition and Classical Asymptotics

Let $X_1,\ldots,X_n$ be i.i.d. observations from a model $\{P_\theta : \theta \in \Theta \subset \mathbb{R}^k\}$, with log-likelihood $\ell_n(\theta) = \sum_{i=1}^n \log p_\theta(X_i)$. For testing $H_0: \theta \in \Theta_0$ versus $H_1: \theta \in \Theta\setminus\Theta_0$ where $\Theta_0\subset\Theta$, the LRT statistic is:

\[
\Lambda_n = -2 \log\frac{\sup_{\theta\in\Theta_0} L_n(\theta)}{\sup_{\theta\in\Theta} L_n(\theta)} = 2 \bigl( \ell_n(\hat\theta_1) - \ell_n(\hat\theta_0) \bigr)
\]

where $\hat\theta_1$ and $\hat\theta_0$ are the unrestricted and restricted maximum likelihood estimators, respectively. Under standard regularity conditions—including interiority of the null, differentiability in quadratic mean, invertible Fisher information, Lipschitz continuity, and consistency of MLEs—Wilks’ theorem holds:

\[
\Lambda_n \xrightarrow{d} \chi^2_{d}, \quad d=\dim(\Theta)-\dim(\Theta_0)
\]

as $n\to\infty$ [2008.03971]. This is pivotal: critical values can be drawn from the $\chi^2_d$ distribution, enabling asymptotic level control.

## 2. Extensions: Dimension-Restricted LRTs and Power

Dimension-restricted submodels arise when restricting alternatives to a strict submanifold $\Theta_1\subset\Theta$, of dimension $d<k$. The corresponding restricted LRT statistic is:

\[
\Lambda_\mathrm{res} = -2\log \frac{L(\theta_0)}{L(\tilde\theta)},\quad \tilde\theta=\operatorname{arg\,max}_{\theta\in\Theta_1}L(\theta)
\]

Under regularity, asymptotic null distribution becomes $\chi^2_d$. As per the “dimension-restricted LRT conjecture,” any restriction lowering the alternative's dimension improves (increases) Pitman asymptotic power against local alternatives. Explicitly, under $\theta_n = \theta_0 + h/\sqrt{n}$, both unrestricted and restricted statistics converge to noncentral chi-squared distributions with the same noncentrality parameter $\lambda$, but the power strictly increases as $d$ decreases [1608.00032, Theorem 1]:

\[
P(\chi^2_d(\lambda)>c_{d,\alpha}) > P(\chi^2_k(\lambda)>c_{k,\alpha}),\quad d<k
\]

Nevertheless, this guarantee is only asymptotic: in finite samples, counterexamples (e.g., multinomial models with Hardy–Weinberg constraints) demonstrate that a dimension restriction can, for certain alternatives, reduce power [1608.00032]. Thus, for small $n$ or discrete sample spaces, exact power analysis is essential.

## 3. LRT Beyond Regularity: Boundaries, Singularities, and Latent Variables

Wilks’ theorem may not hold if regularity conditions fail—e.g., the null is on a boundary, information is singular, or nuisance parameters are unidentifiable under $H_0$. Latent variable models (factor analysis, random effects) exemplify such irregularities [2008.03971]. In these cases, the limiting distribution of $\Lambda_n$ is typically not $\chi^2$ but a more complicated functional of Gaussian processes and tangent cones.

Chernoff–van der Vaart–Drton theory replaces the standard limit with:

\[
\Lambda_n \xrightarrow{d} \min_{\tau \in T_{\Theta_0}(\theta^*)} \|Z - I^{1/2} \tau\|^2
\]

where $T_{\Theta_0}(\theta^*)$ is the tangent cone to the null parameter space at $\theta^*$ and $Z\sim N(0,I)$. If the tangent cone is a subspace, the limit is $\chi^2$; if a convex cone, a mixture (“chi-bar squared,” $\bar\chi^2$) occurs. For boundary points with unidentifiable nuisance and singular information, LRT statistics (e.g., genetic linkage, mixture models) converge to suprema of $\bar\chi^2$-processes [2605.08471]. Proper inference requires characterizing tangent cones and, when necessary, computing quantiles numerically or using parametric bootstraps [2008.03971].

### Special Tables: Asymptotic LRT Law Types

| Scenario                | Limit of $\Lambda_n$                          | Reference Section      |
|-------------------------|-----------------------------------------------|-----------------------|
| Regular, interior null  | $\chi^2_{d}$                                 | Wilks' theorem        |
| Boundary/irregular      | Mixture ($\bar\chi^2$) via tangent cone      | Chernoff–Drton theory |

## 4. High-dimensional Settings and Corrected LRTs

In high-dimensional regimes ($p$, $k$, or group numbers comparable to $n$), classical $\chi^2$ approximations break down: LRT statistics diverge or have non-pivotal, non-Gaussian laws [1306.0254, 1206.0867, 1812.06894, 1302.3302, 1502.00384]. The solution is to recenter and rescale the test statistic—using random matrix theory and CLTs for spectral statistics—so that the limit becomes normal, not chi-squared.

For identity testing in covariance matrices:

\[
L_n = \frac{1}{p}{\rm tr}(S_n) - \frac{1}{p}\log|S_n| - 1 - d(y_n),\quad y_n=p/n<1
\]
has
\[
\frac{p L_n - \mu_n}{\sigma_n} \xrightarrow{d} N(0,1)
\]

where $\mu_n$, $\sigma_n^2$ are explicit in $p$, $n$ [1302.3302, 1502.00384]. For linear regression or MANOVA, analogous corrections apply [1206.0867, 1812.06894, 1905.10354], and testing simultaneous means and covariances similarly requires centering and scaling, with explicit formulas from random matrix theory [2403.05760].

Regularization, e.g., via shrinkage estimators, further improves power and variance control in high-dimensional covariance inference [1502.00384].

## 5. Modern LRT Generalizations: Universality and Irregularity

Universal inference develops LRT-based tests and confidence sets with exact, finite-sample guarantees, valid without any regularity conditions [2104.14676, 2510.23821]. The “split LRT” class splits the data, fits parameters on one subset, and evaluates likelihoods on the other: the resulting statistic is an e-variable, giving finite-sample valid tests. Aggregating over many splits (by averaging e-values) yields “universal” LRT confidence sets and tests. In standard situations (e.g., Gaussian mean), the universal LRT is typically slightly conservative, with confidence sets having at most 50% larger squared radius than classical LRT-based sets, and slightly reduced power [2104.14676]. For non-convex or set-identified nulls, universal LRTs can outperform standard LRT-based approaches (e.g., the "doughnut" annulus example).

## 6. LRTs with Incomplete Data and Imputation

With missing data and multiple imputation (MI), the construction of LRTs is nontrivial. Classic MI-based LRT combinators (Rubin, Meng–Rubin) can be non-invariant, non-monotonic, negative, and inconsistently estimate the fraction of missing information. The “stacked” MI LRT—analyzing the concatenation of all imputed data sets—solves all these issues: it is always nonnegative, reparametrization-invariant, and the associated missing-information estimator is consistent under both null and alternative [1711.08822]. The null distribution takes an $F$-approximation with denominator degrees of freedom determined by the finite-sample fraction of missing information, yielding a principled and implementable approach for likelihood-based inference with missing data.

## 7. Practical Considerations and Choosing LRT Variants

- **Classical $\chi^2$-based LRTs** are appropriate only under standard regularity and when $p, k \ll n$.
- **Corrected and regularized LRTs** are essential when dimension is not negligible relative to sample size; corrections must use high-dimensional central limit theorems and, where beneficial, shrinkage [1206.0867, 1502.00384, 1812.06894].
- **Universal/subsampled LRTs** provide robust, finite-sample valid inference even in highly irregular or misspecified models [2104.14676, 2510.23821].
- **Restrictions on alternatives** can improve asymptotic power but may have counterintuitive finite-sample effects; explicit power calculations (not LRT heuristics) are essential in small samples or discrete models [1608.00032].
- **Boundary and latent variable problems** require the use of $\chi^2$-mixture reference distributions derived from tangent cone geometry [2008.03971, 2605.08471].
- **Imputation-based LRTs** should use the stacking approach for principled missing data inference [1711.08822].

Careful attention to underlying assumptions, dimensionality, and regularity is mandatory for valid LRT-based inference in modern, complex, or high-dimensional data settings.

Source: https://www.emergentmind.com/topics/likelihood-ratio-testing-lrt