Likelihood-Based Inference (LBI)
- Likelihood-Based Inference (LBI) is a statistical framework that utilizes the data's likelihood function for parameter estimation, testing, model comparison, and prediction.
- It incorporates classical methods such as maximum likelihood estimation and likelihood-ratio tests while extending to hierarchical, composite, and simulation-based approaches for complex models.
- Recent advancements include likelihood-preserving embeddings and Monte Carlo techniques that offer prior-free decision rules and efficient computation in high-dimensional and latent variable settings.
Likelihood-based inference (LBI) treats the observed data through the likelihood function , or equivalently the log-likelihood , and uses relative values of that function for estimation, testing, model comparison, and prediction (Reid, 2013). In the formulation associated with the likelihood principle, all information about that can be obtained from an observation is contained in the likelihood function up to a proportional constant (Giang et al., 2012). Contemporary usage spans classical maximum-likelihood and likelihood-ratio procedures, profile and adjusted likelihoods for nuisance-parameter problems, hierarchical likelihoods for models with unobservables, exact Monte Carlo likelihood methods for stochastic processes, and synthetic-likelihood constructions for simulation-based inference when the simulator likelihood is intractable [(Lee et al., 2010); (Gonçalves et al., 2017); (Glaser et al., 2022)].
1. Foundational concepts and the likelihood principle
In the classical parametric setup, has joint density or mass for . The likelihood is defined up to proportionality by , and the log-likelihood is . The score is , the Fisher information is , and the observed information is 0 (Reid, 2013).
The likelihood principle asserts that if two experiments produce data 1 and 2 whose likelihood functions are proportional,
3
then no inference about 4 should depend on which experiment was performed; only the likelihood ratio matters (Giang et al., 2012). Giang and Shenoy formulate an “extended likelihood”
5
and for any 6,
7
Their Theorem 1 shows that 8 is a possibility measure, so inference based on 9 is max-decomposable and normalized to 0 (Giang et al., 2012).
That possibility-theoretic construction leads to a prior-free decision rule. Given an action 1 and data 2, one induces a possibilistic lottery 3 on consequences, and Theorem 2 yields a unique qualitative expected utility representation. The resulting likelihood-based decision rule is
4
which uses only the normalized likelihood and a utility assessment in terms of binary utilities, rather than a prior distribution 5 (Giang et al., 2012). Within the data provided, this positions LBI between Bayesian procedures, which require a prior, and minimax procedures, which optimize worst-case risk without using the observed-data likelihood.
2. Classical inferential constructions
Under standard regularity conditions, the maximum likelihood estimator 6 solves 7 and satisfies
8
which yields the familiar quadratic approximation to confidence regions and the asymptotic normality that underlies likelihood-based testing (Reid, 2013). The central classical test statistics are the likelihood-ratio, Wald, and score statistics: 9 each with an approximate 0 null distribution in regular settings (Reid, 2013).
LBI also supports prediction, not only parameter inference. For iid data 1 with density 2 and an independent future draw 3, the likelihood-ratio statistic for testing 4 is
5
where 6 (Tian et al., 2021). Inversion yields a nominal 7 prediction region
8
with 9 chosen from an exact pivotal law, a Wilks-type 0 approximation, or a parametric bootstrap (Tian et al., 2021).
When 1 is pivotal, the LR-based prediction interval is exact and often recovers standard constructions. The data supplied list the known-2 normal case, the unknown-3 normal case, the exponential model with unknown mean, the 4 model, and more generally any location-scale family (Tian et al., 2021). When no true pivot exists, the same framework still yields asymptotically correct or bootstrap-calibrated intervals for continuous and discrete settings, including Gamma prediction, Binomial and Poisson prediction, and within-sample censored-life problems (Tian et al., 2021). This makes likelihood ratios a unifying device across estimation, testing, and prediction.
3. Nuisance parameters, conditioning, and higher-order likelihood theory
When 5 contains a parameter of interest 6 and nuisance parameter 7, the profile log-likelihood is
8
where 9 maximizes 0 for fixed 1. The corresponding profile likelihood-ratio statistic,
2
is asymptotically 3 under regularity (Reid, 2013). In finite samples, profile likelihood may under-correct for nuisance-parameter estimation, which motivates modified profile likelihoods of the form
4
and higher-order pivots such as
5
with 6 (Reid, 2013).
DiCiccio, Kuffner, Young and Zaretzki analyze the stability of likelihood-based pivots under conditioning on ancillary statistics. For scalar 7, the signed root of the profile-likelihood ratio statistic,
8
with 9, is stable to order 0: under a second-order local ancillary 1,
2
More generally, asymptotically normal pivots with the expansion
3
are stable to 4 provided the leading constant satisfies
5
These results are accompanied by sufficient conditions under which one-sided 6-values derived from two different pivots agree to 7 (DiCiccio et al., 2015).
For vector interest parameter 8, both the uncorrected likelihood-ratio statistic 9 and its Bartlett-corrected version 0 satisfy
1
so marginal approximations accurate to 2 also respect conditioning to that order (DiCiccio et al., 2015). The supplied material further states that pivots built using observed information automatically satisfy the key stability and uniqueness conditions, whereas those built from expected information generally do not (DiCiccio et al., 2015). In practice, this singles out likelihood-ratio and observed-information constructions as the default high-accuracy likelihood pivots.
4. Extensions to unobservables, complex dependence, and sampled structures
A major extension of LBI concerns models with latent variables or random effects. In hierarchical likelihood, or h-likelihood, the complete-data model is written as
3
and the h-likelihood is
4
This treats 5 on the same formal footing that ordinary likelihood treats 6 alone (Lee et al., 2010). Joint score equations 7 and 8 estimate fixed parameters and predict unobservables simultaneously, while the block-partitioned observed information
9
supports variance approximations. Laplace adjustment of 0 provides an approximate marginal log-likelihood for 1, and the full inverse joint Hessian gives an 2 approximation to the conditional mean-squared error of predictors for 3 (Lee et al., 2010).
The data supplied identify the Rasch model and a spatial disease-mapping model as direct examples. In those examples, the h-likelihood yields score equations for fixed effects, random effects, and dispersion parameters, and produces standard errors and CMSE-adjusted predictors without introducing priors on 4 (Lee et al., 2010). This extends classical likelihood reasoning to settings in which integration over 5 is analytically difficult.
Another extension arises when the full joint density is intractable but lower-order components are available. Composite likelihood forms
6
and the maximum composite-likelihood estimator is asymptotically normal with Godambe information 7, where 8 and 9 (Reid, 2013). Pairwise and conditional composite likelihoods, including Besag’s pseudo-likelihood for spatial lattice models, are listed in the supplied material as standard examples (Reid, 2013).
For random graph models, the supplied account emphasizes a distinct point: projectivity is not required for valid likelihood-based superpopulation inference. If a population graph 0 is generated but only an induced subgraph 1 is observed under an ignorable sampling design, the correct likelihood is the marginal probability of the observed subgraph under the full-population model,
2
obtained by summing over all completions 3. This definition does not require projectivity, and lack of projectivity does not invalidate the likelihood or preclude consistency of the MLE (Schweinberger et al., 2017). The non-projective sparse Erdős–Rényi example with 4 is explicitly given as a case where the MLE remains consistent (Schweinberger et al., 2017).
5. Intractable likelihoods, Monte Carlo likelihoods, and simulation-based inference
When 5 is unavailable in closed form but simulation from the model is feasible, LBI can proceed by unbiased or exact Monte Carlo likelihood evaluation. In inverse binomial sampling (IBS), for each observation 6 one simulates from the model until the simulated output matches 7. If 8 is the number of draws required and 9, then the per-trial estimator
00
satisfies 01, so
02
is a uniformly unbiased estimator of the log-likelihood (Opheusden et al., 2020). Its variance is
03
which is uniformly bounded; the supplied material also states that IBS is the unique unbiased estimator under the geometric sampling policy and therefore minimum-variance among unbiased estimators with that policy (Opheusden et al., 2020).
For discretely observed finite-activity jump-diffusions, exact Monte Carlo likelihood-based inference can be built from a complete-data representation. After the Lamperti transform, the complete-data log-likelihood contains drift integrals, jump contributions, and bridge terms. The observed-data likelihood is obtained by integrating out the latent bridge path, jump times, jump sizes, and jump counts. Poisson estimators provide unbiased Monte Carlo estimators for the intractable time-integral terms, and importance sampling on jump-diffusion bridges yields local Monte Carlo estimators with lower variance than global alternatives (Gonçalves et al., 2017). The supplied summary describes both a frequentist MCEM algorithm and a Bayesian Gibbs-within-Metropolis scheme using Barker’s formula, a Two-Coin algorithm, and exact bridge simulation; in that framework, the only sources of error are Monte Carlo error and convergence of EM or MCMC algorithms (Gonçalves et al., 2017).
A different line of work addresses simulation-based inference through learned synthetic likelihoods. In “Maximum Likelihood Learning of Unnormalized Models for Simulation-Based Inference,” the conditional likelihood is parameterized as an energy-based model
04
with 05 represented by a neural network (Glaser et al., 2022). Two variants are introduced. The amortized variant, AUNLE, minimizes a forward KL between the true joint 06 and a tilted EBM joint 07; at the optimum, the learned 08 matches 09 and 10 becomes constant in 11, so the amortized posterior is
12
which is singly intractable and can be handled with standard MCMC or variational inference (Glaser et al., 2022). The sequential variant, SUNLE, maximizes the conditional log-likelihood
13
leading to a doubly intractable posterior
14
for which the supplied material lists exchange MCMC and Doubly-Intractable VI (DIVI) as posterior-sampling strategies (Glaser et al., 2022). The empirical summaries state that AUNLE recovers both modes in a toy bimodal likelihood, is on par with NLE on several SBI benchmarks and better on Two-Moons, while SUNLE matches or exceeds SNLE/SNVI performance and yields tighter posteriors on the pyloric network model with a fraction of the simulation budget (Glaser et al., 2022).
Likelihood-based inference has also been used for latent state estimation in agent-based models. In the bounded-confidence model, one specifies a Bernoulli observation model
15
so that the full likelihood factorizes over times and agent pairs, and the log-likelihood is optimized over the latent trajectory or initial state by automatic differentiation through the ABM dynamics (Kolic et al., 22 Sep 2025). The supplied case study reports that LBI better recovers latent agent-level opinions than data assimilation, even under model mis-specification, while both methods perform comparably at the aggregate level under some parameter settings (Kolic et al., 22 Sep 2025).
6. Likelihood-preserving representations, empirical domains, and current limitations
The increasing use of learned representations has motivated a direct reformulation of LBI under compression. “Likelihood-preserving embeddings” are defined through the likelihood-ratio distortion
16
the worst-case additive error in log-likelihood differences induced by an embedding 17 (Akdemir, 27 Dec 2025). The Hinge Theorem states that if 18, then likelihood-ratio tests are asymptotically preserved, surrogate MLEs are asymptotically equivalent to full-data MLEs, and Bayes factors between any two models have log-error bounded by 19 (Akdemir, 27 Dec 2025). If pointwise error satisfies 20, then absolute quantities such as AIC, BIC, and posterior normalizing constants are preserved up to 21 (Akdemir, 27 Dec 2025).
The same work proves a No Free Lunch theorem: exact ratio preservation for the class of all distributions dominated by 22 requires an embedding that is 23-almost-everywhere injective, and for 24-parameter exponential families any embedding with 25 must have output dimension 26 (Akdemir, 27 Dec 2025). A constructive framework is given using an encoder 27, a dataset summary 28, and a decoder 29 trained either with a pointwise loss or directly on likelihood ratios. The data supplied state the bounds
30
together with a covering-number bound and a sample-complexity statement 31 (Akdemir, 27 Dec 2025).
The experimental summaries supplied for that paper are notable because they isolate inferential rather than predictive fidelity. In the Gaussian toy model, embedding dimension 32 yields 33 and 34, whereas 35 recovers exact sufficiency; in the Cauchy case, 36 and 37 decay smoothly but never reach zero; in a synthetic GMM, a 16-dimensional embedding of 1000 points in 38 achieves 39 per sample and correlation 40 between true and surrogate log-likelihoods; and in a distributed clinical trial, an exact sufficient summary with 41 numbers per site matches pooled analysis exactly while a targeted compression with 42 attains 43 of the pooled power (Akdemir, 27 Dec 2025). These results make likelihood preservation an explicit design criterion for representation learning.
A complementary empirical perspective appears in cosmological simulation-based inference for primordial non-Gaussianity. There, the likelihood-based pipeline assumes a multivariate Gaussian likelihood
44
with 45 fitted by a low-order polynomial and 46 obtained from simulation covariance estimates corrected by the Hartlap and Dodelson–Schneider factors (Alokda et al., 1 May 2026). The posterior 47 is then evaluated on a fine 1D grid under a uniform prior over the training range 48 (Alokda et al., 1 May 2026). Across 1000 realizations at the true value 49, the supplied results report unbiased posterior means, 50, mean skewness 51, and mean excess kurtosis 52, with conditional coverage slightly larger than nominal for all summary statistics considered (Alokda et al., 1 May 2026). At the same time, the study emphasizes that coverage-based diagnostics probe calibration only in an averaged sense and do not constrain posterior behavior at fixed parameter value; larger discrepancies appear in posterior kurtosis, and LBI underestimates the true posterior kurtosis variability across realizations (Alokda et al., 1 May 2026).
Taken together, the supplied literature presents LBI not as a single algorithm but as a family of inferential constructions organized around the likelihood function and its ratios. In regular low-dimensional models, it yields the standard machinery of MLEs, likelihood-ratio tests, profile likelihoods, and higher-order pivots [(Reid, 2013); (DiCiccio et al., 2015)]. In latent-variable and structured-sampling problems, it extends through h-likelihood, composite likelihood, and exact marginalization arguments [(Lee et al., 2010); (Schweinberger et al., 2017)]. In simulator-based and intractable-likelihood settings, it persists through unbiased Monte Carlo estimators, exact bridge-based methods, and learned synthetic likelihoods (Opheusden et al., 2020, Gonçalves et al., 2017, Glaser et al., 2022). Recent work on embeddings and on fixed-truth validation further indicates that preserving likelihood geometry, rather than only predictive performance or average coverage, is central to preserving inferential conclusions (Akdemir, 27 Dec 2025, Alokda et al., 1 May 2026).