Papers
Topics
Authors
Recent
Search
2000 character limit reached

Turnbull NPMLE for Interval-Censored Data

Updated 10 July 2026
  • Turnbull NPMLE is a nonparametric likelihood estimator that computes the unknown distribution function from interval-censored data via self-consistent EM updates.
  • It partitions data into Turnbull intervals where probability mass is allocated solely on observed censoring boundaries, simplifying the infinite-dimensional optimization.
  • Modern extensions leverage this framework in semiparametric regression, censored density estimation, and multi-state models, highlighting its foundational role in incomplete-data likelihood analysis.

Searching arXiv for recent and foundational papers relevant to Turnbull NPMLE and its modern extensions. arxiv_search(query="Turnbull NPMLE interval-censored data nonparametric maximum likelihood estimator", max_results=10, sort_by="relevance") Turnbull’s nonparametric maximum likelihood estimator (NPMLE) is the canonical likelihood-based estimator of an unknown distribution function from interval-censored data. In its classical form, one observes, for each subject, only an interval known to contain an unobserved event time, and estimates the distribution function by maximizing the interval-censoring likelihood over all cumulative distribution functions. The resulting estimator is nonparametric, likelihood-based, and discrete, with mass concentrated on intervals determined by the observed censoring structure. In contemporary statistical research, Turnbull’s estimator also serves as a template for broader NPMLE constructions in semiparametric and shape-constrained models, including generalized linear models with interval-censored covariates, censored log-concave density estimation, and interval-censored multi-state regression (Toloba et al., 13 Jan 2026, Duembgen et al., 2013, Gu et al., 2022).

1. Definition and basic likelihood structure

In the classical interval-censoring setting, one observes for subject ii an interval Li,Ri\lfloor L_i, R_i \rfloor such that an unobserved event time TiT_i satisfies

TiLi,Ri.T_i \in \lfloor L_i, R_i \rfloor.

The inferential target is the distribution function

F(t)=Pr(Tt).F(t) = \Pr(T \le t).

Under independent interval censoring, the likelihood is

L(F)=i=1nPr(TiLi,Ri)=i=1n{F(Ri)F(Li)},L(F) = \prod_{i=1}^n \Pr(T_i \in \lfloor L_i, R_i \rfloor) = \prod_{i=1}^n \{ F(R_i) - F(L_i)\},

with the obvious modifications for left-, right-, or exactly observed times (Toloba et al., 13 Jan 2026).

Turnbull’s NPMLE F^Tb\widehat F^{\rm Tb} is the maximizer of this likelihood over all cdfs. In general NPMLE terminology, it is a maximizer of likelihood over an infinite-dimensional model P\mathcal P of probability measures. A general formulation is

P^nargmaxPPi=1np(Zi),\hat P_n \in \arg\max_{P \in \mathcal{P}} \prod_{i=1}^n p(Z_i),

where p=dP/dμp=dP/d\mu is a density with respect to a reference measure Li,Ri\lfloor L_i, R_i \rfloor0 (Yara et al., 2024). Turnbull’s estimator is a specific instance in which Li,Ri\lfloor L_i, R_i \rfloor1 is the set of all distribution functions on Li,Ri\lfloor L_i, R_i \rfloor2, restricted only by monotonicity and total mass constraints.

A defining property is that the estimator is discrete and places mass only on a finite set induced by the observed censoring intervals. In one modern exposition, this is stated as follows: the classical Turnbull NPMLE for interval-censored data is of this form, and “the estimator is a discrete distribution function with jumps at a subset of the observed interval endpoints” (Yara et al., 2024). In a more refined interval-based description, the NPMLE places mass on the so-called Turnbull intervals, which are the maximal intersections of the observed censoring intervals (Toloba et al., 13 Jan 2026).

2. Turnbull intervals, self-consistency, and EM representation

Turnbull’s central structural insight is that the likelihood can be reduced to a finite-dimensional optimization over masses assigned to data-determined support intervals. If Li,Ri\lfloor L_i, R_i \rfloor3 denotes the Turnbull intervals, then the estimator is characterized by masses

Li,Ri\lfloor L_i, R_i \rfloor4

Let

Li,Ri\lfloor L_i, R_i \rfloor5

The self-consistent equations are

Li,Ri\lfloor L_i, R_i \rfloor6

These equations can be interpreted as EM updates, with the unobserved event times treated as latent data (Toloba et al., 13 Jan 2026).

The corresponding EM algorithm alternates between conditional allocation of each observation across the support intervals and normalization of the resulting expected counts. The E-step computes

Li,Ri\lfloor L_i, R_i \rfloor7

and the M-step updates

Li,Ri\lfloor L_i, R_i \rfloor8

This is equivalent to the self-consistent equations above (Toloba et al., 13 Jan 2026).

This EM interpretation is important beyond the original problem. Later work repeatedly generalizes Turnbull’s architecture—discrete support on a data-dependent partition, latent allocation, and mass updates—to more complex semiparametric likelihoods. A plausible implication is that Turnbull’s enduring influence lies less in a single estimator than in a reusable computational grammar for interval-censored likelihoods.

3. Relation to general NPMLE theory

Turnbull’s estimator is a paradigmatic example of NPMLE in an infinite-dimensional model. Modern NPMLE theory emphasizes several common ingredients: an infinite-dimensional model Li,Ri\lfloor L_i, R_i \rfloor9, likelihood or log-likelihood as objective, approximation error, estimation error, empirical process methods such as bracketing entropy and covering numbers, and convergence metrics including Hellinger distance, total variation, TiT_i0, and sometimes KL divergence (Yara et al., 2024).

A recent general treatment states that Turnbull’s estimator is “precisely of this form” and explicitly identifies it as a member of the same NPMLE family as deep-network-based logistic estimators (Yara et al., 2024). In that framework, the parameter need not be a cdf; it may be any infinite-dimensional conditional distribution model. What remains invariant is likelihood maximization over a nonparametric class.

This perspective is particularly relevant for convergence analysis. In nonparametric logistic regression, direct KL analysis can fail because the estimator may assign zero probability where the truth is positive, making KL divergence infinite. To avoid this, recent work studies Hellinger risk instead and derives oracle inequalities using empirical process arguments (Yara et al., 2024). That work explicitly notes an analogy with Turnbull-type problems: if one can control the relevant bracketing entropy for the feasible class, an oracle inequality of the same flavor can be obtained.

This suggests a conceptual placement of Turnbull’s estimator within a larger theory of likelihood-based estimation under incomplete observation. The classical estimator is not merely a survival-analysis artifact; it is a prototype of a broad NPMLE principle in which inference is driven by likelihood geometry and complexity control rather than finite-dimensional parametrization.

4. Extensions to semiparametric regression with interval-censored covariates

A major recent extension uses Turnbull’s estimator as the nonparametric component of semiparametric generalized linear models with one interval-censored covariate TiT_i1 (Toloba et al., 13 Jan 2026). There, the outcome TiT_i2 and fully observed covariates TiT_i3 are observed, while

TiT_i4

The GLM has linear predictor

TiT_i5

with conditional density

TiT_i6

where TiT_i7 and TiT_i8 (Toloba et al., 13 Jan 2026).

The observed-data likelihood is

TiT_i9

where TiLi,Ri.T_i \in \lfloor L_i, R_i \rfloor.0 is estimated nonparametrically (Toloba et al., 13 Jan 2026).

The key innovation is the augmented Turnbull estimator. Instead of the classical Turnbull intervals, the support TiLi,Ri.T_i \in \lfloor L_i, R_i \rfloor.1 is partitioned using all unique left and right endpoints of the observed intervals, producing augmented Turnbull intervals TiLi,Ri.T_i \in \lfloor L_i, R_i \rfloor.2. Defining

TiLi,Ri.T_i \in \lfloor L_i, R_i \rfloor.3

the estimator is characterized by TiLi,Ri.T_i \in \lfloor L_i, R_i \rfloor.4-dependent self-consistent equations,

TiLi,Ri.T_i \in \lfloor L_i, R_i \rfloor.5

under a piecewise-uniform assumption within each interval TiLi,Ri.T_i \in \lfloor L_i, R_i \rfloor.6 (Toloba et al., 13 Jan 2026).

The paper proves that, for fixed TiLi,Ri.T_i \in \lfloor L_i, R_i \rfloor.7, any cdf TiLi,Ri.T_i \in \lfloor L_i, R_i \rfloor.8 whose interval masses solve the self-consistent equations is an NPMLE of TiLi,Ri.T_i \in \lfloor L_i, R_i \rfloor.9 for the joint likelihood. It further establishes consistency of F(t)=Pr(Tt).F(t) = \Pr(T \le t).0, consistency of F(t)=Pr(Tt).F(t) = \Pr(T \le t).1, and asymptotic normality of F(t)=Pr(Tt).F(t) = \Pr(T \le t).2 under regularity conditions, with covariance estimated by observed information (Toloba et al., 13 Jan 2026).

The structural relation to classical Turnbull estimation is direct.

Feature Classical Turnbull Augmented Turnbull in GLMs
Latent variable Event time F(t)=Pr(Tt).F(t) = \Pr(T \le t).3 Covariate F(t)=Pr(Tt).F(t) = \Pr(T \le t).4
Likelihood contribution F(t)=Pr(Tt).F(t) = \Pr(T \le t).5 F(t)=Pr(Tt).F(t) = \Pr(T \le t).6
Support partition Turnbull intervals Augmented Turnbull intervals
Update mechanism Self-consistency / EM Weighted self-consistency / EM-like

A central methodological point is that ignoring the dependence of F(t)=Pr(Tt).F(t) = \Pr(T \le t).7 on the outcome model can lead to inconsistent estimation of the regression parameter F(t)=Pr(Tt).F(t) = \Pr(T \le t).8 (Toloba et al., 13 Jan 2026). Thus the augmented estimator is not a cosmetic variation; it changes the inferential target from the censoring-only distribution of F(t)=Pr(Tt).F(t) = \Pr(T \le t).9 to a nonparametric distribution estimate that is joint-likelihood compatible with the regression model.

5. Shape-constrained and censored-density variants

Turnbull’s framework has also been extended to shape-restricted density estimation under interval-censoring, right-censoring, and binned observation (Duembgen et al., 2013). In that setting, the unknown distribution L(F)=i=1nPr(TiLi,Ri)=i=1n{F(Ri)F(Li)},L(F) = \prod_{i=1}^n \Pr(T_i \in \lfloor L_i, R_i \rfloor) = \prod_{i=1}^n \{ F(R_i) - F(L_i)\},0 on L(F)=i=1nPr(TiLi,Ri)=i=1n{F(Ri)F(Li)},L(F) = \prod_{i=1}^n \Pr(T_i \in \lfloor L_i, R_i \rfloor) = \prod_{i=1}^n \{ F(R_i) - F(L_i)\},1 may have mass

L(F)=i=1nPr(TiLi,Ri)=i=1n{F(Ri)F(Li)},L(F) = \prod_{i=1}^n \Pr(T_i \in \lfloor L_i, R_i \rfloor) = \prod_{i=1}^n \{ F(R_i) - F(L_i)\},2

and an absolutely continuous part with subdensity

L(F)=i=1nPr(TiLi,Ri)=i=1n{F(Ri)F(Li)},L(F) = \prod_{i=1}^n \Pr(T_i \in \lfloor L_i, R_i \rfloor) = \prod_{i=1}^n \{ F(R_i) - F(L_i)\},3

where L(F)=i=1nPr(TiLi,Ri)=i=1n{F(Ri)F(Li)},L(F) = \prod_{i=1}^n \Pr(T_i \in \lfloor L_i, R_i \rfloor) = \prod_{i=1}^n \{ F(R_i) - F(L_i)\},4 is concave and

L(F)=i=1nPr(TiLi,Ri)=i=1n{F(Ri)F(Li)},L(F) = \prod_{i=1}^n \Pr(T_i \in \lfloor L_i, R_i \rfloor) = \prod_{i=1}^n \{ F(R_i) - F(L_i)\},5

The observed-data log-likelihood is

L(F)=i=1nPr(TiLi,Ri)=i=1n{F(Ri)F(Li)},L(F) = \prod_{i=1}^n \Pr(T_i \in \lfloor L_i, R_i \rfloor) = \prod_{i=1}^n \{ F(R_i) - F(L_i)\},6

where

L(F)=i=1nPr(TiLi,Ri)=i=1n{F(Ri)F(Li)},L(F) = \prod_{i=1}^n \Pr(T_i \in \lfloor L_i, R_i \rfloor) = \prod_{i=1}^n \{ F(R_i) - F(L_i)\},7

(Duembgen et al., 2013).

If one removes the log-concavity restriction and allows arbitrary distributions, this recovers the classical Turnbull-type censored likelihood. The difference is the feasible set: Turnbull maximizes over all cdfs, whereas the shape-restricted estimator maximizes over the smaller class of distributions with log-concave density. The resulting estimator is continuous with piecewise-exponential density, rather than discrete with arbitrary jumps (Duembgen et al., 2013).

An EM algorithm is developed for approximate computation. The E-step constructs a pseudo-sample measure L(F)=i=1nPr(TiLi,Ri)=i=1n{F(Ri)F(Li)},L(F) = \prod_{i=1}^n \Pr(T_i \in \lfloor L_i, R_i \rfloor) = \prod_{i=1}^n \{ F(R_i) - F(L_i)\},8, representing the expected distribution of latent finite event times under the current fit. The M-step then replaces Turnbull’s unrestricted mass allocation by a log-concave density MLE step, together with an update for the cure parameter L(F)=i=1nPr(TiLi,Ri)=i=1n{F(Ri)F(Li)},L(F) = \prod_{i=1}^n \Pr(T_i \in \lfloor L_i, R_i \rfloor) = \prod_{i=1}^n \{ F(R_i) - F(L_i)\},9 (Duembgen et al., 2013). This is a constrained analogue of Turnbull’s EM: the latent allocation logic remains, but the optimization is projected onto a shape-restricted function class.

The paper proves existence of the estimator under mild conditions, describes its shape properties, and establishes consistency results. In simulation comparisons, the log-concave estimator yields smoother survival estimates and lower sup-norm error than Turnbull’s step-function estimator when log-concavity is appropriate (Duembgen et al., 2013). This does not refute Turnbull’s estimator; rather, it shows how shape restrictions trade robustness of model class for efficiency and smoothness under structural assumptions.

6. Multi-state and semiparametric process generalizations

Another broad generalization treats interval-censored multi-state data using semiparametric proportional intensity models with random effects (Gu et al., 2022). Here the event history is a finite-state process with feasible transitions F^Tb\widehat F^{\rm Tb}0, and the transition intensities are

F^Tb\widehat F^{\rm Tb}1

where the baseline intensities F^Tb\widehat F^{\rm Tb}2 are modeled nonparametrically (Gu et al., 2022).

Turnbull’s influence appears in the representation of each baseline cumulative intensity as a step function over the grid of unique examination times: F^Tb\widehat F^{\rm Tb}3 with nonnegative jump sizes F^Tb\widehat F^{\rm Tb}4 (Gu et al., 2022). This is the direct analogue of assigning mass to support intervals in classical interval-censored survival.

The observed-data likelihood integrates over random effects and uses product integrals of transition probability matrices. To maximize it, the paper develops a stable EM algorithm based on latent Poisson variables F^Tb\widehat F^{\rm Tb}5, which encode transition counts at grid points. The M-step updates the nonparametric jumps explicitly as

F^Tb\widehat F^{\rm Tb}6

a count-over-exposure form that is structurally analogous to Turnbull’s expected-count update (Gu et al., 2022).

The paper states that it uses Turnbull’s method to pre-estimate jump sizes and remove time points with estimates smaller than a threshold of order F^Tb\widehat F^{\rm Tb}7, thereby reducing the computational grid (Gu et al., 2022). It proves strong consistency of the NPMLE and asymptotic normality of the finite-dimensional components, with covariance attaining the semiparametric efficiency bound and estimated by profile likelihood (Gu et al., 2022).

In a special case—two states, one absorbing transition, no covariates, no random effects—the model reduces to a Turnbull-type interval-censored event-time problem (Gu et al., 2022). This reduction clarifies that multi-state semiparametric NPMLEs are not separate constructions but layered extensions of the same underlying likelihood principle.

7. Metrics, misconceptions, and modern perspective

A common misconception is that Turnbull’s estimator is only a survival-analysis tool for a narrow censoring setup. Recent work contradicts this restricted view. Turnbull’s NPMLE is used as a building block in GLMs with interval-censored covariates (Toloba et al., 13 Jan 2026), as an unconstrained benchmark for censored log-concave density estimation (Duembgen et al., 2013), and as a conceptual and computational precursor for multi-state semiparametric intensity models (Gu et al., 2022). A further modern perspective places it inside the general NPMLE framework that also includes neural-network-based likelihood estimators (Yara et al., 2024).

A second misconception is that the natural metric for analyzing NPMLEs is always KL divergence. In nonparametric settings, KL may diverge easily because the estimator can assign zero probability where the truth is positive. Recent work therefore advocates direct analysis in Hellinger distance and derives risk bounds in that metric under mild assumptions (Yara et al., 2024). The same work explicitly notes that this style of oracle inequality is conceptually parallel to what one would seek for Turnbull’s estimator, with appropriate entropy calculations (Yara et al., 2024). This suggests that the most stable theoretical language for Turnbull-type estimators may often be Hellinger-based rather than KL-based.

A third misconception is that Turnbull’s discreteness is an incidental computational artifact. In fact, discreteness is an intrinsic consequence of likelihood maximization over all cdfs under interval censoring. Shape-constrained variants modify that property only by shrinking the feasible model class, not by altering the underlying likelihood logic (Duembgen et al., 2013).

From a contemporary viewpoint, Turnbull’s NPMLE occupies a central place in the theory of incomplete-data likelihood. It combines three features that remain methodologically active: exact likelihood under censoring, nonparametric flexibility, and an EM-compatible finite-support representation. Modern extensions differ in target parameter, structural constraints, and asymptotic theory, but they continue to inherit the core Turnbull architecture.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Turnbull Nonparametric Maximum Likelihood Estimator (NPMLE).