Papers
Topics
Authors
Recent
Search
2000 character limit reached

Dependent Tail-Free Process Latent Ensembles

Updated 17 February 2026
  • The paper introduces the DTFP ensemble that applies a Bayesian nonparametric prior to assign input-dependent weights and decompose model selection uncertainty.
  • It employs structured variational inference with a CRPS calibration objective to produce well-calibrated predictive intervals and improve empirical coverage.
  • The method outperforms traditional ensemble techniques by adapting to heterogeneous input domains and smoothly varying model weights through a tree-structured latent process.

A dependent tail-free process (DTFP) latent ensemble is an adaptive, probabilistic ensemble learning methodology that assigns input-dependent, non-deterministic weights to base models through a Bayesian nonparametric prior, enabling interpretable decompositions of predictive and model-selection epistemic uncertainty. The DTFP prior ensures that model weights are functionally dependent on the input x\mathbf{x}, yielding smooth variations in weights across the feature domain, hierarchical grouping of base models, and coherent quantification of selection uncertainty. Calibration of predictive distributions is central, achieved through a variational inference strategy that directly penalizes miscalibration as measured by the continuous ranked probability score (CRPS), resulting in improved empirical coverage and accurate uncertainty assessment across diverse tasks (Liu et al., 2018).

1. Dependent Tail-Free Process Prior for Ensemble Weights

A DTFP prior defines a random measure μ:F×X[0,1]\mu: F \times X \rightarrow [0,1] over a collection of KK base models F={f^1,,f^K}F = \{\hat{f}_1, \ldots, \hat{f}_K\} and feature space XX such that k=1Kμ(f^k,x)=1\sum_{k=1}^K \mu(\hat{f}_k, x) = 1 for each xXx \in X. The DTFP is constructed via a tree-structured partition Π\Pi of FF, allowing for hierarchical model combinations. Each non-leaf node vv in this tree, with child nodes μ:F×X[0,1]\mu: F \times X \rightarrow [0,1]0, is associated with latent functions μ:F×X[0,1]\mu: F \times X \rightarrow [0,1]1 sampled i.i.d. from a Gaussian process prior μ:F×X[0,1]\mu: F \times X \rightarrow [0,1]2 and a sparsity parameter μ:F×X[0,1]\mu: F \times X \rightarrow [0,1]3.

Conditional weights are specified via a softmax transformation:

μ:F×X[0,1]\mu: F \times X \rightarrow [0,1]4

for each μ:F×X[0,1]\mu: F \times X \rightarrow [0,1]5. The overall ensemble weight assigned to leaf model μ:F×X[0,1]\mu: F \times X \rightarrow [0,1]6 is

μ:F×X[0,1]\mu: F \times X \rightarrow [0,1]7

where μ:F×X[0,1]\mu: F \times X \rightarrow [0,1]8 is the ancestor chain in μ:F×X[0,1]\mu: F \times X \rightarrow [0,1]9. This construction ties model weights across KK0, introducing smooth, data-adaptive dependencies.

2. Full Probabilistic Ensemble Model

Given training data KK1, a hierarchical probabilistic model is defined. The ensemble mean function is constructed as

KK2

where KK3 captures residual systematic uncertainty. Observations follow KK4.

The full joint model consists of random variables KK5 for all non-leaves, KK6 for all non-leaves, the residual GP KK7, and noise variance KK8. Priors are assigned as KK9, F={f^1,,f^K}F = \{\hat{f}_1, \ldots, \hat{f}_K\}0, F={f^1,,f^K}F = \{\hat{f}_1, \ldots, \hat{f}_K\}1 (e.g., log-normal), and F={f^1,,f^K}F = \{\hat{f}_1, \ldots, \hat{f}_K\}2 (e.g., inverse-Gamma or half-Cauchy).

The marginal predictive distribution for a new input F={f^1,,f^K}F = \{\hat{f}_1, \ldots, \hat{f}_K\}3 is

F={f^1,,f^K}F = \{\hat{f}_1, \ldots, \hat{f}_K\}4

3. Structured Variational Inference and Calibration Objective

Posterior inference is performed via a structured variational family F={f^1,,f^K}F = \{\hat{f}_1, \ldots, \hat{f}_K\}5, fully factorized as:

F={f^1,,f^K}F = \{\hat{f}_1, \ldots, \hat{f}_K\}6

where each F={f^1,,f^K}F = \{\hat{f}_1, \ldots, \hat{f}_K\}7 and F={f^1,,f^K}F = \{\hat{f}_1, \ldots, \hat{f}_K\}8 is a sparse GP-variational approximation, and F={f^1,,f^K}F = \{\hat{f}_1, \ldots, \hat{f}_K\}9, XX0 are fully-factorized log-normal.

The optimization objective balances regularization and calibration:

XX1

where the continuous ranked probability score (CRPS) for example XX2 and predictive CDF XX3 is:

XX4

with XX5 sampled i.i.d. from the predictive distribution at XX6. The XX7 component is optimized via the evidence lower bound (ELBO) and reparameterization for GPs and log-normals, while the CRPS gradient is estimated via the score-function estimator. Stochastic gradient optimizers such as Adam are used, with optional Rao–Blackwellization for variance reduction.

4. Interpretation and Uncertainty Quantification

The DTFP ensemble supports explicit decomposition of ensemble-level uncertainty:

  • Model-selection uncertainty: The posterior spread of XX8 quantifies uncertainty in which base model predominates at a given XX9.
  • Residual predictive uncertainty: The posterior spread of k=1Kμ(f^k,x)=1\sum_{k=1}^K \mu(\hat{f}_k, x) = 10 reflects irreducible uncertainty after model selection.

Credible predictive intervals are derived from samples k=1Kμ(f^k,x)=1\sum_{k=1}^K \mu(\hat{f}_k, x) = 11 drawn from k=1Kμ(f^k,x)=1\sum_{k=1}^K \mu(\hat{f}_k, x) = 12. For a given k=1Kμ(f^k,x)=1\sum_{k=1}^K \mu(\hat{f}_k, x) = 13,

k=1Kμ(f^k,x)=1\sum_{k=1}^K \mu(\hat{f}_k, x) = 14

with k=1Kμ(f^k,x)=1\sum_{k=1}^K \mu(\hat{f}_k, x) = 15. Empirical quantiles of k=1Kμ(f^k,x)=1\sum_{k=1}^K \mu(\hat{f}_k, x) = 16 yield calibrated credible intervals.

5. Empirical Evaluation and Case Studies

The DTFP approach was evaluated on both synthetic and real-world predictive tasks:

  • Synthetic nonlinear 1D regression: Data comprised k=1Kμ(f^k,x)=1\sum_{k=1}^K \mu(\hat{f}_k, x) = 17, with four RBF-kernel regressors as base models. The DTFP ensemble achieved RMSE k=1Kμ(f^k,x)=1\sum_{k=1}^K \mu(\hat{f}_k, x) = 18 and nearly exact k=1Kμ(f^k,x)=1\sum_{k=1}^K \mu(\hat{f}_k, x) = 19 interval coverage, outperforming simple averaging, stacking, and GAM methods, which incurred higher RMSE xXx \in X0 and exhibited under- or overconfident intervals.
  • Spatio-temporal PMxXx \in X1 fusion in New England: Three state-of-the-art exposure models predicted annual particle pollution across 43 monitors. Leave-one-out RMSE for the DTFP ensemble was xXx \in X2, compared to xXx \in X3 (average) and xXx \in X4–xXx \in X5 (stacking). Spatial maps of posterior xXx \in X6 highlighted spatial nonstationarity in the ensemble weights and increased model-selection uncertainty in regions with heterogeneous base predictions or sparse monitoring. The ensemble produced predictive uncertainties matching empirical variability, enabling well-calibrated xXx \in X7 and xXx \in X8 intervals (Liu et al., 2018).

6. Comparative Analysis and Scope

DTFP latent ensembles extend ensemble methodologies by modeling adaptive, input-dependent weights with a coherent Bayesian nonparametric prior. Unlike conventional ensembles with fixed weights, DTFP ensembles address variable base model accuracy across subgroups and explicitly quantify uncertainty both in model selection and prediction. Calibration, achieved through direct penalization of miscalibration (CRPS), distinguishes the approach from deterministic or likelihood-only ensemble constructions, which can yield overconfident or miscalibrated intervals.

A plausible implication is that the DTFP approach is particularly well-suited for applications with heterogeneous input domains and diverse model error profiles, where both predictive performance and credible quantification of selection uncertainty are critical.

7. Interpretations, Limitations, and Directions

The DTFP framework provides rigorously calibrated predictive inference and interpretable model weight learning, even in hierarchical or grouped ensemble scenarios. It enables the fusion of diverse models with spatially- or feature-varying reliability and has demonstrated efficacy in both controlled and real-world spatio-temporal tasks. Limitations include the computational challenges inherent in Gaussian process-based variational inference and the scalability of sampling-based credible intervals for high-dimensional xXx \in X9. Progress in sparse GP techniques and optimizing structured variational objectives is expected to further broaden the applicability of DTFP latent ensembles (Liu et al., 2018).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Dependent Tail-Free Process Latent Ensembles.