Papers
Topics
Authors
Recent
Search
2000 character limit reached

SurvDiff: Diffusion for Survival Analysis

Updated 12 July 2026
  • SurvDiff is a diffusion model designed to generate synthetic survival datasets by jointly modeling mixed-type covariates, event times, and censoring indicators.
  • It employs a dual-block approach with continuous diffusion for covariates and event times and discrete diffusion for categorical variables, using a survival-tailored loss that integrates weighted Cox partial likelihood.
  • SurvDiff sets a benchmark against methods like SDPM and SurvivalGAN by achieving superior fidelity and robust downstream survival performance, especially in high-censoring scenarios.

SurvDiff denotes a diffusion-based approach to survival analysis centered on right-censored time-to-event data. In the strict sense of the named model, "SurvDiff: A Diffusion Model for Generating Synthetic Data in Survival Analysis" is an end-to-end diffusion model specifically designed for generating synthetic survival data by jointly generating mixed-type covariates, event times, and right-censoring indicators, with a survival-tailored loss function that encodes the time-to-event structure and directly optimizes for downstream survival tasks (Brockschmidt et al., 26 Sep 2025). In a broader usage that appears in subsequent work, the label also serves as a reference point for “survival via diffusion” methods more generally; SDPM, for example, explicitly contrasts itself with SurvDiff and presents a conditional diffusion model for survival prediction rather than full synthetic-data generation (Kirpichenko et al., 21 May 2026).

1. Scope, terminology, and problem setting

Survival analysis models time-to-event outcomes such as metastasis, disease relapse, or patient death, under the complication that the event may not be observed for all patients because of dropout, loss to follow-up, or study end. A standard observed-data representation uses covariates, an observed time, and an event indicator. In the SurvDiff formulation, one conceptually distinguishes latent event and censoring times and then writes

Y=min⁡(Tevent,C),δ=1{Tevent≤C}.Y = \min(T^{\text{event}}, C), \quad \delta = \mathbf{1}\{T^{\text{event}} \le C\}.

The model itself uses TT for the observed time and EE for the indicator, and aims to learn the joint distribution

PX(cont),X(disc),T,E.P_{X^{(\text{cont})}, X^{(\text{disc})}, T, E}.

This target is broader than a conventional survival regressor, because it includes both the covariate distribution and the censoring mechanism, not only a survival curve conditional on covariates (Brockschmidt et al., 26 Sep 2025).

A useful terminological distinction separates three research directions that are sometimes conflated because of similar names or overlapping motivation.

Term What is modeled Primary use
SurvDiff p(X,T,E)p(X, T, E) Synthetic survival data generation
SDPM P(T,δ∣x)\mathbb{P}(T,\delta \mid \mathbf{x}) Individual survival prediction
Diffsurv Risk-score ordering under censoring Set-based survival ranking

This suggests a broader “SurvDiff-style” family of diffusion-based survival methods, but the named model SurvDiff remains specifically an unconditional generator of synthetic survival datasets, whereas SDPM is a conditional outcome diffusion model and Diffsurv is a differentiable sorting method rather than a diffusion model (Kirpichenko et al., 21 May 2026).

2. Generative formulation and diffusion architecture

SurvDiff is a score-based diffusion model tailored to survival analysis. It jointly models continuous covariates and observed time in a continuous block,

z=(x(cont),T)∈Rdcont+1,z = (x^{(\text{cont})}, T) \in \mathbb{R}^{d_{\text{cont}} + 1},

and discrete covariates together with the event indicator in a discrete block, z~=(X(disc),E)\tilde{z} = (X^{(\text{disc})}, E). Each discrete covariate is represented by a one-hot vector including an extra “mask” category, and the event indicator is treated as a discrete variable and embedded analogously (Brockschmidt et al., 26 Sep 2025).

For the continuous block, SurvDiff uses a variance-exploding SDE with

dz=f(z,u) du+g(u) dWu,f(z,u)≡0,\mathrm{d}z = f(z,u)\,\mathrm{d}u + g(u)\,\mathrm{d}W_u, \quad f(z,u) \equiv 0,

which yields the closed-form forward corruption

zu=z0+σcont(u) ε,ε∼N(0,I),z_u = z_0 + \sigma^{\text{cont}}(u)\,\varepsilon, \quad \varepsilon \sim \mathcal{N}(0, I),

and

TT0

For the discrete block, it uses masked categorical diffusion with retention probability TT1, so that

TT2

where TT3 is the mask vector. As TT4, the continuous variables approach isotropic Gaussian noise and the discrete variables converge to the masked state (Brockschmidt et al., 26 Sep 2025).

The reverse process alternates Gaussian denoising for TT5 and categorical unmasking for TT6. SurvDiff uses the TabDiff architecture as backbone: a Transformer/MLP-based denoising network TT7 that takes noisy inputs TT8 and outputs a continuous prediction TT9 and a discrete prediction EE0, together with a separate survival head EE1 that maps reconstructed covariates into a scalar risk score (Brockschmidt et al., 26 Sep 2025).

3. Survival-tailored objective and treatment of censoring

The core conceptual novelty of SurvDiff is that the diffusion objective is augmented with a survival-specific loss. For the continuous block, training uses the standard VE SDE score-matching objective,

EE2

while the discrete block is trained with the continuous-time ELBO for masked diffusion,

EE3

These combine into

EE4

The survival head then produces a scalar risk score EE5, and SurvDiff adds a weighted Cox partial negative log-likelihood,

EE6

where

EE7

Only uncensored events contribute directly to the numerator, while censored observations enter indirectly through the risk sets and remain in risk sets until their censoring time. The weighting scheme is sparsity-aware: EE8 so that late sparse events are exponentially down-weighted. The total objective is

EE9

with PX(cont),X(disc),T,E.P_{X^{(\text{cont})}, X^{(\text{disc})}, T, E}.0 calibrated adaptively during a warm-up phase (Brockschmidt et al., 26 Sep 2025).

A notable feature of this construction is that SurvDiff does not model a latent censoring time PX(cont),X(disc),T,E.P_{X^{(\text{cont})}, X^{(\text{disc})}, T, E}.1 explicitly. Instead, it learns the joint distribution of the observed time PX(cont),X(disc),T,E.P_{X^{(\text{cont})}, X^{(\text{disc})}, T, E}.2 and event indicator PX(cont),X(disc),T,E.P_{X^{(\text{cont})}, X^{(\text{disc})}, T, E}.3 alongside covariates. The event indicator is generated as part of the discrete block, the observed time as part of the continuous block, and the censoring structure is represented through the empirical relationship between PX(cont),X(disc),T,E.P_{X^{(\text{cont})}, X^{(\text{disc})}, T, E}.4, PX(cont),X(disc),T,E.P_{X^{(\text{cont})}, X^{(\text{disc})}, T, E}.5, and PX(cont),X(disc),T,E.P_{X^{(\text{cont})}, X^{(\text{disc})}, T, E}.6, together with the Cox-style survival loss. The paper does not impose explicit independence assumptions such as non-informative censoring in the generative model; rather, it implicitly learns whatever relationship exists in real data (Brockschmidt et al., 26 Sep 2025).

4. Sampling procedure, evaluation protocol, and empirical behavior

Once trained, SurvDiff generates synthetic samples by initializing the continuous block at PX(cont),X(disc),T,E.P_{X^{(\text{cont})}, X^{(\text{disc})}, T, E}.7, setting all discrete entries to the mask state, and then running the reverse diffusion process over a discretized schedule, such as 300 sampling steps, until PX(cont),X(disc),T,E.P_{X^{(\text{cont})}, X^{(\text{disc})}, T, E}.8. The output is a synthetic tuple

PX(cont),X(disc),T,E.P_{X^{(\text{cont})}, X^{(\text{disc})}, T, E}.9

and optional administrative censoring can be applied post hoc via

p(X,T,E)p(X, T, E)0

Training details reported in the appendix include 5 transformer hidden layers, 2 MLP hidden layers, 1 hidden layer for the survival head, 300 sampling steps, training for 4000 epochs, learning rate 0.002, and batch size 256 (Brockschmidt et al., 26 Sep 2025).

SurvDiff is evaluated against NFlow, TVAE, CTGAN, TabDiff, and SurvivalGAN. Covariate distribution fidelity is assessed using Jensen–Shannon distance and Wasserstein distance; downstream survival utility is evaluated under the “train on synthetic, test on real” paradigm with Cox proportional hazards, Weibull accelerated failure time, random survival forest, Survival XGBoost, and DeepHit; and survival-specific fidelity is evaluated using restricted mean survival time gap and KM MSE. Across AIDS, GBSG2, and METABRIC, SurvDiff achieves the lowest JS distance across all datasets, and its Wasserstein distance is competitive and generally better than SurvivalGAN. On AIDS, the Wasserstein distance is reported as approximately p(X,T,E)p(X, T, E)1 for SurvivalGAN versus p(X,T,E)p(X, T, E)2 for SurvDiff, about a 44–50% reduction (Brockschmidt et al., 26 Sep 2025).

The strongest gains appear in downstream survival performance under heavier censoring. On AIDS, which has 91.7% censored observations, SurvDiff reaches a C-index of p(X,T,E)p(X, T, E)3 versus p(X,T,E)p(X, T, E)4 for SurvivalGAN, and a Brier score of p(X,T,E)p(X, T, E)5 versus p(X,T,E)p(X, T, E)6. On GBSG2, with 56.4% censored observations, SurvDiff attains a C-index of p(X,T,E)p(X, T, E)7 and a Brier score of p(X,T,E)p(X, T, E)8, both best among the baselines. On METABRIC, with 42% censored observations, it is on par with TabDiff for C-index and Brier score, while both outperform SurvivalGAN. The reported training dynamics show that diffusion and survival losses decrease smoothly and stably, adaptive weighting keeps the survival loss at an appropriate scale relative to diffusion, and training times are modest, at less than 13 minutes on an A100 per experiment (Brockschmidt et al., 26 Sep 2025).

5. Relation to SDPM and the emergence of “SurvDiff-style” survival prediction

SDPM explicitly contrasts itself with SurvDiff and clarifies a methodological split within diffusion-based survival research. SurvDiff is described there as an end-to-end diffusion model for synthetic survival data, modeling p(X,T,E)p(X, T, E)9 to generate full synthetic datasets that preserve survival and censoring structure for privacy and data sharing. SDPM, by contrast, is a conditional outcome diffusion model for survival prediction, modeling P(T,δ∣x)\mathbb{P}(T,\delta \mid \mathbf{x})0 and using the generated samples to estimate individual survival functions (Kirpichenko et al., 21 May 2026).

SDPM treats the observable pair P(T,δ∣x)\mathbb{P}(T,\delta \mid \mathbf{x})1 as the target random variable in a diffusion model, learns a conditional generative model P(T,δ∣x)\mathbb{P}(T,\delta \mid \mathbf{x})2 via DDPM-style training, and at test time samples many pairs P(T,δ∣x)\mathbb{P}(T,\delta \mid \mathbf{x})3 for a new P(T,δ∣x)\mathbb{P}(T,\delta \mid \mathbf{x})4. It then reconstructs the individual survival function P(T,δ∣x)\mathbb{P}(T,\delta \mid \mathbf{x})5 nonparametrically by applying the Kaplan–Meier estimator to these samples. The model operates in a transformed target space: time is log-transformed and standardized,

P(T,δ∣x)\mathbb{P}(T,\delta \mid \mathbf{x})6

and the censoring indicator is represented continuously as a Gaussian mixture with

P(T,δ∣x)\mathbb{P}(T,\delta \mid \mathbf{x})7

This avoids parametric assumptions on the event-time distribution and does not require a discretization of the output time space (Kirpichenko et al., 21 May 2026).

Empirically, SDPM is evaluated on ten real survival datasets against five strong baselines, including tree-based, boosting-based, and neural survival models. Results show that it achieves competitive predictive performance across C-index, integrated time-dependent AUC, and integrated Brier score; it is best on 7/10 datasets for IBS and has the best average rank across all metrics, especially for IBS. In a synthetic Cox–Weibull study, SDPM can recover the shape of an underlying continuous survival distribution more accurately than random survival forest when sufficiently many samples are generated: at P(T,δ∣x)\mathbb{P}(T,\delta \mid \mathbf{x})8, average KS distances are P(T,δ∣x)\mathbb{P}(T,\delta \mid \mathbf{x})9, z=(x(cont),T)∈Rdcont+1,z = (x^{(\text{cont})}, T) \in \mathbb{R}^{d_{\text{cont}} + 1},0, and z=(x(cont),T)∈Rdcont+1,z = (x^{(\text{cont})}, T) \in \mathbb{R}^{d_{\text{cont}} + 1},1 for event rates 25%, 50%, and 75%, compared with z=(x(cont),T)∈Rdcont+1,z = (x^{(\text{cont})}, T) \in \mathbb{R}^{d_{\text{cont}} + 1},2, z=(x(cont),T)∈Rdcont+1,z = (x^{(\text{cont})}, T) \in \mathbb{R}^{d_{\text{cont}} + 1},3, and z=(x(cont),T)∈Rdcont+1,z = (x^{(\text{cont})}, T) \in \mathbb{R}^{d_{\text{cont}} + 1},4 for RSF (Kirpichenko et al., 21 May 2026).

This suggests that SurvDiff has become both a specific model name and a reference point for a broader design pattern in which denoising diffusion is used to learn survival-relevant distributions and classical survival estimators, especially Kaplan–Meier, are applied to generated samples. The distinction between synthetic data generation and individual prediction remains central.

6. Distinction from adjacent methods and recurrent ambiguities

A recurrent ambiguity in the literature is terminological rather than methodological. Diffsurv is not a diffusion model for survival data generation. Instead, it is a method for differentiable sorting for censored time-to-event data: a neural network maps covariates to scalar risk scores, a differentiable sorting network produces a predicted permutation matrix z=(x(cont),T)∈Rdcont+1,z = (x^{(\text{cont})}, T) \in \mathbb{R}^{d_{\text{cont}} + 1},5, and censoring is handled through a possible permutation matrix z=(x(cont),T)∈Rdcont+1,z = (x^{(\text{cont})}, T) \in \mathbb{R}^{d_{\text{cont}} + 1},6 that encodes feasible ranks under partial order constraints. The loss is set-based rather than pairwise, and in the base case z=(x(cont),T)∈Rdcont+1,z = (x^{(\text{cont})}, T) \in \mathbb{R}^{d_{\text{cont}} + 1},7, Diffsurv is equivalent to the pairwise ranking loss and Cox partial likelihood. On survSVHN, Diffsurv reaches C-index values of .918, .934, .940, .943, and .941 for risk set sizes z=(x(cont),T)∈Rdcont+1,z = (x^{(\text{cont})}, T) \in \mathbb{R}^{d_{\text{cont}} + 1},8, compared with .913, .925, .931, .933, and .930 for Cox partial likelihood; for top-10% prediction, its top-z=(x(cont),T)∈Rdcont+1,z = (x^{(\text{cont})}, T) \in \mathbb{R}^{d_{\text{cont}} + 1},9 variant reports z~=(X(disc),E)\tilde{z} = (X^{(\text{disc})}, E)0 versus Cox’s z~=(X(disc),E)\tilde{z} = (X^{(\text{disc})}, E)1 (Vauvelle et al., 2023).

Another adjacent but distinct use of the string “SurvDiff” appears in cumulative-difference methodology for subgroup comparison. Tygert’s graphical method of cumulative differences between two subpopulations is a bin-free technique for comparing outcome distributions along an ordering variable such as a score. It builds local differences z~=(X(disc),E)\tilde{z} = (X^{(\text{disc})}, E)2, accumulates them into a cumulative curve

z~=(X(disc),E)\tilde{z} = (X^{(\text{disc})}, E)3

and summarizes the result with KS-style and Kuiper-style statistics,

z~=(X(disc),E)\tilde{z} = (X^{(\text{disc})}, E)4

The paper notes that this is conceptually parallel to survival-analysis cumulative sums, including log-rank-style comparisons, but it is neither a diffusion model nor a generator of synthetic survival datasets (Tygert, 2021).

The resulting conceptual map is therefore precise. SurvDiff, in its named form, is a score-based diffusion generator for mixed-type covariates, observed survival times, and event indicators. SDPM inherits the diffusion-based logic but reorients it toward conditional survival prediction. Diffsurv replaces diffusion with differentiable sorting for censored ranking. Tygert’s cumulative-difference framework offers a graphical and statistical comparison tool over ordered scores. These methods address different objects—synthetic data generation, individual prediction, ranking, and graphical comparison—even when they share notation, censoring-aware reasoning, or cumulative survival-analysis intuition.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SurvDiff.