---
title: Multivariate Shared Frailty Cure-Rate Model
url: https://www.emergentmind.com/topics/multivariate-shared-frailty-cure-rate-model
type: topic
---

# Multivariate Shared Frailty Cure-Rate Model

Searching arXiv for recent and foundational papers on multivariate shared frailty cure-rate models and related frailty diagnostics.
A multivariate shared frailty cure-rate model is a clustered survival model in which multiple time-to-event outcomes within a cluster share an unobserved multiplicative random effect and a non-susceptible subpopulation is explicitly represented through a cure mechanism. In this framework, dependence among margins arises through the shared frailty, while long-term survivors are accommodated either by a mixture cure specification or, in discrete frailty formulations, by an atom at zero in the frailty distribution. The model has been developed for clustered current status data with semiparametric generalized odds-rate structure [1912.00295], for familial breast cancer risk prediction using families as the unit of analysis [2508.16350], and for broader analysis of shared frailty cure-rate behavior through the relative frailty variance and cross-ratio function in discrete frailty systems [2303.04915].

## 1. Conceptual definition and model scope

The defining feature of the model is the simultaneous treatment of two forms of latent structure: a shared frailty that induces within-cluster dependence, and a cure component that separates susceptible from non-susceptible subjects. In clustered settings, units within the same cluster are conditionally independent given the frailty, but are marginally associated because the frailty is common to all margins in that cluster [2303.04915].

In the clustered current status formulation, clusters are indexed by \(i = 1,\dots,m\), units by \(j = 1,\dots,n_i\), and each unit contributes a single inspection time \(C_{ij}\) and current-status indicator
\[
\Delta_{ij} = I(T_{ij} \le C_{ij}),
\]
with \(T_{ij}\) denoting latent event time [1912.00295]. In the familial breast cancer formulation, families \(j = 1,\dots,J\) replace generic clusters, and individuals \(i\) within family \(j\) contribute observed time \(t_{ij}\) and event indicator \(\delta_{ij}\) under right-censoring [2508.16350]. These two data structures represent different censoring regimes, but the common modeling principle is identical: cluster-level latent heterogeneity is encoded by a shared positive frailty, and the cure fraction accounts for subjects who never experience the event.

The cure component may be expressed through a latent susceptibility indicator, with \(Z_{ij}=1\) denoting susceptible and \(Z_{ij}=0\) denoting cured, and cure probability
\[
\pi(x_{ij}) = P(Z_{ij}=0 \mid x_{ij}),
\]
often linked through
\[
\pi(x_{ij}) = \operatorname{expit}(x_{ij}^{\top}\gamma)
\]
in mixture formulations [1912.00295]. In discrete frailty models, cure may instead be represented structurally by \(P(Z=0)>0\), so that subjects with zero frailty have null hazard and contribute a survival plateau [2303.04915]. This yields two related but distinct meanings of “cure-rate model”: one based on a latent mixture over susceptibility, the other based on an atom at zero in the frailty law.

## 2. Core probabilistic structure

The shared frailty mechanism introduces a cluster-specific random effect that scales the hazard of all susceptible margins in the cluster. In the semiparametric current status formulation, the within-subject correlation is accounted for by a random frailty effect \(U_i>0\), commonly taken to follow a gamma distribution with mean \(1\) and variance \(\theta>0\),
\[
U_i \sim \operatorname{Gamma}\!\left(\frac{1}{\theta}, \theta\right),
\]
with moments
\[
\mathbb{E}[U_i]=1,\qquad \operatorname{Var}(U_i)=\theta
\]
[1912.00295]. In the breast cancer family-history model, the analogous family-level frailty \(U_j\) is distributed as
\[
U_j \sim \text{Gamma}(k,k),\qquad E[U_j]=1,\qquad \text{Var}(U_j)=1/k
\]
[2508.16350]. The parameterization differs, but both formulations use mean-one gamma frailty to encode latent cluster risk.

Conditional on susceptibility and frailty, the model specifies the event-time law among susceptibles. In the proportional hazards family-history formulation,
\[
h_{ij}(t \mid U_j, X_{ij}) = U_j\, h_0(t)\, \exp(X_{ij}^\top \beta),
\]
with cumulative baseline hazard
\[
\Lambda_0(t)=\int_0^t h_0(s)\,ds
\]
and susceptible survival
\[
S_{ij}^s(t \mid U_j, X_{ij}) = \exp\{-U_j \Lambda_0(t)\exp(X_{ij}^\top\beta)\}
\]
[2508.16350]. In the generalized odds-rate formulation for current status data, the susceptible survival is
\[
S_0(t \mid x_{ij}, U_i)=\left\{1+\alpha U_i\Lambda_0(t)\exp(\eta_{ij})\right\}^{-1/\alpha},
\]
where \(\eta_{ij}=x_{ij}^\top\beta\) and \(\alpha \in \mathbb{R}_+\) is the GOR shape parameter [1912.00295]. This family nests proportional odds at \(\alpha=1\) and converges to proportional hazards as \(\alpha \to 0^+\) [1912.00295].

The overall survival function then combines cure and susceptibility. Under the logistic mixture representation,
\[
S_{ij}(t \mid U_j, X_{ij}, Z_{ij})=\pi_{ij}+(1-\pi_{ij})S_{ij}^s(t \mid U_j, X_{ij})
\]
[2508.16350]. In the current-status setting, the corresponding event-status probabilities at inspection are
\[
P(\Delta_{ij}=1 \mid C_{ij},x_{ij},U_i)=(1-\pi(x_{ij}))F_0(C_{ij}\mid x_{ij},U_i),
\]
\[
P(\Delta_{ij}=0 \mid C_{ij},x_{ij},U_i)=\pi(x_{ij})+(1-\pi(x_{ij}))S_0(C_{ij}\mid x_{ij},U_i)
\]
[1912.00295]. These expressions formalize the key ambiguity for a non-event observation: it may indicate cure or merely survival among susceptibles.

## 3. Major formulations

Several formulations fall under the label “multivariate shared frailty cure-rate model,” and the distinction among them is substantive rather than cosmetic.

The semiparametric generalized odds-rate model for clustered current status data combines a logistic cure fraction, a shared gamma frailty, and a nonparametric baseline cumulative hazard approximated by penalized splines [1912.00295]. In this setting, the baseline is written as
\[
\Lambda_0(t)=\sum_{k=1}^{K} b_k I_k(t), \qquad b_k \ge 0,
\]
using I-spline basis functions \(\{I_k(t)\}_{k=1}^K\), with smoothness controlled by a quadratic penalty
\[
\mathcal{P}(b)=\lambda b^\top P b
\]
[1912.00295]. This construction preserves monotonicity of the cumulative hazard and supports semiparametric efficiency arguments.

The familial breast cancer model adopts a shared frailty cure-rate perspective centered on families rather than individuals [2508.16350]. Its stated objective is to model “families rather than individuals as unit of analysis” and to capture “the latent familial risk underlying breast cancer diagnoses” through a shared frailty among family members while “explicitly account[ing] for a fraction of women not susceptible to breast cancer” [2508.16350]. In one version, the model uses the conventional logistic mixture cure specification. In the paper’s Lehmann-type cure-rate specification, however, the survival is written as
\[
S_{ij}(t \mid U_j)=\big[p+(1-p)\widetilde S(t)\big]^{U_j},
\]
where \(p\) is the cure fraction and \(\widetilde S(t)\) is the susceptible survival under a chosen parametric distribution [2508.16350]. Under this structure, the effective family-specific cure fraction becomes \(p^{U_j}\), so the frailty modulates both the hazard among susceptibles and the cure rate [2508.16350].

A third formulation arises in discrete shared frailty models, where cure is induced through a frailty distribution \(\pi_Z\) with an atom at zero [2303.04915]. With \(m\) margins and shared frailty \(Z \ge 0\), the hazard specification is
\[
\lambda^{(j)}(t \mid Z=z)= z \lambda_0^{(j)}(t),\qquad j=1,\dots,m,
\]
and the marginal joint survival is
\[
S(t_1,\dots,t_m)=\mathcal{L}_Z\big(\Lambda_0(t_1,\dots,t_m)\big),
\qquad
\Lambda_0(t_1,\dots,t_m)=\sum_{j=1}^m \Lambda_0^{(j)}(t_j),
\]
where \(\mathcal{L}_Z\) is the Laplace transform of \(Z\) [2303.04915]. If \(P(Z=0)=p_0\), then
\[
S(\infty)=p_0
\]
provided at least one cumulative baseline hazard diverges [2303.04915]. This suggests a structural equivalence between cure and a zero-frailty subpopulation in discrete time-invariant shared frailty systems.

## 4. Likelihood construction and estimation

Likelihood construction depends on the censoring scheme and cure representation. For clustered current status data, the unit-level contribution conditional on frailty \(U_i=u\) is
\[
L_{ij}(u)
=
\big[(1-\pi(x_{ij}))F_0(C_{ij}\mid x_{ij},u)\big]^{\Delta_{ij}}
\big[\pi(x_{ij})+(1-\pi(x_{ij}))S_0(C_{ij}\mid x_{ij},u)\big]^{1-\Delta_{ij}},
\]
and the cluster-level likelihood is
\[
L_i=\int_0^\infty \left\{\prod_{j=1}^{n_i} L_{ij}(u)\right\} p(u)\,du
\]
[1912.00295]. The full log-likelihood is then
\[
\ell(\beta,\gamma,\alpha,\theta,\Lambda_0)=\sum_{i=1}^m \log L_i
\]
[1912.00295].

For right-censored family data under the mixture cure PH formulation, the cluster-wise likelihood conditional on frailty is
\[
L_j(U_j)=\prod_{i=1}^{n_j}
\left\{
\big[(1-\pi_{ij})f_{ij}(t_{ij}\mid U_j)\big]^{\delta_{ij}}
\big[\pi_{ij}+(1-\pi_{ij})S_{ij}^s(t_{ij}\mid U_j)\big]^{1-\delta_{ij}}
\right\},
\]
with marginal likelihood
\[
L_j=\int_0^\infty L_j(U_j)g(U_j)\,dU_j
\]
[2508.16350]. The paper notes that with a logistic mixture cure link, this integral typically requires numerical quadrature or Laplace approximation because the terms \(\pi_{ij}+(1-\pi_{ij})S^s(t\mid U_j)\) do not factorize conveniently [2508.16350].

The estimation strategy for the semiparametric current-status model is based on the EM algorithm [1912.00295]. The latent variables are the susceptibility indicators \(Z_{ij}\) and the cluster frailties \(U_i\). In the E-step, one computes posterior expectations of the susceptibility indicators. If \(\Delta_{ij}=1\), then \(\tau_{ij}=\mathbb{E}[Z_{ij}\mid \text{data}]=1\). If \(\Delta_{ij}=0\), then conditional on \(U_i=u\),
\[
\mathbb{E}[Z_{ij}\mid \Delta_{ij}=0,C_{ij},x_{ij},u]
=
\frac{(1-\pi_{ij})S_0(C_{ij}\mid x_{ij},u)}{\pi_{ij}+(1-\pi_{ij})S_0(C_{ij}\mid x_{ij},u)},
\]
followed by integration over the frailty posterior [1912.00295]. In the M-step, \(\gamma\) is updated through a weighted logistic regression, while \(\beta\), \(b\), \(\alpha\), and \(\theta\) are updated by maximizing the corresponding expected complete-data criterion, with frailty integration handled numerically, for example by adaptive Gauss-Laguerre quadrature [1912.00295].

The family-history paper also describes EM estimation with latent susceptibility indicators \(Y_{ij}\) in the mixture formulation [2508.16350]. For censored individuals,
\[
E[Y_{ij}\mid t_{ij},\delta_{ij}=0,\text{data}]
=
\frac{(1-\pi_{ij})\,\mathbb{E}_{U_j}[S_{ij}^s(t_{ij}\mid U_j)]}
{\pi_{ij}+(1-\pi_{ij})\,\mathbb{E}_{U_j}[S_{ij}^s(t_{ij}\mid U_j)]},
\]
where under gamma frailty
\[
\mathbb{E}[S_{ij}^s(t\mid U)] = \left(1+\frac{a_{ij}}{k}\right)^{-k},
\qquad a_{ij}=\Lambda_0(t_{ij})\exp(X_{ij}^\top \beta)
\]
[2508.16350]. In the Lehmann MSF-CRM, maximum likelihood is simplified because the gamma moment generating function yields a closed-form multivariate likelihood and a gamma posterior for \(U_i\) [2508.16350]. This removes the need for numerical integration over frailty in that specific specification.

## 5. Dependence, identifiability, and diagnostics

Dependence in a shared frailty cure-rate model is not merely a nuisance feature; it is one of the principal inferential targets. In the current-status and family-history formulations, the frailty captures unobserved cluster-level heterogeneity such as patient-level or family-level risk factors [1912.00295; 2508.16350]. Positive frailty values above one identify high-risk clusters, while values below one identify low-risk clusters [2508.16350].

In discrete shared frailty systems, dependence is characterized by the relative frailty variance (RFV) among survivors and its one-to-one relationship with the cross-ratio function (CRF) [2303.04915]. For generic time \(\Lambda=\sum_j \Lambda_0^{(j)}(t_j)\), the RFV is
\[
\operatorname{RFV}(\Lambda)=
\frac{\mathcal{L}_Z''(\Lambda)\mathcal{L}_Z(\Lambda)}{[\mathcal{L}_Z'(\Lambda)]^2}-1,
\]
and the multivariate CRF satisfies
\[
\operatorname{CRF}(\Lambda)=
\frac{\mathcal{L}_Z''(\Lambda)\mathcal{L}_Z(\Lambda)}{[\mathcal{L}_Z'(\Lambda)]^2}
=1+\operatorname{RFV}(\Lambda)
\]
[2303.04915]. Because the CRF depends only on \(\Lambda\), pair-specific cross-ratios are identical at the same generic time. This provides a direct bridge between survivor heterogeneity and within-cluster association.

Identifiability of the cure fraction requires long enough follow-up or inspection-time support to distinguish cure from delayed failure. In the current-status model, identification relies on susceptible survival decaying toward zero while a nonzero fraction remains unfailed even at large inspection times [1912.00295]. In the family-history setting, cure fraction identification requires “sufficiently long follow-up to observe a survival plateau”; otherwise, cure and frailty variance can trade off [2508.16350]. Frailty identification further depends on nondegenerate cluster sizes and informative variation in event times or current-status observations [1912.00295; 2508.16350].

Diagnostics reflect these two model components. For clustered current-status analysis, the proposed methodology is accompanied by “diagnostic checks to identify influential observations” [1912.00295]. The detailed exposition includes case-deletion influence diagnostics, posterior frailty summaries, and inspection of posterior susceptibility probabilities \(\tau_{ij}\) for \(\Delta_{ij}=0\) [1912.00295]. In the family-history context, the recommended diagnostics include residuals for PH among susceptibles, posterior predictive checks for family-level event counts and survival curves, calibration plots, and comparisons across frailty distributions and baseline hazards [2508.16350]. In discrete frailty models, RFV and CRF trajectories are themselves diagnostic tools: an increasing RFV tail supports a cure-rate mechanism through an atom at zero, whereas a decreasing RFV tail supports the absence of cure [2303.04915].

## 6. Applications and comparative evidence

One motivating application is periodontal disease, where patients are clusters and sites or teeth are units [1912.00295]. In a current-status design, each site is examined once at time \(C_{ij}\), and \(\Delta_{ij}=1\) records whether onset has occurred by that time [1912.00295]. Multiple sites per patient naturally induce within-patient dependence, making shared frailty a natural cluster-level construct. The proposed methodology was illustrated on oral health data, with estimation of cure fractions, frailty variance, regression effects, and a smooth baseline cumulative hazard [1912.00295].

A second major application is breast cancer family history, where the cluster is the family rather than the individual [2508.16350]. The stated goal is to identify “high-risk families” for screening and prevention by using complete family history and latent familial frailty [2508.16350]. The model outputs include posterior family risk \(E[U_j\mid \text{data}]\), individualized survival, and predicted cure probabilities or implied family-specific cure levels in the Lehmann specification [2508.16350].

The comparative evidence reported for the breast cancer application is explicit. In simulation studies, the paper evaluates MSF-CRM under susceptible survival distributions including Weibull, Gamma, Lognormal, and 3-parameter Gamma, and compares it against MSF-Cox and univariate models [2508.16350]. The reported results include the following.

| Setting | Reported result | Source |
|---|---|---|
| Parameter recovery | \(p\) is recovered around \(0.84\)–\(0.86\); \(\theta\) is accurately estimated | [2508.16350] |
| Continuous-risk prediction | MSF-CRM achieves \(C \approx 0.93\)–\(0.96\) across scenarios | [2508.16350] |
| Real data, Swedish registry | MSF-CRM achieves \(C \approx 0.97\) across susceptible distributions | [2508.16350] |

The same paper reports that univariate models based on frailty cure-rate or family-history indicators perform poorly on the Swedish Multi-Generational Breast Cancer registry, with \(C \approx 0.39\)–\(0.51\), whereas MSF-CRM outperforms or matches MSF-Cox at approximately \(0.96\)–\(0.97\) [2508.16350]. The interpretation offered is that complete family history matters materially for identifying the latent high-risk group, and that joint modeling of cure and shared frailty yields broader representation of the disease process than models based only on individual histories or binary family-risk indicators [2508.16350].

## 7. Extensions, limitations, and related methodological directions

The model admits several extensions already described in the literature summarized here. In the current-status framework, interval-censored data can be handled by replacing current-status probabilities with interval probabilities, while the EM and penalized spline machinery remains applicable with minor changes [1912.00295]. Competing risks, multivariate margins with multiple causes, alternative frailty distributions such as log-normal or inverse Gaussian, time-varying covariates, and different links for the cure fraction are all described as natural extensions of the same framework [1912.00295].

In familial applications, possible extensions include competing risks for other cancers or non-cancer death, multi-state processes such as onset \(\to\) progression \(\to\) death, recurrent event settings with multiple primaries, time-varying frailty, and hierarchical frailties for maternal or paternal lines [2508.16350]. The same source also notes the possibility of separate family-level random effects in the cure model and hazard model, rather than a single frailty affecting both [2508.16350].

At the same time, several limitations are explicitly recognized. Conditional independence among susceptibles given frailty and covariates may be violated by shared time-varying exposures, which can inflate frailty variance [2508.16350]. Frailty distribution misspecification is a recurrent concern, especially when gamma frailty is adopted primarily for tractability [2508.16350]. In semiparametric cure-rate models, increasing baseline flexibility may aggravate non-identifiability when combined with frailty and cure [2508.16350]. In clustered current-status data, identifiability requires adequate support of inspection times and nondegenerate cluster structure [1912.00295].

A distinct methodological perspective comes from discrete frailty distributions. The analysis of RFV trajectories shows that discrete time-invariant shared frailty models with an atom at zero have RFV tails diverging to infinity, while those without an atom at zero have RFV tails collapsing to zero [2303.04915]. This result provides a model-selection criterion for cure-rate presence and suggests that RFV or CRF diagnostics can distinguish structural cure from no-cure regimes in multivariate shared frailty models [2303.04915]. A plausible implication is that dependence diagnostics can inform not only frailty specification but also the substantive interpretation of cure mechanisms.

Overall, the multivariate shared frailty cure-rate model occupies the intersection of clustered survival analysis, latent heterogeneity modeling, and long-term survival modeling. Across its semiparametric GOR, proportional hazards mixture, Lehmann-type, and discrete frailty variants, it is unified by a single objective: to model event-time dependence and a non-susceptible fraction jointly rather than treating either as an afterthought [1912.00295; 2508.16350; 2303.04915].

Source: https://www.emergentmind.com/topics/multivariate-shared-frailty-cure-rate-model