---
title: 'HySurvPred: Hyperbolic Multimodal Survival Model'
url: https://www.emergentmind.com/topics/hysurvpred
type: topic
---

# HySurvPred: Hyperbolic Multimodal Survival Model

Searching arXiv for the HySurvPred paper and closely related multimodal survival-prediction work.
HySurvPred is a multimodal survival-prediction framework that combines histopathology images and genomic data in hyperbolic space, with the explicit aim of modeling hierarchical structure, preserving the continuous and ordinal character of survival time, and incorporating censored samples into optimization. It was introduced as “Multimodal Hyperbolic Embedding with Angle-Aware Hierarchical Contrastive Learning and Uncertainty Constraints for Survival Prediction” and is described as the first multimodal survival model to embed whole-slide images and pathway-level omics jointly in hyperbolic space [2503.13862].

## 1. Conceptual framing

HySurvPred addresses three limitations identified in prior multimodal survival-prediction methods. First, those methods rely on multimodal mapping and metrics in Euclidean space, which, in the formulation of the paper, cannot fully capture the hierarchical structures in histopathology and genomics data. Second, they discretize survival time into independent risk intervals, thereby ignoring its continuous and ordinal nature. Third, they treat censorship as a binary indicator and exclude censored samples from model optimization, rather than making full use of them [2503.13862].

The framework is motivated by the observation that the two constituent modalities exhibit distinct hierarchies. Whole-slide images carry a tissue-to-cell-cluster-to-cell hierarchy, while transcriptomic profiles follow a network-to-pathway-to-gene hierarchy. Hyperbolic geometry is used because it expands exponentially with radius and naturally embeds trees or hierarchies with low distortion; points near the origin represent high-level abstractions, while points farther out encode finer details. This geometric choice underpins the model’s attempt to represent both modalities within a single latent space while preserving their hierarchical organization [2503.13862].

A plausible implication is that HySurvPred belongs simultaneously to multimodal survival modeling, geometry-aware representation learning, and contrastive survival learning. Its design is not limited to feature fusion; it redefines the geometry of multimodal alignment and the supervision signal attached to survival outcomes.

## 2. Multimodal hyperbolic representation

At the core of HySurvPred is the Multimodal Hyperbolic Mapping module, abbreviated MHM. The model adopts the \(n\)-dimensional Poincaré ball of curvature \(c>0\),

$$
D^n_c \;=\;\bigl\{\,z\in\mathbb{R}^n\;\big|\;c\|z\|^2<1\bigr\}\,,
$$

with distance

$$
d_c(u,v)\;=\;\operatorname{arccosh}\!\Bigl(1+2\,\frac{\|u-v\|^2}{(1-\|u\|^2)(1-\|v\|^2)}\Bigr)\,.
$$

A Euclidean feature \(x\in\mathbb{R}^n\), whether a whole-slide-image patch embedding or a pathway-level genomic vector, is mapped into hyperbolic space through the exponential map at the origin,

$$
\exp^c_{0}(x) \;=\; \tanh\!\Bigl(\tfrac{\sqrt{c}\,\|x\|}{2}\Bigr)\,\frac{x}{\sqrt{c}\,\|x\|}\,.
$$

The curvature parameter \(c\) is learnable, shared across modalities, and optimized by gradient descent [2503.13862].

For histopathology, each whole-slide image is treated as a bag \(\{x^p_i\}_{i=1}^{M_p}\), processed by an attention-based MIL extractor to yield \(M_p\) vectors, each projected into \(D^n_c\). For genomics, gene expression values are first aggregated into \(M_g\) Hallmark pathways, producing \(\{x^g_j\}_{j=1}^{M_g}\), which are likewise mapped by \(\exp^c_0\). The two point sets then coexist in a single Poincaré ball and are fused by channel-wise interactions and multi-head attention within \(D^n_c\) [2503.13862].

This construction encodes an explicit asymmetry between abstraction levels across modalities. The paper’s later angle-aware constraint formalizes the expectation that genomic embeddings are more abstract and should therefore lie closer to the ball center, while histopathology embeddings are more detailed and should lie farther outward. This suggests that HySurvPred treats cross-modal fusion as a hierarchical alignment problem rather than merely a correlation-maximization problem.

## 3. Survival-aware contrastive learning and angle-aware hierarchy

HySurvPred introduces the Angle-aware Ranking-based Contrastive Loss, or ARCL, to preserve the continuous and ordinal nature of survival time. For a query patient embedding \(q\in D^n_c\), all other patients are sorted by observed survival time, yielding ordered positives \(P_1,P_2,\dots,P_r\) from earliest death to latest death, together with negatives \(\mathcal N\). Similarity is measured by the critic

$$
t(q,p)= -d_c(q,p).
$$

At rank \(i\), the model pulls \(q\) closer to samples in \(P_i\) than to any sample in \(\bigcup_{j\ge i}P_j\cup\mathcal N\). The per-rank loss is

$$
\ell_i \;=\;  -\,\log\, \frac{\sum_{p\in P_i}\exp\bigl(t(q,p)/\tau_i\bigr)}
     {\sum_{p\in\cup_{j\ge i}P_j}\exp\bigl(t(q,p)/\tau_i\bigr)
      +\sum_{n\in\mathcal N}\exp\bigl(t(q,n)/\tau_i\bigr)}\,,
$$

and the overall ranking loss is

$$
\mathcal L_{\rm ARCL} \;=\;\sum_{i=1}^r \ell_i\,.
$$

In this formulation, the survival label is not reduced to a single class or interval index; instead, relative ordering among patients becomes the organizing principle of contrastive supervision [2503.13862].

ARCL also incorporates an angle-aware regularizer via an “entailment cone” to enforce cross-modal hierarchy. Genomic embeddings \(x_g\) should lie closer to the ball center than their paired histopathology embeddings \(x_p\). The cone half-aperture is defined as

$$
\mathrm{aper}(x_g) =\sin^{-1}\!\Bigl(\tfrac{2K}{\sqrt{c}\|x_g\|}\Bigr) \quad(K=0.1),
$$

and the exterior angle as

$$
\mathrm{ext}(x_g,x_p) =\cos^{-1}\!\Bigl(\frac{\langle x_g,x_p\rangle_L}
                                    {\|x_g\|\|x_p\|}\Bigr)\!,
$$

where \(\langle\cdot,\cdot\rangle_L\) is the Lorentzian inner product in the hyperboloid model. Violations are penalized by

$$
a(x_g,x_p)\;=\;\max\!\Bigl(0,\,1-\tfrac{\mathrm{ext}(x_g,x_p)-\mathrm{aper}(x_g)}{\pi}\Bigr),
$$

and \(\sum a(x_g,x_p)\) is incorporated into \(\mathcal L_{\rm ARCL}\) as an angle-aware regularizer [2503.13862].

A common misconception would be to interpret this component as a generic contrastive objective. In fact, the design is explicitly ranking-based and hierarchy-constrained. The model does not merely separate positive from negative pairs; it encodes ordinal survival relations and a modality-specific abstraction ordering within hyperbolic geometry.

## 4. Censor-conditioned uncertainty and objective function

HySurvPred introduces a third module, the Censor-Conditioned Uncertainty Constraint, abbreviated CUC, to use right-censored samples more fully. The framework treats censored cases as carrying partial information and encourages their embeddings to remain nearer the origin, which the paper associates with higher epistemic uncertainty, while uncensored cases are pushed outward. The loss is

$$
\mathcal L_{\rm CUC} = \sum_{i\mid c_i=1}\bigl\|\, z_i - r_{\rm high}\bigr\| + \sum_{i\mid c_i=0}\bigl\|\, z_i - r_{\rm low}\bigr\|,
$$

where \(z_i\) is the final hyperbolic embedding of patient \(i\), and

$$
r_{\rm low} =\sqrt{\tfrac{1}{c}\cosh^{-1}(c_{\rm low})}\,,\quad
r_{\rm high} =\sqrt{\tfrac{1}{c}\cosh^{-1}(c_{\rm high})},\quad
c_{\rm low}>c_{\rm high}.
$$

This softly anchors censored embeddings near radius \(r_{\rm high}\), closer to the origin, and uncensored embeddings near \(r_{\rm low}\) [2503.13862].

The overall training objective combines three terms:

$$
\mathcal L \;=\; \underbrace{\mathcal L_{\rm surv}_{\text{negative log-likelihood over discrete risk intervals}
\;+\; \lambda\;\underbrace{\mathcal L_{\rm ARCL}_{\text{ranking + angle regularizer}
\;+\; \gamma\;\underbrace{\mathcal L_{\rm CUC}_{\text{uncertainty anchoring}\,,
$$

with \(\lambda=0.1\) and \(\gamma=0.001\), chosen via ablation [2503.13862].

There is a noteworthy tension in the formulation. The paper criticizes the discretization of survival time into independent risk intervals, yet the final loss still includes a negative log-likelihood over discrete risk intervals. The intended resolution, as stated in the model description, is that ARCL supplements this term with ranking-based supervision that preserves ordinal structure. This suggests that HySurvPred should be understood as a hybrid objective rather than a complete replacement of interval-based survival losses.

## 5. Empirical evaluation on TCGA cohorts

HySurvPred was benchmarked on five TCGA cohorts with paired whole-slide images and mRNA profiles: BLCA (\(N=373\)), BRCA (\(N=956\)), UCEC (\(N=480\)), LUAD (\(N=453\)), and GBMLGG (\(N=569\)). Genomic features were aggregated into six Hallmark pathways. Performance was reported by concordance index on held-out test splits using standard TCGA protocols [2503.13862].

The paper compares HySurvPred with 13 state-of-the-art models, including unimodal MIL, self-normalizing nets, multimodal co-attention, optimal transport, and factorized bilinear fusion. HySurvPred achieved the highest mean C-index on four of the five cohorts and the second best on GBMLGG [2503.13862].

| Cohort | HySurvPred C-index | Previous best |
|---|---:|---:|
| BLCA | 0.711 | 0.691 |
| BRCA | 0.757 | 0.684 |
| UCEC | 0.726 | 0.703 |
| LUAD | 0.709 | 0.696 |
| GBMLGG | 0.859 | 0.861 |

Across all five datasets, paired \(t\)-tests against the prior top performer, MOTCat, yielded \(p<0.05\) [2503.13862]. On the reported numbers, the largest absolute gain appears on BRCA, while GBMLGG is the sole cohort where HySurvPred does not achieve the top score. This suggests that the framework’s inductive biases are broadly effective but not universally dominant across all disease contexts.

The study also reports Kaplan–Meier analyses using median predicted risk to stratify patients into high- and low-risk groups. HySurvPred yielded significantly clearer separation with log-rank \(p\ll0.05\) on all five cohorts, whereas baselines showed mixed strata [2503.13862]. In the context of survival modeling, this indicates that the learned representations support not only pairwise ranking metrics such as the C-index but also clinically interpretable risk stratification.

## 6. Ablations, representation analysis, and implications

An ablation study on BRCA begins from the MOTCat baseline with \(\mathrm{C\!-\!index}\approx0.673\) and incrementally adds model components. MHM only yields 0.740, ARCL only yields 0.741, MHM+ARCL yields 0.739, and the full HySurvPred model with MHM+ARCL+CUC yields 0.757 [2503.13862].

These results support the claim that each module contributes to performance, with the greatest single gain attributed to ARCL and an additional uplift from CUC. At the same time, the MHM+ARCL combination being slightly below each individual module on BRCA indicates that component interactions are not trivially additive. A plausible implication is that the uncertainty constraint acts as a stabilizing mechanism when the hierarchical hyperbolic representation and ranking-based contrastive objective are combined.

The paper further reports t-SNE visualizations of the final hyperbolic embeddings. Before training, pathology and genomic points intermix around the origin. After training, the two modality clouds pull apart radially, with genomics inward and histology outward, and risk groups corresponding to early versus late death arrange along monotonic gradations in radius [2503.13862]. This visualization is consistent with the intended semantics of the hyperbolic space: radial position is used to encode abstraction level and, after optimization, also correlates with survival-related structure.

The model’s conclusions outline several future directions: jointly learning the curvature \(c\) per cohort or per modality, extending to deeper hierarchies such as full gene networks and multi-resolution tissue graphs, and incorporating Bayesian hyperbolic Gaussian processes for calibrated uncertainty [2503.13862]. These directions suggest that HySurvPred is best viewed as a framework rather than a fixed architecture. Its central contribution lies in linking hyperbolic multimodal representation, ordinal survival supervision, and censor-aware uncertainty constraints within one optimization scheme.

Source: https://www.emergentmind.com/topics/hysurvpred