Stability of the Competing-Risk Survival-Head Ordering

Determine whether the observed ordering among Cox, MTLR, and DeepHit survival-head interfaces remains stable when evaluated on larger competing-risk benchmarks.

Background

The paper evaluates competing-risk adaptation using four data sets and finds that cause-specific MTLR ranks highest among the Tabular Foundation Model survival heads, while cause-specific Cox performs less strongly and jointly normalized DeepHit occupies an intermediate position. The authors emphasize that this ordering is preliminary because the competing-risk benchmark is substantially smaller than the single-risk benchmark and may be sensitive to individual data sets.

A larger and more diverse competing-risk evaluation is therefore needed to determine whether the reported relative performance ordering reflects a reproducible pattern or is an artifact of the limited benchmark. This unresolved question concerns the general stability of the comparison among survival-head interfaces, rather than the performance of any single data set.

References

Larger competing-risk benchmarks would also be needed to determine whether the observed ordering among survival heads is stable.

Adaptation Interfaces for In-Context Tabular Foundation Models in Time-to-Event Prediction  (2609.04901 - Pham et al., 4 Sep 2026) in Section Discussion, paragraph beginning “Several limitations should be considered” / final paragraph of Discussion