Disentangle the contributions of temporal modeling, label refinement, and multimodal learning

Determine the individual contributions of temporal modeling, label refinement, and multi-task learning to the macro-phase recognition improvement by conducting dedicated ablation studies.

Background

The reported macro-phase improvement compares the dual-head MS-TCN++ model with refined labels against a zero-shot vision-language baseline. This configuration changes several components simultaneously: it introduces temporal modeling, refined phase labels, and multi-task learning. Because these factors are not evaluated separately, the relative contribution of each component to the observed performance gain remains unresolved. The paper identifies dedicated ablation experiments as necessary to disentangle their effects.

References

This configuration jointly introduces temporal modeling, label refinement, and multi-task learning relative to the Stage~A baseline; disentangling their individual contributions requires dedicated ablations, which we leave to future work.

Multimodal Shared Latent Representation of Narration, Microscope and iOCT Images for Phase Recognition in Vitreoretinal Surgery  (2608.31065 - Izmitlioglu et al., 31 Aug 2026) in Section Experimental Results, paragraph beginning “For macro-phase recognition”