Mutual-information estimation without the sufficient encoder assumption

Estimate or approximate the mutual information I(X;\tilde{X}) for SMILE's selected multimodal medical representations without relying on the sufficient encoder assumption, thereby addressing the information-theoretic difficulty of performing this estimation directly.

Background

SMILE optimizes modality-specific feature explanations using an information-bottleneck objective that includes the compression term I(\tilde{X}m;Xm). To make this term tractable for heterogeneous medical data, the method assumes sufficiently expressive, lossless modality-specific encoders and estimates mutual information in the resulting latent spaces.

The authors explicitly identify this sufficient encoder assumption as over-optimistic. Removing it would require a principled method for estimating or approximating the mutual information between the original medical input X and the selected representation \tilde{X}, including in settings involving high-dimensional, noisy, and structured multimodal data.

References

First, the sufficient encoder assumption in (\ref{eq: multi_loss2}) simplifies optimization but remains over-optimistic; estimating or approximating $I(X;\tilde{X})$ without it is an open problem, even from an information theory perspective.

SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis  (2609.05174 - Yang et al., 4 Sep 2026) in Section Conclusion and Future Work