Formal characterization of training-time weighting alignment regimes

Characterize formally the regimes in which training-time modality weighting is aligned with test-time modality utility, including low complementarity, dominance alignment, and low sample-level variation.

Background

The paper identifies three conditions under which optimization-based weighting may be harmless or may accidentally correlate with modality utility: one modality strictly dominates with little complementarity, the fastest-learning modality is also the most informative at test time, or optimal fusion weights are approximately constant across samples.

These conditions are presented as alignment regimes rather than guarantees. The authors state that a formal characterization is still unresolved, particularly because substantial complementarity, overfitting by fast learners, and sample-dependent modality utility can make training-time fitting signals diverge from discriminative contribution.

References

These are alignment regimes rather than guarantees. When complementarity is substantial, when fast learners overfit, or when modality utility varies across samples, the mismatch between fitting and contributing becomes decisive. Characterizing these regimes more formally remains an open problem.

The Illusion of Balanced Multimodal Sentiment Analysis: Beyond the Limits of Optimization-Based Methods  (2609.11247 - Kaffeza et al., 10 Sep 2026) in Section 5.2, “When Should We Expect Training-Time Weighting to Work?”