Generalization beyond English humor and sarcasm benchmarks

Determine whether the modality hierarchies observed in the MultiHuSE, UR-FUNNY, and MUStARD benchmarks and the performance benefits of Inverted Asymmetric Fusion generalize to other languages and to tasks such as emotion recognition, visual question answering, and action recognition.

Background

The evaluation is restricted to English-language datasets focused on humor and sarcasm detection. Although the use of the multilingual E5 encoder suggests possible cross-lingual applicability, the paper does not evaluate Inverted Asymmetric Fusion on non-English data or on other multimodal tasks. Consequently, whether the reported modality hierarchies and architectural benefits transfer to different languages and task domains remains unresolved.

References

It therefore remains unclear whether the modality hierarchies observed here, and the benefits of IAF, generalise to other languages or to tasks such as emotion recognition, visual question answering, or action recognition.

Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion  (2608.26879 - Kenneth et al., 27 Aug 2026) in Section “Limitations”