Representation-compression explanation for vision–language performance differences
Investigate whether the stronger degradation of auxiliary-network-based Split Federated Learning on vision tasks than on language tasks is caused by differences in intermediate-representation compression, specifically because vision representations progressively discard label-irrelevant information while masked language models preserve more input information.
References
We conjecture that this difference stems from how strongly the interme- diate representations are compressed.
— CoeF-SFL: Preserving Collaborative Server-Client Learning with Enhanced Communication Efficiency
(2609.34360 - Bae et al., 28 Sep 2026) in Section 5, Comparison results (page 8)