Representation-compression explanation for vision–language performance differences

Investigate whether the stronger degradation of auxiliary-network-based Split Federated Learning on vision tasks than on language tasks is caused by differences in intermediate-representation compression, specifically because vision representations progressively discard label-irrelevant information while masked language models preserve more input information.

Background

The paper compares auxiliary-network-based Split Federated Learning with vanilla SFL, gradient-reuse baselines, and the proposed CoeF-SFL methods across vision and language tasks. The experiments show that auxiliary-network-based methods lose substantially more accuracy on vision benchmarks, particularly with ViT-Tiny, than on language benchmarks using DistilRoBERTa or RoBERTa.

The authors offer a conjectural explanation rather than a demonstrated conclusion: vision representations may progressively discard label-irrelevant information, and local objectives imposed at shallow cut layers may accelerate that compression before the server-side model can exploit the discarded information. By contrast, masked LLMs may preserve more information about input tokens, making them less vulnerable to this effect. Determining whether this representation-compression mechanism actually explains the observed modality-dependent performance gap remains unresolved.

References

We conjecture that this difference stems from how strongly the interme- diate representations are compressed.

— CoeF-SFL: Preserving Collaborative Server-Client Learning with Enhanced Communication Efficiency  (2609.34360 - Bae et al., 28 Sep 2026) in Section 5, Comparison results (page 8)