Scaling RegionFed to Larger Transformer Models
Empirically validate whether RegionFed maintains its personalization effectiveness and stable adaptation intensity when applied to larger transformer models, specifically T5-Base with 220 million parameters and T5-Large with 770 million parameters.
References
We hypothesize that RegionFed will maintain effectiveness on T5-Base (220M) and T5-Large (770M) because: (1) gradient conflicts remain informative regardless of model size; (2) the optimal $\rho$ discovered via golden section search will self-adjust to appropriate magnitudes; (3) architecture-agnostic operations avoid the parameter-count-dependent failures observed in SCAFFOLD/pFedMe. Empirical validation on larger models remains future work.
— RegionFed: Federated Learning for Personalized Query Understanding in Heterogeneous Retail Environments
(2609.05403 - Nguyen et al., 4 Sep 2026) in Appendix, Section 'Supplementary Algorithms and Extensions', subsection 'Future Directions and Extensions', subsubsection 'Scaling to Larger Transformers'