Papers
Topics
Authors
Recent
Search
2000 character limit reached

FedMIX-P: Mixing Local and Global Preconditioners for Federated Vision and Language Model Training

Published 1 Oct 2026 in cs.LG and cs.AI | (2610.01515v1)

Abstract: Adaptive preconditioners accelerate model training, but heterogeneous client geometries can bias federated updates even when gradients are evaluated at the same model. Round-start synchronization alone cannot prevent this mismatch from reappearing during local training. We propose \texttt{FedMIX-P}, which mixes shared and local preconditioners at every local step, retaining local adaptation while reducing mean-squared operator mismatch by a factor of λ<sup>2λ<sup>2. For smooth nonconvex objectives with stochastic gradients and partial participation, we establish an O(R<sup>−1/2)O(R<sup>{-1/2}) stationarity bound using suitable stepsizes and a horizon-dependent mixing weight, without requiring local preconditioners to converge to one another. A two-client counterexample shows that fixed positive mixing can preserve a nonstationary fixed point. The theory covers bounded linear symmetric positive-definite preconditioners. Experiments with SOAP, Sophia, and Muon variants across vision and language tasks show improvements over corresponding local optimizers, including accuracy gains of up to $19.47$ percentage points and lower validation loss for 60M--350M LLMs. Full nonlinear and momentum-based updates require separate analysis.

Authors (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.