Convergence analysis under asymmetric adapter rejection

Derive a convergence bound for decentralized large language model fine-tuning with Chorus when receiver-specific rejection decisions produce a row-stochastic but non-doubly-stochastic mixing matrix.

Background

Chorus allows each node to reject adapters from particular neighbors before aggregation. Because these rejection decisions need not be symmetric, the resulting mixing matrix remains row-stochastic but is not necessarily doubly stochastic. Standard convergence analyses for decentralized stochastic gradient descent typically assume double stochasticity, so the paper evaluates utility empirically rather than proving convergence. Establishing a convergence guarantee for this asymmetric aggregation setting would provide a theoretical foundation for Chorus and clarify the conditions under which its decentralized fine-tuning process converges.

References

The resulting mixing matrix remains row-stochastic but is no longer doubly stochastic, which is the condition that standard convergence analyses of decentralized SGD assume~\citep{lian2017dpsgd,koloskova2020unified}. We therefore do not derive a convergence bound, and instead track utility empirically through the held-out cross-entropy per round (\Cref{sec:eval}).

— Backdoor Mitigation in Decentralized LLM Fine-Tuning  (2609.37367 - Biswas et al., 29 Sep 2026) in Section 3, subsection “Stage 2: Update trust and aggregate adapters,” paragraph “Cost of rejecting adapters”