Papers
Topics
Authors
Recent
Search
2000 character limit reached

FedSUM: Efficient Federated Learning Algorithms

Updated 27 December 2025
  • FedSUM family of algorithms is a suite of federated learning methods that integrate delay metrics and a stochastic uplink-merge to manage irregular client participation and data heterogeneity.
  • They employ rigorous delay metrics to capture per-round, maximum, and average staleness, ensuring robust convergence guarantees under standard smoothness and variance assumptions.
  • The three variants—FedSUM-B, FedSUM, and FedSUM-CR—offer practical trade-offs between local computation, communication overhead, and memory usage for real-world deployments.

The FedSUM family comprises a suite of federated learning (FL) algorithms that address the challenge of arbitrary client participation patterns in practical distributed optimization. Unlike prior FL methods restricted to idealized participation assumptions or requiring additional constraints on client data heterogeneity, the FedSUM family unifies handling of both temporally irregular client activity and non-i.i.d. data with minimal assumptions. Central to these methods are delay metrics that quantify staleness due to intermittent client participation, and a “Stochastic Uplink-Merge” technique that robustly merges stale and fresh gradient contributions to correct local update bias and reduce communication overhead. The FedSUM family includes three variants—FedSUM-B (basic, without local steps), FedSUM (with local updates), and FedSUM-CR (communication-reduced)—that collectively generalize, and in specific regimes, recover, the behavior of popular FL algorithms such as FedAvg and SCAFFOLD. Convergence guarantees for all FedSUM variants hold under standard smoothness and bounded-variance assumptions, regardless of participation heterogeneity (You et al., 20 Dec 2025).

1. Delay Metrics and Participation Modeling

The FedSUM family introduces a rigorous framework for capturing variability in client participation using two delay metrics:

  • Per-round delay τt\tau_t: For round tt, τt=max⁡i(t−ai,t)\tau_t = \max_{i} (t - a_{i,t}), where ai,ta_{i,t} is the last round in which client ii was active.
  • Maximum delay τmax⁡\tau_{\max}: Over TT rounds, τmax⁡=max⁡0≤t<Tτt\tau_{\max} = \max_{0 \leq t < T} \tau_t bounds the greatest inactivity gap among all clients.
  • Average delay τavg\tau_{\text{avg}}: Given by τavg=(1/T)∑t=0T−1τt\tau_{\text{avg}} = (1/T)\sum_{t=0}^{T-1} \tau_t, it quantifies the typical staleness in the system.

Arbitrary participation, deterministic or random, is modeled entirely through these delay measures, allowing the framework to subsume nonuniform, cyclic, or adversarial dropouts as long as delays remain sub-linear (You et al., 20 Dec 2025).

2. Algorithmic Structure and Variants

Each FedSUM variant maintains a global model tt0, a server control vector tt1, and per-client control vectors tt2. Their principal distinction lies in their local computation strategies and communication patterns.

Variant Local Steps Uplink Downlink
FedSUM-B None (single-step) tt3 tt4
FedSUM tt5 SGD steps tt6 tt7
FedSUM-CR tt8 SGD steps tt9 τt=max⁡i(t−ai,t)\tau_t = \max_{i} (t - a_{i,t})0
  • FedSUM-B: Active clients compute a mini-batch gradient at τt=max⁡i(t−ai,t)\tau_t = \max_{i} (t - a_{i,t})1, transmit the increment τt=max⁡i(t−ai,t)\tau_t = \max_{i} (t - a_{i,t})2, and update their controls; the server updates the global model solely from aggregated deltas. No downlink of server control is needed.
  • FedSUM: Clients perform τt=max⁡i(t−ai,t)\tau_t = \max_{i} (t - a_{i,t})3 local SGD steps at each round, using server-provided τt=max⁡i(t−ai,t)\tau_t = \max_{i} (t - a_{i,t})4 for variance correction. Clients return a single delta encoding both model and control adjustments.
  • FedSUM-CR: Eliminates the extra downlink by letting clients reconstruct the correction direction locally from stored τt=max⁡i(t−ai,t)\tau_t = \max_{i} (t - a_{i,t})5. Local control overhead increases, but downlink cost matches FedSUM-B.

The shared “Stochastic Uplink-Merge” protocol underlies all methods, robustly combining fresh and stale client contributions regardless of participation irregularity (You et al., 20 Dec 2025).

3. Convergence Guarantees and Theoretical Properties

Under the sole requirements that each τt=max⁡i(t−ai,t)\tau_t = \max_{i} (t - a_{i,t})6 is τt=max⁡i(t−ai,t)\tau_t = \max_{i} (t - a_{i,t})7-smooth and each stochastic gradient has variance at most τt=max⁡i(t−ai,t)\tau_t = \max_{i} (t - a_{i,t})8, the FedSUM algorithms converge to a stationary point of τt=max⁡i(t−ai,t)\tau_t = \max_{i} (t - a_{i,t})9 at a quantifiably delay-dependent rate. Explicitly, for properly chosen stepsizes

ai,ta_{i,t}0

the following holds for any (possibly adversarial) client activity sequence: ai,ta_{i,t}1 where ai,ta_{i,t}2 and ai,ta_{i,t}3. For random participation, these bounds hold in expectation over the delay metrics.

This result unifies prior FL convergence analyses, recovers established rates for uniform or cyclical participation, and quantifies the precise effect of staleness via ai,ta_{i,t}4 and ai,ta_{i,t}5 (You et al., 20 Dec 2025).

4. Assumption Minimality and Heterogeneity

Contrasting with FedAvg and SCAFFOLD, which require constraints such as bounded gradient dissimilarity or two-vector communication, FedSUM imposes no restriction on ai,ta_{i,t}6 for client distributions. The only requirements are ai,ta_{i,t}7-smoothness and bounded variance, alongside explicit tracking of delay metrics. This applicability encompasses:

  • Uniform and independent sampling (with ai,ta_{i,t}8),
  • Cyclic or reshuffled scheduling,
  • Adversarially determined dropouts, as long as delay growth is sub-linear.

Consequently, FedSUM addresses a broader range of practical scenarios with minimal modeling overhead (You et al., 20 Dec 2025).

5. Communication and Computational Complexity

Communication overhead per round across variants is summarized as follows:

  • FedSUM-B: Uplink—1 vector ai,ta_{i,t}9; Downlink—1 vector ii0.
  • FedSUM: Uplink—1 vector ii1; Downlink—2 vectors ii2.
  • FedSUM-CR: Uplink—1 vector ii3; Downlink—1 vector ii4.

FedSUM-B and FedSUM-CR match FedAvg’s minimal downlink cost; FedSUM matches SCAFFOLD for uplink, but doubles the downlink bandwidth. FedSUM-CR trades greater local client memory—storing ii5 of model dimension—for reduced communication, typically favorable when downlink bandwidth is constrained. Computation per round is dominated by ii6 local gradient evaluations, which is uniform across the family. Global server-side storage is limited to maintenance of the control vector ii7 (You et al., 20 Dec 2025).

6. Parameterization and Practical Deployment

Recommended parameter choices for the FedSUM family are as follows:

  • Batch size ii8: Increasing ii9 reduces the variance component τmax⁡\tau_{\max}0 but exacerbates staleness error if clients are inactive across many rounds. Typical values: τmax⁡\tau_{\max}1–τmax⁡\tau_{\max}2.
  • τmax⁡\tau_{\max}3: Set as τmax⁡\tau_{\max}4; conservatively, use an upper bound on τmax⁡\tau_{\max}5 (e.g., τmax⁡\tau_{\max}6 for cyclic regimes).
  • τmax⁡\tau_{\max}7: Not to exceed τmax⁡\tau_{\max}8; initial values often τmax⁡\tau_{\max}9 or TT0, scaled relative to FedAvg by TT1.
  • Local steps TT2: As in FedAvg, generally TT3–TT4; increasing trades off higher computation for fewer communications.
  • Initialization: TT5.
  • Client memory: FedSUM-CR requires TT6 for model dimension TT7, typically negligible compared to the overall model size.

Selection among FedSUM-B, FedSUM, and FedSUM-CR is dictated by communication and memory trade-offs, with FedSUM-CR commonly preferred under tight downlink constraints (You et al., 20 Dec 2025).

7. Significance and Scope

The FedSUM family provides an extensible and communication-efficient foundation for federated nonconvex optimization under fully arbitrary client availability, without imposing restrictive data or participation assumptions. Its delay-metric analysis unifies prior special-case results, and its flexible “stochastic uplink-merge” protocol integrates contributions from both fresh and stale client updates. Its convergence analysis directly addresses the general federated setting encountered in real-world deployments, permitting practitioners to select algorithmic variants suited to system bandwidth and memory constraints, as well as desired computation-to-communication ratios (You et al., 20 Dec 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to FedSUM Family of Algorithms.