Papers
Topics
Authors
Recent
Search
2000 character limit reached

FedNova: Unbiased Federated Optimization

Updated 26 May 2026
  • FedNova is a federated optimization algorithm that eliminates objective inconsistency by normalizing local updates across heterogeneous clients.
  • It decouples local update counts from aggregation weights, ensuring unbiased convergence to the true global objective regardless of client variability.
  • Empirical results show that FedNova outperforms FedAvg and FedProx on non-IID datasets, achieving faster convergence and higher test accuracy.

FedNova is a federated optimization algorithm designed to address the objective inconsistency problem prevalent in heterogeneous federated learning environments where clients possess varying local data distributions and computation speeds. Standard federated learning protocols such as FedAvg and FedProx aggregate weighted client updates but can converge to a biased solution—minimizing a surrogate rather than the true intended global objective—when client participation or local update counts differ. FedNova introduces a normalized averaging mechanism that precisely eliminates this objective inconsistency while retaining the fast convergence characteristics associated with local-update federated schemes (Wang et al., 2020).

1. Objective Inconsistency in Federated Learning

The federated learning paradigm targets minimization of the true global objective: F(x)=∑i=1mpiFi(x),F(x) = \sum_{i=1}^m p_i F_i(x), where pi=ni/np_i = n_i/n denotes the data proportion for client ii, and Fi(x)F_i(x) is the local empirical risk. In the standard FedAvg procedure, clients start from x(t,0)x^{(t,0)} and perform τi\tau_i local SGD steps before submitting cumulative model updates Δi(t)\Delta_i^{(t)}, which are weighted via pip_i at the server: x(t+1,0)=x(t,0)+∑i=1mpiΔi(t).x^{(t+1,0)} = x^{(t,0)} + \sum_{i=1}^m p_i \Delta_i^{(t)}. When τi\tau_i is heterogeneous, FedAvg provably converges to the minimizer of a surrogate objective: pi=ni/np_i = n_i/n0 rather than pi=ni/np_i = n_i/n1. For quadratic objectives pi=ni/np_i = n_i/n2, the limiting solution is weighted by pi=ni/np_i = n_i/n3 instead of pi=ni/np_i = n_i/n4. This inconsistency can result in arbitrary solution bias and incorrect global minimization if local steps differ [(Wang et al., 2020), Lemma 2.1].

2. Normalization Mechanism of FedNova

FedNova remedies objective inconsistency by decoupling the effect of per-client local update counts (pi=ni/np_i = n_i/n5) from aggregation weights. On each round, client pi=ni/np_i = n_i/n6 computes a normalized local update: pi=ni/np_i = n_i/n7 which, under SGD, evaluates to: pi=ni/np_i = n_i/n8 The server then aggregates using the original data weights pi=ni/np_i = n_i/n9: ii0 where ii1 (often set as ii2) is a server-side stepsize matching the aggregate update magnitude of FedAvg. By normalizing updates by ii3 and re-weighting solely via ii4, FedNova ensures aggregation precisely tracks the true global objective, eliminating the bias arising from unnormalized, ii5-dependent contributions (Wang et al., 2020).

3. FedNova Procedure

FedNova is compatible with any client-side solver whose updates are linear combinations of local gradients. The procedure, specialized for SGD, is as follows:

  1. Each round ii6:
    • The server broadcasts ii7 to all clients.
    • Each client ii8 initializes ii9, performs Fi(x)F_i(x)0 local SGD steps, and computes Fi(x)F_i(x)1.
    • Each client sends Fi(x)F_i(x)2 to the server.
  2. Server aggregates:
    • Fi(x)F_i(x)3
    • Fi(x)F_i(x)4

Only Fi(x)F_i(x)5 need to be communicated per client; aggregation and normalization are handled server-side (Wang et al., 2020).

4. Convergence Guarantees

FedNova’s convergence analysis assumes:

  • (A1) Fi(x)F_i(x)6 is Fi(x)F_i(x)7-smooth (Fi(x)F_i(x)8 is Fi(x)F_i(x)9-Lipschitz),
  • (A2) Stochastic local gradients are unbiased with bounded variance x(t,0)x^{(t,0)}0,
  • (A3) Bounded dissimilarity: for any weights x(t,0)x^{(t,0)}1,

x(t,0)x^{(t,0)}2

With local stepsize x(t,0)x^{(t,0)}3 (where x(t,0)x^{(t,0)}4) and server stepsize x(t,0)x^{(t,0)}5, the FedNova update ensures: x(t,0)x^{(t,0)}6 recovering the standard x(t,0)x^{(t,0)}7 rate of nonconvex SGD for large x(t,0)x^{(t,0)}8, and moderate x(t,0)x^{(t,0)}9. The solution bias vanishes because effective weights τi\tau_i0 exactly match the original objective [(Wang et al., 2020), Theorem 4.1].

5. Comparison with FedAvg and FedProx

Method Solution Bias in Heterogeneous τi\tau_i1 Convergence Rate Additional Mechanism
FedAvg Nonvanishing; optimizes surrogate τi\tau_i2 τi\tau_i3 Weighted average, unnormalized steps
FedProx Reduced with stronger proximal term Slower as τi\tau_i4 Adds τi\tau_i5 locally
FedNova None; τi\tau_i6 τi\tau_i7 Normalization by τi\tau_i8, reweighted

FedAvg, if τi\tau_i9, converges to the minimizer of Δi(t)\Delta_i^{(t)}0, not Δi(t)\Delta_i^{(t)}1, incurring persistent bias. FedProx appends a proximal penalty, which can bring weights closer to Δi(t)\Delta_i^{(t)}2 but only asymptotically as the penalty grows—imposing a convergence speed penalty. FedNova achieves unbiased minimization with Δi(t)\Delta_i^{(t)}3 and retains fast communication-efficient convergence (Wang et al., 2020). Empirically, FedNova maintains solution accuracy even with randomly varying Δi(t)\Delta_i^{(t)}4 and outperforms both alternatives in non-IID benchmark settings, e.g., yielding 6–9% higher test accuracy on non-IID CIFAR-10 using VGG-11 after equal communication rounds.

6. Practical Compatibility and Extensions

FedNova’s update is compatible with momentum, server-side variance reduction schemes (e.g., SCAFFOLD), and adaptive optimizers. It can be applied to any client-side local solver producing aggregate updates as linear combinations of gradients, making it broadly usable in modern federated learning environments. The normalization mechanism allows deployment in settings with extreme system heterogeneity, variable participation, and variable communication rates without risking degraded convergence or mismatched objectives (Wang et al., 2020).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to FedNova Algorithm.