FedNova: Unbiased Federated Optimization
- FedNova is a federated optimization algorithm that eliminates objective inconsistency by normalizing local updates across heterogeneous clients.
- It decouples local update counts from aggregation weights, ensuring unbiased convergence to the true global objective regardless of client variability.
- Empirical results show that FedNova outperforms FedAvg and FedProx on non-IID datasets, achieving faster convergence and higher test accuracy.
FedNova is a federated optimization algorithm designed to address the objective inconsistency problem prevalent in heterogeneous federated learning environments where clients possess varying local data distributions and computation speeds. Standard federated learning protocols such as FedAvg and FedProx aggregate weighted client updates but can converge to a biased solution—minimizing a surrogate rather than the true intended global objective—when client participation or local update counts differ. FedNova introduces a normalized averaging mechanism that precisely eliminates this objective inconsistency while retaining the fast convergence characteristics associated with local-update federated schemes (Wang et al., 2020).
1. Objective Inconsistency in Federated Learning
The federated learning paradigm targets minimization of the true global objective: where denotes the data proportion for client , and is the local empirical risk. In the standard FedAvg procedure, clients start from and perform local SGD steps before submitting cumulative model updates , which are weighted via at the server: When is heterogeneous, FedAvg provably converges to the minimizer of a surrogate objective: 0 rather than 1. For quadratic objectives 2, the limiting solution is weighted by 3 instead of 4. This inconsistency can result in arbitrary solution bias and incorrect global minimization if local steps differ [(Wang et al., 2020), Lemma 2.1].
2. Normalization Mechanism of FedNova
FedNova remedies objective inconsistency by decoupling the effect of per-client local update counts (5) from aggregation weights. On each round, client 6 computes a normalized local update: 7 which, under SGD, evaluates to: 8 The server then aggregates using the original data weights 9: 0 where 1 (often set as 2) is a server-side stepsize matching the aggregate update magnitude of FedAvg. By normalizing updates by 3 and re-weighting solely via 4, FedNova ensures aggregation precisely tracks the true global objective, eliminating the bias arising from unnormalized, 5-dependent contributions (Wang et al., 2020).
3. FedNova Procedure
FedNova is compatible with any client-side solver whose updates are linear combinations of local gradients. The procedure, specialized for SGD, is as follows:
- Each round 6:
- The server broadcasts 7 to all clients.
- Each client 8 initializes 9, performs 0 local SGD steps, and computes 1.
- Each client sends 2 to the server.
- Server aggregates:
- 3
- 4
Only 5 need to be communicated per client; aggregation and normalization are handled server-side (Wang et al., 2020).
4. Convergence Guarantees
FedNova’s convergence analysis assumes:
- (A1) 6 is 7-smooth (8 is 9-Lipschitz),
- (A2) Stochastic local gradients are unbiased with bounded variance 0,
- (A3) Bounded dissimilarity: for any weights 1,
2
With local stepsize 3 (where 4) and server stepsize 5, the FedNova update ensures: 6 recovering the standard 7 rate of nonconvex SGD for large 8, and moderate 9. The solution bias vanishes because effective weights 0 exactly match the original objective [(Wang et al., 2020), Theorem 4.1].
5. Comparison with FedAvg and FedProx
| Method | Solution Bias in Heterogeneous 1 | Convergence Rate | Additional Mechanism |
|---|---|---|---|
| FedAvg | Nonvanishing; optimizes surrogate 2 | 3 | Weighted average, unnormalized steps |
| FedProx | Reduced with stronger proximal term | Slower as 4 | Adds 5 locally |
| FedNova | None; 6 | 7 | Normalization by 8, reweighted |
FedAvg, if 9, converges to the minimizer of 0, not 1, incurring persistent bias. FedProx appends a proximal penalty, which can bring weights closer to 2 but only asymptotically as the penalty grows—imposing a convergence speed penalty. FedNova achieves unbiased minimization with 3 and retains fast communication-efficient convergence (Wang et al., 2020). Empirically, FedNova maintains solution accuracy even with randomly varying 4 and outperforms both alternatives in non-IID benchmark settings, e.g., yielding 6–9% higher test accuracy on non-IID CIFAR-10 using VGG-11 after equal communication rounds.
6. Practical Compatibility and Extensions
FedNova’s update is compatible with momentum, server-side variance reduction schemes (e.g., SCAFFOLD), and adaptive optimizers. It can be applied to any client-side local solver producing aggregate updates as linear combinations of gradients, making it broadly usable in modern federated learning environments. The normalization mechanism allows deployment in settings with extreme system heterogeneity, variable participation, and variable communication rates without risking degraded convergence or mismatched objectives (Wang et al., 2020).