---
title: 'FedNova: Unbiased Federated Optimization'
url: https://www.emergentmind.com/topics/fednova-algorithm
type: topic
---

# FedNova: Unbiased Federated Optimization

FedNova is a federated optimization algorithm designed to address the objective inconsistency problem prevalent in heterogeneous federated learning environments where clients possess varying local data distributions and computation speeds. Standard federated learning protocols such as FedAvg and FedProx aggregate weighted client updates but can converge to a biased solution—minimizing a surrogate rather than the true intended global objective—when client participation or local update counts differ. FedNova introduces a normalized averaging mechanism that precisely eliminates this objective inconsistency while retaining the fast convergence characteristics associated with local-update federated schemes [2007.07481].

## 1. Objective Inconsistency in Federated Learning

The federated learning paradigm targets minimization of the true global objective:
\[ F(x) = \sum_{i=1}^m p_i F_i(x), \]
where $p_i = n_i/n$ denotes the data proportion for client $i$, and $F_i(x)$ is the local empirical risk. In the standard FedAvg procedure, clients start from $x^{(t,0)}$ and perform $\tau_i$ local SGD steps before submitting cumulative model updates $\Delta_i^{(t)}$, which are weighted via $p_i$ at the server:
\[ x^{(t+1,0)} = x^{(t,0)} + \sum_{i=1}^m p_i \Delta_i^{(t)}. \]
When $\tau_i$ is heterogeneous, FedAvg provably converges to the minimizer of a surrogate objective:
\[ \widetilde F(x) = \sum_{i=1}^m w_i F_i(x), \quad w_i = \frac{p_i \tau_i}{\sum_j p_j \tau_j}, \]
rather than $F(x)$. For quadratic objectives $F_i(x) = \frac{1}{2}\|x-e_i\|^2$, the limiting solution is weighted by $\tau_i$ instead of $p_i$. This inconsistency can result in arbitrary solution bias and incorrect global minimization if local steps differ [2007.07481, Lemma 2.1].

## 2. Normalization Mechanism of FedNova

FedNova remedies objective inconsistency by decoupling the effect of per-client local update counts ($\tau_i$) from aggregation weights. On each round, client $i$ computes a normalized local update:
\[ d_i^{(t)} = \frac{x_i^{(t,\tau_i)} - x^{(t,0)}}{\tau_i}, \]
which, under SGD, evaluates to:
\[ d_i^{(t)} = -\eta \frac{1}{\tau_i} \sum_{k=0}^{\tau_i-1} g_i(x_i^{(t,k)}). \]
The server then aggregates using the original data weights $p_i$:
\[ x^{(t+1,0)} = x^{(t,0)} - \eta_g \sum_{i=1}^m p_i d_i^{(t)}, \]
where $\eta_g$ (often set as $\sum_i p_i \tau_i$) is a server-side stepsize matching the aggregate update magnitude of FedAvg. By normalizing updates by $\tau_i$ and re-weighting solely via $p_i$, FedNova ensures aggregation precisely tracks the true global objective, eliminating the bias arising from unnormalized, $\tau_i$-dependent contributions [2007.07481].

## 3. FedNova Procedure

FedNova is compatible with any client-side solver whose updates are linear combinations of local gradients. The procedure, specialized for SGD, is as follows:

1. Each round $t$: 
   - The server broadcasts $x^{(t,0)}$ to all clients.
   - Each client $i$ initializes $x_i = x^{(t,0)}$, performs $\tau_i$ local SGD steps, and computes $d_i = (x_i - x^{(t,0)})/\tau_i$.
   - Each client sends $(p_i d_i, p_i \tau_i)$ to the server.
2. Server aggregates:
   - $\bar{\tau} = \sum_{i} p_i \tau_i$
   - $x^{(t+1,0)} = x^{(t,0)} - (\bar{\tau}) \left( \sum_i p_i d_i \right )$

Only $(p_i d_i, p_i \tau_i)$ need to be communicated per client; aggregation and normalization are handled server-side [2007.07481].

## 4. Convergence Guarantees

FedNova’s convergence analysis assumes:

- (A1) $F_i$ is $L$-smooth ($\nabla F_i$ is $L$-Lipschitz),
- (A2) Stochastic local gradients are unbiased with bounded variance $\sigma^2$,
- (A3) Bounded dissimilarity: for any weights $\{w_i\}$,
  \[ \sum_i w_i \|\nabla F_i(x)\|^2 \leq \| \sum_i w_i \nabla F_i(x) \|^2 + \Delta. \]

With local stepsize $\eta = O(1/\sqrt{m \tilde\tau T})$ (where $\tilde\tau = (1/m) \sum_{t,i} \tau_i(t)$) and server stepsize $\eta_g = \sum_i p_i \tau_i$, the FedNova update ensures:
\[
\min_{t=0,\ldots,T-1} \mathbb{E}\|\nabla F(x^{(t,0)})\|^2 = O\left( 
\frac{\tilde \tau \sigma^2}{\sqrt{m \tilde\tau T}} + 
\frac{L \Delta}{\sqrt{m \tilde\tau T}} +
\frac{m \Delta}{\tilde\tau T} 
\right),
\]
recovering the standard $O(1/\sqrt{mT})$ rate of nonconvex SGD for large $T, m$, and moderate $\Delta$. The solution bias vanishes because effective weights $w_i = p_i$ exactly match the original objective [2007.07481, Theorem 4.1].

## 5. Comparison with FedAvg and FedProx

| Method   | Solution Bias in Heterogeneous $\tau_i$ | Convergence Rate   | Additional Mechanism                  |
|----------|-----------------------------------------|--------------------|---------------------------------------|
| FedAvg   | Nonvanishing; optimizes surrogate $\widetilde F$ | $O(1/\sqrt{mT})$  | Weighted average, unnormalized steps  |
| FedProx  | Reduced with stronger proximal term     | Slower as $\mu\to\infty$ | Adds $\mu\|x-x^{(t,0)}\|^2$ locally  |
| FedNova  | None; $w_i = p_i$                       | $O(1/\sqrt{mT})$   | Normalization by $\tau_i$, reweighted |

FedAvg, if $\tau_i \neq \tau_j$, converges to the minimizer of $\sum_i (p_i \tau_i/\sum_j p_j \tau_j) F_i(x)$, not $F(x)$, incurring persistent bias. FedProx appends a proximal penalty, which can bring weights closer to $p_i$ but only asymptotically as the penalty grows—imposing a convergence speed penalty. FedNova achieves unbiased minimization with $w_i = p_i$ and retains fast communication-efficient convergence [2007.07481]. Empirically, FedNova maintains solution accuracy even with randomly varying $\tau_i$ and outperforms both alternatives in non-IID benchmark settings, e.g., yielding 6–9% higher test accuracy on non-IID CIFAR-10 using VGG-11 after equal communication rounds.

## 6. Practical Compatibility and Extensions

FedNova’s update is compatible with momentum, server-side variance reduction schemes (e.g., SCAFFOLD), and adaptive optimizers. It can be applied to any client-side local solver producing aggregate updates as linear combinations of gradients, making it broadly usable in modern federated learning environments. The normalization mechanism allows deployment in settings with extreme system heterogeneity, variable participation, and variable communication rates without risking degraded convergence or mismatched objectives [2007.07481].

Source: https://www.emergentmind.com/topics/fednova-algorithm