Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Unified Framework for Gradient Aggregation in Multi-Objective Optimization

Published 28 May 2026 in cs.LG, cs.AI, and math.OC | (2605.30452v1)

Abstract: Many machine learning problems involve multiple inherent trade-offs that are best addressed by gradient-based multi-objective optimization (MOO) algorithms. Existing methods are often proposed with various motivations, analyzed case by case, and differ algorithmically in how the component gradients are aggregated at each step. In this work, we develop a unifying framework for gradient aggregation in MOO, establishing (optimal) rates of convergence to Pareto stationarity, the standard measure of performance in MOO. Central to our analysis is a sufficient alignment condition, from which we derive a theorem showing that non-conflicting directions, when chosen within the convex hull of gradients, form a fundamental sufficient condition for convergence. We further show that feasibility can be ensured through projection onto the dual cone, broadening the scope of methods that admit convergence guarantees. In parallel, we present a primal optimization perspective of gradient aggregation that encompasses established algorithms, clarifies their theoretical relationships, and enables the design of new variants. As an illustration, we introduce capped MGDA, derived from a CVaR-based formulation, and demonstrate its robustness in adversarial federated learning. Finally, we validate our theory through experiments on synthetic problems and practical benchmarks.

Authors (3)

Summary

  • The paper develops an alignment-based framework showing that convex-hull, non-conflicting gradient directions converge to Pareto stationarity at the optimal O(1/√t) rate under smoothness assumptions.
  • Its primal subproblem formulation unifies MGDA, Nash-MTL, FairGrad, UPGrad, CAGrad, and related methods while identifying when projection or repeated projection restores convergence guarantees.
  • The proposed capped MGDA improves robustness to adversarial gradient conflicts, achieving 0.515 accuracy versus MGDA’s 0.211 in the reported adversarial federated-learning experiment, while matching MGDA without attacks.

Overview

This paper develops a unifying theoretical framework for gradient-based multi-objective optimization (MOO), addressing a long-standing gap: gradient aggregation methods such as MGDA, Nash-MTL, FairGrad, UPGrad, and CAGrad were each proposed under distinct motivations and analyzed case by case, with method-specific assumptions and proofs. The authors—Hu, Ho, and Yu—establish a general alignment condition on update directions that guarantees convergence to Pareto stationarity at an optimal O(1/t)O(1/\sqrt{t}) rate, derive from it a fundamental sufficient condition based on non-conflicting directions within the convex hull of gradients, and show that feasibility can be restored via projection onto the dual cone. They further introduce a primal optimization subproblem perspective that subsumes many existing aggregators as special cases of a single parametric family, and use it to design new methods, most notably capped MGDA, derived from a CVaR formulation. The framework is validated on synthetic MOO problems, a fairness classification benchmark, and adversarial federated learning.

The alignment condition and convergence guarantees

The core technical device is Theorem 1 (the sufficient alignment condition). For the standard MOO update $\wv_{t+1} = \wv_t - \eta_t \dv_t$, suppose there exists an LL-smooth nonnegative surrogate FF such that

$\langle \dv_t, \nabla F(\wv_t) \rangle \geq c_t \Gamma_t \|\dv_t\|, \quad c_t \geq 0,$

where Γt\Gamma_t is a quantity of interest chosen by the analyst. With step size $\eta_t = c_t \Gamma_t / (L\|\dv_t\|)$ and ctc>0c_t \geq c > 0, telescoping the descent lemma yields $\sum_t \Gamma_t^2 \leq 2LF(\wv_0)/c^2$, hence mintTΓt=O(1/T)\min_{t \leq T} \Gamma_t = O(1/\sqrt{T}) and $\wv_{t+1} = \wv_t - \eta_t \dv_t$0. The theorem is deliberately permissive: it imposes essentially no structure on how $\wv_{t+1} = \wv_t - \eta_t \dv_t$1 or $\wv_{t+1} = \wv_t - \eta_t \dv_t$2 is constructed. In the single-objective case with $\wv_{t+1} = \wv_t - \eta_t \dv_t$3, it reduces to the classical angle constraint of Zoutendijk's method of feasible directions; in the multi-objective setting, $\wv_{t+1} = \wv_t - \eta_t \dv_t$4 serves purely as a proof surrogate (typically a linear combination of component objectives), and the guarantee on $\wv_{t+1} = \wv_t - \eta_t \dv_t$5 need not correspond to minimizing $\wv_{t+1} = \wv_t - \eta_t \dv_t$6 itself.

Setting $\wv_{t+1} = \wv_t - \eta_t \dv_t$7 yields the easily verifiable alignment condition (A): $\wv_{t+1} = \wv_t - \eta_t \dv_t$8. Corollary 1 then shows that if $\wv_{t+1} = \wv_t - \eta_t \dv_t$9 lies in the convex hull of the component gradients, LL0 holds trivially, where LL1 is the standard measure of Pareto stationarity. Consequently, any convex-hull direction satisfying (A) converges to Pareto stationarity at rate LL2 with constant step size LL3—a rate the authors note is optimal even for single-objective nonconvex optimization.

The pivotal specialization is Theorem 2: if LL4 and LL5 (i.e., LL6 is non-conflicting, meaning LL7), then condition (A) holds automatically with LL8 and LL9. This is a strong claim: convex-hull membership plus non-conflictedness alone—both easy to check and construct—suffice for optimal-rate convergence to Pareto stationarity, without smoothness assumptions beyond FF0-smoothness or any independence conditions on gradients. The proof is a three-line verification using only non-negativity of inner products. This result retroactively explains the success of MGDA, Nash-MTL, UPGrad, DualProj, and FairGrad, all of which produce such directions, and it removes assumptions required by prior analyses: notably, the original Nash-MTL analysis needed linearly independent gradients and bounded sublevel sets to establish mere subsequence convergence, whereas this framework yields an explicit FF1 rate from Lipschitz smoothness alone.

Two extensions complete the picture. First, Proposition 1 shows via Moreau's decomposition that projecting any conic-combination pre-direction onto the dual cone FF2 yields a direction that is simultaneously in the cone and the dual cone; Corollary 2 thereby extends convergence guarantees to projection-based schemes (DualProj, UPGrad, and the newly proposed Greedy-DCP). Second, because condition (A) is stable under convex combinations, mixed aggregator scheduling (MAS)—alternating among different non-conflicting aggregators within one run—is itself provably convergent, a practical consequence not previously formalized.

In the convex regime, Theorem 3 strengthens the guarantee: for FF3-smooth convex objectives and directions in the convex hull inducing monotone descent, there exists FF4 such that the averaged function value gap satisfies FF5 convergence to the minimum of the scalarized objective FF6. This generalizes Fliege et al.'s MGDA-specific analysis to any interior dual-cone direction, and reduces to classical gradient descent when FF7.

The primal subproblem perspective

The second pillar is a parametric subproblem for constructing update directions:

FF8

with FF9 and increasing convex $\langle \dv_t, \nabla F(\wv_t) \rangle \geq c_t \Gamma_t \|\dv_t\|, \quad c_t \geq 0,$0. Theorem 4 establishes that when $\langle \dv_t, \nabla F(\wv_t) \rangle \geq c_t \Gamma_t \|\dv_t\|, \quad c_t \geq 0,$1 is decreasing convex with $\langle \dv_t, \nabla F(\wv_t) \rangle \geq c_t \Gamma_t \|\dv_t\|, \quad c_t \geq 0,$2 and $\langle \dv_t, \nabla F(\wv_t) \rangle \geq c_t \Gamma_t \|\dv_t\|, \quad c_t \geq 0,$3, the solution lies in the conic hull of gradients (hence its normalization lies in the convex hull) and satisfies condition (A) with $\langle \dv_t, \nabla F(\wv_t) \rangle \geq c_t \Gamma_t \|\dv_t\|, \quad c_t \geq 0,$4, where the dual variables come from the Fenchel conjugate $\langle \dv_t, \nabla F(\wv_t) \rangle \geq c_t \Gamma_t \|\dv_t\|, \quad c_t \geq 0,$5. Since $\langle \dv_t, \nabla F(\wv_t) \rangle \geq c_t \Gamma_t \|\dv_t\|, \quad c_t \geq 0,$6 for quadratic regularization and $\langle \dv_t, \nabla F(\wv_t) \rangle \geq c_t \Gamma_t \|\dv_t\|, \quad c_t \geq 0,$7 whenever the dual constraint forces $\langle \dv_t, \nabla F(\wv_t) \rangle \geq c_t \Gamma_t \|\dv_t\|, \quad c_t \geq 0,$8 into (a superset of) the simplex, the $\langle \dv_t, \nabla F(\wv_t) \rangle \geq c_t \Gamma_t \|\dv_t\|, \quad c_t \geq 0,$9 rate follows immediately.

This single template recovers a striking range of existing methods through the choice of Γt\Gamma_t0: uniform linear scalarization (Γt\Gamma_t1 power mean), Nash-MTL (Γt\Gamma_t2, geometric mean), MGDA (Γt\Gamma_t3, minimax), FairGrad/PIVRG (Γt\Gamma_t4), CAGrad (MGDA's Γt\Gamma_t5 with a ball constraint around the average gradient), IMTL-G, and the projection-based DualProj and UPGrad. The paper also documents where the framework does not apply cleanly: PCGrad's single-pass projections are provably non-conflicting only for Γt\Gamma_t6 (where it coincides with UPGrad); for Γt\Gamma_t7 the authors exhibit an explicit counterexample where PCGrad's output conflicts with a gradient, and they propose PCGrad+, which projects repeatedly until dual-cone membership holds, restoring convergence guarantees for arbitrary Γt\Gamma_t8. IMTL-G is shown via counterexamples to produce directions outside both the cone and dual cone, and RGW admits no fixed point at all, so neither fits the framework.

New methods: capped MGDA and Greedy-DCP

Choosing Γt\Gamma_t9 from the family of convex risk measures—which restrict the dual domain to the simplex—yields new aggregators with automatic convergence guarantees. The headline example is capped MGDA, based on CVaR$\eta_t = c_t \Gamma_t / (L\|\dv_t\|)$0: averaging the tail of the per-objective losses $\eta_t = c_t \Gamma_t / (L\|\dv_t\|)$1 beyond the $\eta_t = c_t \Gamma_t / (L\|\dv_t\|)$2 quantile. Its dual is exactly MGDA's min-norm QP with an additional cap constraint $\eta_t = c_t \Gamma_t / (L\|\dv_t\|)$3. The method interpolates between MGDA ($\eta_t = c_t \Gamma_t / (L\|\dv_t\|)$4) and linear scalarization ($\eta_t = c_t \Gamma_t / (L\|\dv_t\|)$5), and inherits the $\eta_t = c_t \Gamma_t / (L\|\dv_t\|)$6 rate since $\eta_t = c_t \Gamma_t / (L\|\dv_t\|)$7. A second new method, Greedy-DCP, selects the largest-norm dual-cone-projected gradient and converges by Corollary 2.

The motivation for capping comes from an adversarial robustness argument: MGDA's min-norm objective assigns weights near $\eta_t = c_t \Gamma_t / (L\|\dv_t\|)$8 to opposing gradients, collapsing the update norm toward zero. Experiments on CIFAR-10 federated learning with 10 non-i.i.d. clients and one injected adversarial gradient ($\eta_t = c_t \Gamma_t / (L\|\dv_t\|)$9) show a stark contrast: MGDA's global test accuracy drops to 0.211 versus a no-adversary baseline of 0.541, while capped MGDA with ctc>0c_t \geq c > 00 achieves 0.515—comparable to the strongest projection baselines (UPGrad* at 0.511, DualProj* at 0.515). Notably, a naive "MGDA + coefficient clipping" variant with the same threshold performs substantially worse, indicating that the constrained CVaR formulation, not merely bounding coefficients post hoc, is what matters. The paper is candid that capped MGDA is generally not a non-conflicting aggregator—smaller caps push ctc>0c_t \geq c > 01 below zero—so its convergence rests on Theorem 4 rather than Theorem 2. Under no attack, capped MGDA matches MGDA, so the cap carries no cost in benign settings.

Empirical validation of the theory

Experiments serve to validate the theoretical claims rather than rank methods—a caveat the authors state explicitly, noting that Pareto stationary solutions are generally incomparable. On VLMOP2 and Omnitest, all four non-conflicting aggregators drive ctc>0c_t \geq c > 02 to zero, consistent with Theorem 2. A subtle but important empirical observation concerns normalization: normalized variants (rescaled into the convex hull) converge at similar asymptotic rates, while unnormalized variants appear faster only because their larger ctc>0c_t \geq c > 03 acts as an implicit larger step size—Nash-MTL's fixed unit norm causes overshooting near stationarity, which convex-hull normalization fixes both theoretically (restoring Theorem 2 applicability) and empirically. MAS experiments confirm that switching among aggregators mid-run preserves convergence, supporting the closure property of condition (A). On the Adult fairness benchmark (three objectives: cross-entropy, DEO1, DEO2), MAS achieves intermediate performance across accuracy and fairness metrics—for example, MAS-Rand attains 76.42% accuracy with DEO1 of 60.90%, between DualProj (79.02%/53.97%) and MGDA (75.70%/61.04%)—positioning it as a representative baseline rather than a dominant method.

Limitations and open questions

The framework has clear boundaries. The rates are for full-batch deterministic settings; extension to stochastic gradients, constrained problems, and non-smooth objectives is left open. The ctc>0c_t \geq c > 04 nonconvex rate, while optimal in general, is not tightened under stronger curvature—the authors list PL/strong-convexity refinements as future work. Theorem 3 requires monotone function-value descent, which for boundary-of-dual-cone directions needs auxiliary mechanisms (line search or perturbation) not analyzed here. Capped MGDA's cap parameter ctc>0c_t \geq c > 05 trades off conflict-freeness against robustness without a principled selection rule, and its convergence constant depends on ctc>0c_t \geq c > 06 regardless of cap tightness. Finally, MAS is validated only with simple random and round-robin schedules; whether adaptive schedules yield better trade-offs remains unexamined.

Conclusion

This paper provides a coherent account of when gradient aggregation in MOO converges to Pareto stationarity. Its central contributions are the alignment condition underlying Theorem 1, the identification of convex-hull non-conflicting directions as a fundamental sufficient condition (Theorem 2) with optimal rates and minimal assumptions, the dual-cone projection mechanism restoring feasibility, and the primal subproblem template (Theorem 4) that unifies existing methods and generates new ones. The capped MGDA results demonstrate that the framework is not merely taxonomic: it produces practically useful algorithms with automatic guarantees, as evidenced by the adversarial federated learning experiments where capping restores performance lost to gradient-collision attacks.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.