Papers
Topics
Authors
Recent
Search
2000 character limit reached

Multiobjective Balanced Gradient Flow (MBGF)

Updated 7 July 2026
  • The paper introduces MBGF as a continuous-time dynamical system that leverages the minimum-norm convex combination of normalized gradients for multiobjective optimization.
  • It establishes global well-posedness and convergence to weak Pareto points in convex settings using L-smoothness, bounded gradients, and a regularization parameter η.
  • MBGF ensures simultaneous descent across all objectives and provides explicit convergence rates of O(1/t) in convex and O(1/√t) in non-convex scenarios.

Searching arXiv for the MBGF paper and closely related gradient-flow work to support the article. Multiobjective Balanced Gradient Flow (MBGF) is a continuous-time dynamical system for unconstrained multiobjective optimization that models a class of normalized-gradient methods through a projection onto the convex hull of normalized objective gradients. It is introduced for problems of the form

minxRnF(x)=(f1(x),,fm(x)),\min_{x\in\mathbb{R}^n} F(x)=(f_1(x),\dots,f_m(x))^\top,

with smooth objective functions, and is analyzed as a flow whose trajectories admit existence results, convergence to weak Pareto points in the convex case, and explicit convergence rates in both convex and non-convex settings (Yin, 3 Aug 2025). The construction is “balanced” because it selects the minimum-norm convex combination of normalized gradients, thereby reducing the influence of disparate gradient magnitudes across objectives rather than following a single objective or a scalarized surrogate (Yin, 3 Aug 2025).

1. Definition and geometric structure

MBGF is defined through the differential inclusion, equivalently differential equation,

x˙(t)+projCη(x(t))(0)=0,x˙(t)=projCη(x(t))(0),\dot{x}(t)+\operatorname{proj}_{C_\eta(x(t))}(0)=0, \qquad \dot{x}(t)=-\operatorname{proj}_{C_\eta(x(t))}(0),

where

Cη(x):=conv{fi(x)fi(x)+η  |  i=1,,m},η0.C_\eta(x):=\operatorname{conv}\left\{\frac{\nabla f_i(x)}{\|\nabla f_i(x)\|+\eta}\; \middle|\; i=1,\dots,m\right\}, \qquad \eta\ge 0.

Here, projCη(x)(0)\operatorname{proj}_{C_\eta(x)}(0) is the minimum-norm element of the convex hull of the normalized gradients (Yin, 3 Aug 2025).

A useful equivalent condition is

x˙(t)Cη(x(t)),\dot{x}(t)\in -C_\eta(x(t)),

with the actual velocity chosen as the projection of $0$ onto Cη(x(t))C_\eta(x(t)) (Yin, 3 Aug 2025). This formulation makes the balancing mechanism explicit: the dynamics do not select one objective gradient, nor an arbitrary convex combination, but rather the minimum-norm compromise among the normalized gradients.

The flow is described as the continuous-time analogue of normalized multiobjective gradient methods. The normalization is intended to reduce the effect of disparate gradient magnitudes across objectives. At the same time, MBGF is not merely a multiobjective restatement of the single-objective normalized gradient flow

x˙(t)+f(x(t))f(x(t))=0,\dot{x}(t)+\frac{\nabla f(x(t))}{\|\nabla f(x(t))\|}=0,

because in the multiobjective case the speed x˙(t)\|\dot{x}(t)\| is generally not identically $1$ (Yin, 3 Aug 2025).

The regularization parameter x˙(t)+projCη(x(t))(0)=0,x˙(t)=projCη(x(t))(0),\dot{x}(t)+\operatorname{proj}_{C_\eta(x(t))}(0)=0, \qquad \dot{x}(t)=-\operatorname{proj}_{C_\eta(x(t))}(0),0 controls the normalization. For x˙(t)+projCη(x(t))(0)=0,x˙(t)=projCη(x(t))(0),\dot{x}(t)+\operatorname{proj}_{C_\eta(x(t))}(0)=0, \qquad \dot{x}(t)=-\operatorname{proj}_{C_\eta(x(t))}(0),1, it avoids division by small gradient norms; for x˙(t)+projCη(x(t))(0)=0,x˙(t)=projCη(x(t))(0),\dot{x}(t)+\operatorname{proj}_{C_\eta(x(t))}(0)=0, \qquad \dot{x}(t)=-\operatorname{proj}_{C_\eta(x(t))}(0),2, it yields exact normalization (Yin, 3 Aug 2025). This suggests that x˙(t)+projCη(x(t))(0)=0,x˙(t)=projCη(x(t))(0),\dot{x}(t)+\operatorname{proj}_{C_\eta(x(t))}(0)=0, \qquad \dot{x}(t)=-\operatorname{proj}_{C_\eta(x(t))}(0),3 plays both an analytical and numerical regularization role.

2. Existence theory and well-posedness

For the Cauchy problem

x˙(t)+projCη(x(t))(0)=0,x˙(t)=projCη(x(t))(0),\dot{x}(t)+\operatorname{proj}_{C_\eta(x(t))}(0)=0, \qquad \dot{x}(t)=-\operatorname{proj}_{C_\eta(x(t))}(0),4

the existence analysis assumes three conditions: each objective x˙(t)+projCη(x(t))(0)=0,x˙(t)=projCη(x(t))(0),\dot{x}(t)+\operatorname{proj}_{C_\eta(x(t))}(0)=0, \qquad \dot{x}(t)=-\operatorname{proj}_{C_\eta(x(t))}(0),5 is x˙(t)+projCη(x(t))(0)=0,x˙(t)=projCη(x(t))(0),\dot{x}(t)+\operatorname{proj}_{C_\eta(x(t))}(0)=0, \qquad \dot{x}(t)=-\operatorname{proj}_{C_\eta(x(t))}(0),6-smooth, each gradient is uniformly bounded by a constant x˙(t)+projCη(x(t))(0)=0,x˙(t)=projCη(x(t))(0),\dot{x}(t)+\operatorname{proj}_{C_\eta(x(t))}(0)=0, \qquad \dot{x}(t)=-\operatorname{proj}_{C_\eta(x(t))}(0),7, and every level set x˙(t)+projCη(x(t))(0)=0,x˙(t)=projCη(x(t))(0),\dot{x}(t)+\operatorname{proj}_{C_\eta(x(t))}(0)=0, \qquad \dot{x}(t)=-\operatorname{proj}_{C_\eta(x(t))}(0),8 is bounded (Yin, 3 Aug 2025).

The x˙(t)+projCη(x(t))(0)=0,x˙(t)=projCη(x(t))(0),\dot{x}(t)+\operatorname{proj}_{C_\eta(x(t))}(0)=0, \qquad \dot{x}(t)=-\operatorname{proj}_{C_\eta(x(t))}(0),9-smoothness condition is

Cη(x):=conv{fi(x)fi(x)+η  |  i=1,,m},η0.C_\eta(x):=\operatorname{conv}\left\{\frac{\nabla f_i(x)}{\|\nabla f_i(x)\|+\eta}\; \middle|\; i=1,\dots,m\right\}, \qquad \eta\ge 0.0

and the bounded-gradient assumption is

Cη(x):=conv{fi(x)fi(x)+η  |  i=1,,m},η0.C_\eta(x):=\operatorname{conv}\left\{\frac{\nabla f_i(x)}{\|\nabla f_i(x)\|+\eta}\; \middle|\; i=1,\dots,m\right\}, \qquad \eta\ge 0.1

Under these assumptions, and for Cη(x):=conv{fi(x)fi(x)+η  |  i=1,,m},η0.C_\eta(x):=\operatorname{conv}\left\{\frac{\nabla f_i(x)}{\|\nabla f_i(x)\|+\eta}\; \middle|\; i=1,\dots,m\right\}, \qquad \eta\ge 0.2, the set-valued map Cη(x):=conv{fi(x)fi(x)+η  |  i=1,,m},η0.C_\eta(x):=\operatorname{conv}\left\{\frac{\nabla f_i(x)}{\|\nabla f_i(x)\|+\eta}\; \middle|\; i=1,\dots,m\right\}, \qquad \eta\ge 0.3 is shown to be Hausdorff-Lipschitz:

Cη(x):=conv{fi(x)fi(x)+η  |  i=1,,m},η0.C_\eta(x):=\operatorname{conv}\left\{\frac{\nabla f_i(x)}{\|\nabla f_i(x)\|+\eta}\; \middle|\; i=1,\dots,m\right\}, \qquad \eta\ge 0.4

This yields upper semicontinuity of Cη(x):=conv{fi(x)fi(x)+η  |  i=1,,m},η0.C_\eta(x):=\operatorname{conv}\left\{\frac{\nabla f_i(x)}{\|\nabla f_i(x)\|+\eta}\; \middle|\; i=1,\dots,m\right\}, \qquad \eta\ge 0.5, which allows an existence theorem for differential inclusions to be applied (Yin, 3 Aug 2025).

The resulting statement is that, under the stated assumptions, MBGF admits global absolutely continuous trajectories for Cη(x):=conv{fi(x)fi(x)+η  |  i=1,,m},η0.C_\eta(x):=\operatorname{conv}\left\{\frac{\nabla f_i(x)}{\|\nabla f_i(x)\|+\eta}\; \middle|\; i=1,\dots,m\right\}, \qquad \eta\ge 0.6 (Yin, 3 Aug 2025). In the terminology of the paper, local existence is extended to a global solution on Cη(x):=conv{fi(x)fi(x)+η  |  i=1,,m},η0.C_\eta(x):=\operatorname{conv}\left\{\frac{\nabla f_i(x)}{\|\nabla f_i(x)\|+\eta}\; \middle|\; i=1,\dots,m\right\}, \qquad \eta\ge 0.7 (Yin, 3 Aug 2025). A plausible implication is that the regularized normalization Cη(x):=conv{fi(x)fi(x)+η  |  i=1,,m},η0.C_\eta(x):=\operatorname{conv}\left\{\frac{\nabla f_i(x)}{\|\nabla f_i(x)\|+\eta}\; \middle|\; i=1,\dots,m\right\}, \qquad \eta\ge 0.8 is central to the present well-posedness theory, since the Hausdorff-Lipschitz estimate is stated only in that regime.

3. Descent mechanism and level-set dynamics

A central inequality along MBGF trajectories is

Cη(x):=conv{fi(x)fi(x)+η  |  i=1,,m},η0.C_\eta(x):=\operatorname{conv}\left\{\frac{\nabla f_i(x)}{\|\nabla f_i(x)\|+\eta}\; \middle|\; i=1,\dots,m\right\}, \qquad \eta\ge 0.9

for each objective index projCη(x)(0)\operatorname{proj}_{C_\eta(x)}(0)0 (Yin, 3 Aug 2025). Since

projCη(x)(0)\operatorname{proj}_{C_\eta(x)}(0)1

this implies that every objective function is nonincreasing along the trajectory (Yin, 3 Aug 2025).

This simultaneous monotonicity is one of the defining features of MBGF. The velocity is chosen from the convex hull of normalized gradients, but the projection structure ensures that the resulting direction remains a common descent direction in the sense encoded by the inequality above. The paper further deduces the nesting property

projCη(x)(0)\operatorname{proj}_{C_\eta(x)}(0)2

showing that the trajectory evolves through a decreasing family of multiobjective level sets (Yin, 3 Aug 2025).

The balancing interpretation is tied directly to the geometry of

projCη(x)(0)\operatorname{proj}_{C_\eta(x)}(0)3

Two effects are emphasized. First, magnitude normalization prevents large gradient norms from dominating purely because of scale. Second, the projection projCη(x)(0)\operatorname{proj}_{C_\eta(x)}(0)4 selects the convex combination of normalized gradients with smallest norm, described as the best common compromise direction (Yin, 3 Aug 2025). This distinguishes MBGF from methods that optimize one objective at a time or from fixed scalarization procedures.

4. Weak Pareto convergence in the convex regime

In the convex case, MBGF trajectories converge to weak Pareto points (Yin, 3 Aug 2025). The relevant notion is weak Pareto optimality: a point projCη(x)(0)\operatorname{proj}_{C_\eta(x)}(0)5 is weakly Pareto optimal if there is no projCη(x)(0)\operatorname{proj}_{C_\eta(x)}(0)6 such that

projCη(x)(0)\operatorname{proj}_{C_\eta(x)}(0)7

meaning projCη(x)(0)\operatorname{proj}_{C_\eta(x)}(0)8 for every objective projCη(x)(0)\operatorname{proj}_{C_\eta(x)}(0)9 (Yin, 3 Aug 2025). This is weaker than ordinary Pareto optimality, which excludes x˙(t)Cη(x(t)),\dot{x}(t)\in -C_\eta(x(t)),0 with at least one strict inequality (Yin, 3 Aug 2025).

The analysis uses the merit function

x˙(t)Cη(x(t)),\dot{x}(t)\in -C_\eta(x(t)),1

The cited properties are

x˙(t)Cη(x(t)),\dot{x}(t)\in -C_\eta(x(t)),2

and x˙(t)Cη(x(t)),\dot{x}(t)\in -C_\eta(x(t)),3 is lower semicontinuous (Yin, 3 Aug 2025). Thus x˙(t)Cη(x(t)),\dot{x}(t)\in -C_\eta(x(t)),4 serves as a weak-Pareto gap function.

For smooth convex objectives with Lipschitz gradients, the convergence proof combines monotonic decay of each x˙(t)Cη(x(t)),\dot{x}(t)\in -C_\eta(x(t)),5, boundedness of trajectories, Opial’s lemma, and the fact that x˙(t)Cη(x(t)),\dot{x}(t)\in -C_\eta(x(t)),6 (Yin, 3 Aug 2025). The conclusion is that any MBGF trajectory converges to a point x˙(t)Cη(x(t)),\dot{x}(t)\in -C_\eta(x(t)),7 satisfying

x˙(t)Cη(x(t)),\dot{x}(t)\in -C_\eta(x(t)),8

hence x˙(t)Cη(x(t)),\dot{x}(t)\in -C_\eta(x(t)),9 (Yin, 3 Aug 2025).

This places MBGF within the class of continuous-time Pareto-descent systems whose asymptotic behavior is characterized through merit-function decay rather than through scalarized objective convergence. A plausible implication is that the flow provides a continuous-time explanation for why normalized multiobjective descent schemes can converge to weak Pareto configurations without privileging any single objective scale.

5. Convergence rates and stationarity bounds

MBGF admits distinct rate statements in convex and non-convex settings (Yin, 3 Aug 2025). In the convex case, the rate is given for the weak-Pareto merit function. For smooth convex objectives,

$0$0

where $0$1 is a radius bound coming from the initial level-set geometry and $0$2 is a bound used in the analysis of the normalized flow (Yin, 3 Aug 2025). In the special case $0$3, this becomes

$0$4

Accordingly, MBGF achieves an $0$5 decay of the merit function measuring distance to weak Pareto optimality (Yin, 3 Aug 2025).

In the non-convex case, the result is formulated as a stationarity-type bound on the projected norm:

$0$6

For $0$7, under the stated nondegeneracy assumption, an analogous $0$8 estimate is proved with $0$9 replaced by a positive quantity Cη(x(t))C_\eta(x(t))0 derived from gradient bounds on the initial level set:

Cη(x(t))C_\eta(x(t))1

Thus the non-convex rate controls the smallest projected descent norm along the trajectory up to time Cη(x(t))C_\eta(x(t))2, interpreted in the paper as a measure of approximate first-order stationarity (Yin, 3 Aug 2025).

The two rate statements separate global optimality geometry from local stationarity geometry. In convex problems, the bound is on a merit function that vanishes exactly at weak Pareto points. In non-convex problems, the bound concerns the vanishing of the projection norm of the balanced normalized-gradient hull. This suggests a direct correspondence between MBGF’s vector field and the natural stationarity notion for normalized multiobjective descent.

6. Relation to adjacent balanced and gradient-flow formulations

MBGF belongs to a broader landscape of multiobjective flow-based methods, but its construction is specific: Euclidean continuous time, unconstrained optimization, normalized gradients, and minimum-norm projection on their convex hull (Yin, 3 Aug 2025). Closely related work uses different geometries or different balancing mechanisms.

A useful comparison is with "Multi-Objective Optimization via Wasserstein-Fisher-Rao Gradient Flow" (Ren et al., 2023), which does not use the term MBGF but is described as conceptually aligned with balanced gradient flow. Its balance is between a Wasserstein transport component and a Fisher-Rao birth-death component, implemented by a splitting scheme alternating overdamped Langevin dynamics and birth-death dynamics (Ren et al., 2023). In that framework, the dominance potential

Cη(x(t))C_\eta(x(t))3

is used to eliminate dominated particles and support global Pareto optimality (Ren et al., 2023). The similarity to MBGF lies in the balancing intent; the difference lies in the state space and in the mechanism, since WFR methods operate over distributions rather than trajectories in Cη(x(t))C_\eta(x(t))4.

A second adjacent line is "Multiple Wasserstein Gradient Descent Algorithm for Multi-Objective Distributional Optimization" (Nguyen et al., 24 May 2025). That method is not called MBGF, but it constructs a multiobjective Wasserstein gradient flow over probability measures by combining objective-wise Wasserstein gradients through simplex weights chosen by a min-norm oracle:

Cη(x(t))C_\eta(x(t))5

The resulting velocity is

Cη(x(t))C_\eta(x(t))6

which is explicitly a convex combination chosen to maximize the minimal improvement across objectives (Nguyen et al., 24 May 2025). This is structurally close to the balancing principle of MBGF, although it is formulated in Wasserstein space rather than Euclidean space.

Acceleration also appears in a later Euclidean flow, "Time Scaling Makes Accelerated Gradient Flow and Proximal Method Faster in Multiobjective Optimization" (Yin, 10 Aug 2025). That paper studies a second-order inertial multiobjective flow with a time-scaled convex-hull gradient term and proves rates faster than Cη(x(t))C_\eta(x(t))7 under appropriate parameter choices (Yin, 10 Aug 2025). In contrast, MBGF is first-order and normalized, with convergence rates of Cη(x(t))C_\eta(x(t))8 in the convex case and Cη(x(t))C_\eta(x(t))9 in the non-convex case (Yin, 3 Aug 2025). The comparison indicates that “balanced gradient flow” is not a single formal template but a family of constructions sharing Pareto-aware compromise directions while differing in normalization, geometry, inertia, and state space.

7. Significance, scope, and interpretation

The main contributions attributed to MBGF are the introduction of a new dynamical system for normalized multiobjective descent, a well-posedness theory for x˙(t)+f(x(t))f(x(t))=0,\dot{x}(t)+\frac{\nabla f(x(t))}{\|\nabla f(x(t))\|}=0,0, convergence to weak Pareto points in the convex setting, and explicit rates in both convex and non-convex regimes (Yin, 3 Aug 2025). The framework is presented as a continuous-time explanation of normalized gradient strategies for unbalanced multiobjective problems, especially where objective gradients have very different scales (Yin, 3 Aug 2025).

Its significance lies in the combination of three elements. First, the flow is built directly from the convex hull of normalized gradients, rather than from scalarization. Second, it preserves simultaneous descent of all objectives along trajectories. Third, it admits a rigorous bridge between heuristic normalization and continuous-time convergence analysis (Yin, 3 Aug 2025). This suggests that MBGF can serve as a theoretical foundation for algorithm design in settings where scale imbalance among objectives is a primary obstacle.

A common misconception would be to identify MBGF with a unit-speed normalized flow by analogy with the single-objective equation x˙(t)+f(x(t))f(x(t))=0,\dot{x}(t)+\frac{\nabla f(x(t))}{\|\nabla f(x(t))\|}=0,1. The paper explicitly rejects that equivalence: in the multiobjective case, the norm of the velocity is generally not identically x˙(t)+f(x(t))f(x(t))=0,\dot{x}(t)+\frac{\nabla f(x(t))}{\|\nabla f(x(t))\|}=0,2 (Yin, 3 Aug 2025). Another possible misconception would be to equate MBGF with arbitrary convex combinations of gradients. MBGF instead uses the minimum-norm projection from the convex hull of normalized gradients, which is the specific source of its “balanced” character (Yin, 3 Aug 2025).

Within the current literature, MBGF is most precisely understood as a first-order Euclidean normalized-gradient flow for multiobjective optimization, with convergence to weak Pareto points under convexity and stationarity-type guarantees beyond convexity (Yin, 3 Aug 2025). Related methods in Wasserstein space, accelerated multiobjective flows, and scaled discrete proximal methods illuminate neighboring approaches to objective balancing, but they do not replace the specific role of MBGF as a continuous-time normalized compromise flow (Ren et al., 2023, Nguyen et al., 24 May 2025, Yin, 10 Aug 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Multiobjective Balanced Gradient Flow (MBGF).