---
title: Primal-Dual Dynamical Formulation
url: https://www.emergentmind.com/topics/primal-dual-dynamical-formulation
type: topic
---

# Primal-Dual Dynamical Formulation

Primal-dual dynamical formulation denotes a class of continuous-time constructions in which primal variables and dual variables evolve jointly so that equilibria coincide with Karush–Kuhn–Tucker points, saddle points of a Lagrangian, or zeros of a monotone inclusion. In recent work, the term covers first-order saddle and differential-inclusion models, mixed second-order primal and first-order dual inertial systems, projected product-space flows, preconditioned monotone-inclusion dynamics, prediction-correction systems for time-varying optimization, and even infinite-dimensional port-Hamiltonian PDEs for convex optimal control [2103.12931], [1905.08290], [2102.03807], [2312.10203], [2506.00501], [2606.18818].

## 1. Mathematical core

The common starting point is a constrained convex optimization problem or a structured monotone inclusion written in primal-dual form. For linear equality constrained convex minimization, a standard model is
\[
\min_x f(x)\qquad \text{s.t. }Ax=b,
\]
with Lagrangian
\[
\mathcal L(x,\lambda)=f(x)+\langle \lambda,Ax-b\rangle,
\]
and KKT conditions
\[
\nabla f(x^*)+A^\top\lambda^*=0,\qquad Ax^*-b=0.
\]
In this setting, a primal-dual dynamical system is constructed so that its stationary points reproduce exactly these optimality relations; the same pattern appears in augmented-Lagrangian variants, separable models with auxiliary primal blocks, and saddle formulations in Hilbert spaces [2103.12931], [2106.12294], [2408.06884], [2603.29124].

A second canonical route is operator-theoretic. Instead of starting from a differentiable Lagrangian, one writes a coupled inclusion such as
\[
0\in Ap+L^*BLp,
\]
or equivalently the product-space Kuhn–Tucker system
\[
-L^*v^*\in Ap,\qquad Lp\in B^{-1}v^*.
\]
The corresponding dynamics then evolve directly in \(\mathcal H\times\mathcal G\), frequently through resolvents, Yosida approximations, and projection operators rather than explicit gradients [2102.03807], [2506.00501].

A third extension replaces static optimality by time-varying KKT tracking. For inequality-constrained time-varying convex programs, the optimizer trajectory \((x^\star(t),\lambda^\star(t))\) itself satisfies a differentiated KKT system
\[
J(x^\star,\lambda^\star,t)\begin{bmatrix}\dot x^\star\\ \dot\lambda^\star\end{bmatrix}+H(x^\star,\lambda^\star,t)=0,
\]
and the primal-dual dynamics are designed to track this moving solution rather than a fixed saddle point [2312.10203].

## 2. Principal architectures

The literature does not reduce to a single archetype. Several distinct continuous-time designs coexist.

| Family | Representative form | Distinctive feature |
|---|---|---|
| Augmented-Lagrangian first-order flow | Resolvent/proximal differential inclusion in \((x,z,y)\) | Full splitting for \(f\), \(h\), \(g\), \(A\) |
| Mixed-order inertial flow | Second-order primal, first-order dual | Acceleration concentrated on the primal side |
| Fully second-order saddle flow | Second-order equations in both primal and dual variables | Vanishing damping and \(O(1/t^2)\)-type behavior |
| Projection-based product-space flow | \(\dot x=Q(\bar w,x,Tx)-x\) | Haugazeau best-approximation geometry |
| Preconditioned monotone-inclusion flow | \(P(t)\dot z(t)+Mz(t)\ni 0\) | Continuous analogue of Chambolle–Pock preconditioning |
| Projected time-varying KKT flow | Prediction-correction with \(\dot\lambda=[\cdot]^+_\lambda\) | Tracking of moving primal-dual optima |
| Port-Hamiltonian PDE flow | \(\partial_\tau(Xx,Uu,Pp)^T=\cdots\) | Infinite-dimensional optimal control formulation |

Structured first-order primal-dual flows arise naturally in composite convex minimization. One representative system evolves \((x,z,y)\) through proximal subproblems for \(f\) and \(g\), a forward term for \(\nabla h\), and a dual ascent law on \(Ax-z\); it is explicitly formulated as a continuous-time counterpart of a combination of the linearized proximal method of multipliers and proximal ADMM [1905.08290].

Mixed-order architectures have become especially prominent. In one line of work, the primal variable obeys
\[
\ddot{x}(t)+\gamma\dot{x}(t)=-\beta(t)\nabla_x\mathcal L^\sigma(x(t),\lambda(t)),
\]
while the dual variable follows
\[
\dot{\lambda}(t)=\beta(t)\nabla_\lambda\mathcal L^\sigma(x(t)+\delta\dot x(t),\lambda(t)),
\]
producing a deliberately asymmetric “second-order primal + first-order dual” system [2103.12931]. Closely related separable and regularized variants use two second-order primal equations and one first-order dual equation [2408.06884], while more elaborate versions add variable mass, Hessian-driven damping, extrapolation, and time-dependent Tikhonov regularization [2603.29124]. Closed-loop controlled models retain the same mixed-order structure but generate the scaling and damping coefficients from a feedback law involving the current Lagrangian gradient [2602.16402].

A different family remains fully second order in the saddle variables. For equality-constrained convex optimization with augmented Lagrangian \(\mathcal L_\beta\), one representative flow is
\[
\ddot x(t)+\frac{\alpha}{t}\dot x(t)+\nabla_x\mathcal L_\beta(x(t),\lambda(t)+\theta t\dot\lambda(t))=0,
\]
\[
\ddot \lambda(t)+\frac{\alpha}{t}\dot \lambda(t)-\nabla_\lambda\mathcal L_\beta(x(t)+\theta t\dot x(t),\lambda(t))=0,
\]
which extends the Su–Boyd–Candès/Nesterov vanishing-damping paradigm to a primal-dual augmented-Lagrangian setting [2106.12294].

Projection-based and preconditioned flows depart further from classical saddle-gradient intuition. The Haugazeau-type system
\[
\dot x(t)=Q(\bar w,x(t),Tx(t))-x(t)
\]
is a first-order autonomous ODE in product space whose asymptotic state is the best approximation \(P_Z(\bar w)\) from the primal-dual solution set [2102.03807]. The preconditioned inclusion
\[
P(t)\dot z(t)+Mz(t)\ni 0,\qquad
P(t)=\begin{pmatrix}\alpha(t)I&\beta(t)A^*\\ \gamma(t)A&\delta(t)I\end{pmatrix},
\]
is instead the continuous-time counterpart of a preconditioned proximal-point interpretation of Chambolle–Pock [2506.00501].

## 3. Structural design mechanisms

A large part of the subject is the systematic addition of structural terms that alter asymptotic selection, damping, or metric geometry.

Tikhonov regularization is used to bias the dynamics toward minimal-norm solutions. In one recent equality-constrained model, the dynamics are built from the time-dependent saddle function
\[
\mathcal L_t(x,\lambda)=\mathcal L(x,\lambda)+\frac{c}{2t^p}\left(\|x\|^2-\|\lambda\|^2\right),
\]
which is strongly convex in the primal variable and strongly concave in the dual variable for each fixed \(t\). The associated unique saddle point \((x_t,\lambda_t)\) converges to the minimal norm saddle point of the original problem, and the actual trajectory is analyzed relative to \((x_t,\lambda_t)\) rather than directly relative to a fixed optimizer [2603.29124]. A separable analogue adds \(\epsilon(t)x\) and \(\epsilon(t)y\) to the primal equations and uses slow vanishing regularization to select \(\operatorname{proj}_S0\) [2408.06884].

Hessian-driven damping appears when the primal equation contains \(\frac{d}{dt}\nabla_x\mathcal L\) or \(\frac{d}{dt}\nabla_x\mathcal L_t\). Since
\[
\frac{d}{dt}\nabla_x\mathcal L_t(x(t),\lambda(t))
=\nabla^2 f(x(t))\dot x(t)+A^\top\dot\lambda(t)+\cdots,
\]
this produces curvature-adapted friction and explicit primal-dual velocity coupling. The variable-mass system of [2603.29124] and the closed-loop controlled system of [2602.16402] both use this mechanism, and both interpret it as a way to reduce oscillations.

Time scaling is another recurrent device. In mixed-order equality-constrained flows, a factor \(\beta(t)\) or \(t^s\) multiplies the primal and dual driving terms, so that weighted Lyapunov estimates convert bounded energies into explicit rates such as \(O(1/\beta(t))\) or better [2103.12931], [2408.06884], [2603.29124]. In closed-loop designs, the scaling is no longer prescribed externally: \(\mu(t)\) is defined by
\[
[\mu(t)]^p\|\nabla_x\mathcal L(x(t),\lambda(t))\|^{p-1}=1,
\]
and \(\tau(t)\) is generated from \(\mu(t)\), so damping, time scaling, and extrapolation all depend on the current state [2602.16402].

Preconditioning changes the metric of the flow rather than the objective itself. In the continuous preconditioned primal-dual inclusion, the off-diagonal velocity couplings \(\beta(t)A^*\dot y\) and \(\gamma(t)A\dot x\) alter the balance between dissipation and skew coupling. The antisymmetric and asymptotically antisymmetric regimes, characterized by \(\gamma=-\beta\) or \(\beta+\gamma\to0\), are singled out because they cancel the mixed kinetic term that obstructs non-ergodic decay in the unpreconditioned Arrow–Hurwicz flow [2506.00501].

Projected dynamics introduce geometric constraints at the vector-field level. In time-varying inequality-constrained optimization, dual feasibility is preserved by projecting the dual velocity onto the tangent cone of \(\mathbb R_{\ge0}^m\), while an additional augmentation term is introduced specifically to prevent multipliers from remaining stuck at zero when a constraint becomes active again [2312.10203]. In distributed optimization, adaptive synchronization laws can also be interpreted as closed-loop modifications of the interconnection weights, strengthening consensus couplings when primal disagreement is large [1905.00837].

## 4. Lyapunov analysis and convergence regimes

The dominant proof technique is Lyapunov analysis, but the asymptotic conclusions vary substantially with the geometry of the flow.

Projection-based Haugazeau dynamics are designed for strong convergence to a selected point. Under weak-cluster uniqueness, the trajectory converges strongly to \(\bar z=P_Z(\bar w)\), not merely to an arbitrary primal-dual solution, and explicit Euler discretization preserves the same best-approximation interpretation [2102.03807]. This is a markedly different selection principle from saddle-gradient or augmented-Lagrangian flows, where the limit point is usually any KKT pair unless extra regularization is added.

For structured convex minimization with full splitting, one obtains weaker asymptotics but explicit ergodic complexity. The first-order differential inclusion in \((x,z,y)\) converges weakly to a saddle point of the Lagrangian, the velocities vanish, and the ergodic trajectories satisfy \(O(1/t)\) bounds for feasibility violation and objective residual [1905.08290].

Mixed-order inertial systems exhibit several rate regimes. With constant viscous damping and time scaling, the second-order primal plus first-order dual flow yields
\[
\mathcal L(x(t),\lambda^*)-\mathcal L(x^*,\lambda(t))=O\!\left(\frac1{\beta(t)}\right),\qquad
\|Ax(t)-b\|=O\!\left(\beta(t)^{-1/2}\right),
\]
and when \(\beta(t)=\mu e^{t/\delta}\), both objective residual and feasibility violation decay like \(O(e^{-t/\delta})\), without strong convexity [2103.12931]. The accelerated APD flow based on the variables \((x,v,\lambda)\) goes further in the smooth case: its tailored Lyapunov function decays exponentially in continuous time, and the resulting discretizations inherit nonergodic bounds for primal-dual gap, feasibility violation, and objective residual [2109.12604].

Vanishing-damping second-order saddle systems recover the classical \(O(1/t^2)\) behavior of accelerated gradient dynamics in a constrained setting. For the augmented-Lagrangian flow with damping \(\alpha/t\), one has
\[
\mathcal L(x(t),\lambda^*)-\mathcal L(x^*,\lambda(t))=O\!\left(\frac1{t^2}\right),\qquad
\|Ax(t)-b\|=O\!\left(\frac1{t^2}\right),
\]
together with a two-sided \(O(1/t^2)\) objective estimate; if \(\nabla f\) is Lipschitz continuous, the full primal-dual trajectory converges weakly to a primal-dual optimal solution [2106.12294].

Tikhonov regularization changes the convergence mode. In the separable second-order-plus-first-order system, one obtains
\[
\mathcal L(x(t),y(t),\lambda^*)-\mathcal L(x^*,y^*,\lambda^*)=\mathcal O\!\left(\frac1{\beta(t)}\right),
\]
while objective error, feasibility residual, and gradient decay are \(\mathcal O(1/\sqrt{\beta(t)})\). Under \(\beta(t)\epsilon(t)\to+\infty\), the primal trajectory is driven toward the minimum norm solution, although the full strong-convergence statement requires an additional geometric condition on the tail of the trajectory [2408.06884]. The later variable-mass Hessian-damped system removes that extra geometric assumption and proves strong convergence of the full primal-dual trajectory to the minimal norm saddle point under explicit parameter inequalities, together with parameter-dependent asymptotic rates for primal-dual gap, objective residual, and feasibility violation [2603.29124].

Time-varying projected dynamics are formulated as tracking laws rather than optimizer-seeking flows. The asymptotic version is locally asymptotically stable with Lyapunov function \(V=\frac12\|\nabla_xL(x,\lambda,t)\|^2\), while the nonlinear correction
\[
-\frac{c_1\nabla_xL}{\|\nabla_xL\|^{\gamma_1}}-\frac{c_2\nabla_xL}{\|\nabla_xL\|^{\gamma_2}}
\]
yields fixed-time convergence of trajectories to the optimizer trajectory, with a uniform settling-time upper bound [2312.10203].

Preconditioned flows sharpen the distinction between ergodic and non-ergodic behavior. In the symmetric preconditioned case, the theory yields only ergodic gap estimates, effectively recovering time-rescaled Arrow–Hurwicz behavior. In contrast, exact or asymptotically antisymmetric preconditioners yield non-ergodic gap decay of the form
\[
\Delta_{u,v}(t)\le C\exp\!\left(-\int_0^t \frac{ds}{Z(s)}\right),
\]
and weak cluster points are primal-dual optimal [2506.00501]. This makes the choice of preconditioner a structural, not merely technical, issue.

Passivity-based convergence arguments appear in distributed networked settings. The adaptively synchronized distributed primal-dual dynamics are represented as a feedback interconnected passive system, and LaSalle’s invariance principle for hybrid systems is used to obtain asymptotic convergence and stability of the saddle-point solution; the same framework is then used to study accelerated convergence of the primal dynamics and \(L_2\)-gain robustness [1905.00837].

## 5. Discretization and algorithmic consequences

One of the main reasons primal-dual dynamical formulations are studied is that they often act as design templates for numerical algorithms, but the extent of this link differs widely across papers.

In some cases the continuous-time and discrete-time objects are almost interchangeable. The Haugazeau flow
\[
\dot x(t)=Q(\bar w,x(t),Tx(t))-x(t)
\]
has explicit discretization with step size \(1\) equal to the best-approximation algorithm
\[
x_{n+1}=Q(\bar w,x_n,Tx_n),
\]
so the dynamical system is literally the infinitesimal counterpart of an existing strongly convergent projection algorithm [2102.03807].

In structured convex minimization, explicit discretization of the primal-dual differential inclusion yields a scheme that the paper identifies as a combination of the linearized proximal method of multipliers and proximal ADMM; in a specific specialization, elimination of an auxiliary variable gives exactly the Chambolle–Pock primal-dual algorithm [1905.08290]. The continuous model is therefore not only interpretive but also algorithm-generating.

The APD framework was developed expressly as a dynamics-to-algorithm pipeline. Implicit, semi-implicit, and corrected explicit discretizations of the accelerated continuous model produce families of primal-dual methods analyzed through a discrete Lyapunov function mirroring the continuous one. These schemes give nonergodic rates for primal-dual gap, objective residual, and feasibility violation, and were further specialized to decentralized distributed optimization [2109.12604].

Closed-loop controlled dynamics also admit a direct discretization. For \(q=1\), time discretization of the controlled second-order primal plus first-order dual system yields the accelerated autonomous primal-dual algorithm AAPDA, whose step size is defined adaptively by
\[
\mu_k=\|\nabla_x\mathcal L(x_k,\lambda_k)\|^{-\frac{p-1}{p}},
\]
and whose analysis establishes discrete-time rates for the same three quantities as in the continuous model [2602.16402].

Other formulations remain only conceptually connected to algorithms. The variable-mass Hessian-damped Tikhonov system proves strong convergence and rate refinements in continuous time, but its conclusion explicitly presents explicit discretization into an inertial numerical algorithm as future work rather than as a developed result [2603.29124]. The preconditioned continuous flow is motivated by Chambolle–Pock as its discrete ancestor, yet its contribution is primarily asymptotic analysis of the continuous system rather than a new discretization theorem [2506.00501]. In infinite-dimensional optimal control, finite algorithmic time \(\tau\) is interpreted as producing sub-optimal controls, but the paper does not provide a complete discretized online control law [2606.18818].

## 6. Scope, applications, and conceptual boundaries

The range of applications is broad but still structurally coherent. Equality-constrained convex optimization is the dominant testbed, including nonseparable problems, separable block-structured models, decentralized distributed optimization, and linearly constrained quadratic programs [2103.12931], [2106.12294], [2408.06884], [2109.12604]. Monotone inclusions and best-approximation problems extend the framework beyond classical Lagrangian convex programs [2102.03807], [2506.00501]. Time-varying inequality-constrained convex optimization motivates projected prediction-correction flows capable of tracking moving KKT points and handling active/inactive switching of constraints [2312.10203]. Distributed multi-agent optimization introduces adaptive synchronization and passivity-based robustness analysis [1905.00837]. Convex optimal control pushes the framework into PDE form, with physical time \(t\) and algorithmic time \(\tau\) coexisting in an infinite-dimensional port-Hamiltonian system [2606.18818].

Several misconceptions recur in the literature. First, primal-dual dynamical formulation is not synonymous with classical Arrow–Hurwicz saddle flow. Projection-based Haugazeau dynamics, preconditioned monotone-inclusion flows, and port-Hamiltonian PDE formulations are all genuinely primal-dual but not of plain gradient descent-ascent type [2102.03807], [2506.00501], [2606.18818]. Second, strong convergence is not automatic; many systems yield only weak convergence or ergodic guarantees unless strengthened by Tikhonov regularization, Haugazeau geometry, fixed-time nonlinear correction, or special antisymmetric preconditioning [1905.08290], [2408.06884], [2312.10203], [2506.00501]. Third, dual projection for inequality constraints can create a “sticking at zero” pathology for multipliers, which is why augmentation terms such as \(G^\top\nabla_xL\) are introduced in time-varying projected dynamics [2312.10203]. Fourth, continuous-time acceleration does not by itself imply a useful discrete algorithm; some works derive full algorithmic counterparts, whereas others remain continuous-time analyses or conceptual templates [2109.12604], [2603.29124].

A final boundary is terminological. Several papers develop primal-dual formulations that are not dynamical systems at all. The proximal duality theory for non-convex variational optimization is explicitly a static variational formulation, not a flow [2103.01855]. The kernelizable primal-dual formulation of multilinear SVD is a KKT-based primal/dual optimization equivalence rather than a continuous-time model [2410.10504]. The primal-dual pricing method for short-circuit current in unit commitment is likewise a mixed-integer primal-dual optimization reformulation, not a primal-dual dynamics [2510.05293]. This suggests that “primal-dual dynamical formulation” should be reserved for constructions where time evolution itself is part of the mathematical object, whether as an ODE, differential inclusion, projected dynamical system, hybrid networked system, or PDE in algorithmic time.

Source: https://www.emergentmind.com/topics/primal-dual-dynamical-formulation