---
title: Zero-Sum Multichain Lagrangian Formulation
url: https://www.emergentmind.com/topics/zero-sum-multichain-lagrangian-formulation
type: topic
---

# Zero-Sum Multichain Lagrangian Formulation

The Feasible Actor-Critic (FAC) algorithm is a model-free framework for constrained reinforcement learning (RL) that enforces statewise safety. Unlike conventional constrained RL methods, which apply constraints only in expectation over initial states, FAC imposes safety for every feasible initial state individually. FAC achieves this by leveraging a statewise Lagrange multiplier, parameterized by a neural network, to dynamically adapt constraint enforcement at the state level and distinguish inherently infeasible states from those for which safe policies exist [2105.10682].

## 1. Statewise Feasibility and Constraint Formalism

Let $\mathcal S$ denote the state space and $\pi$ a stochastic policy. The cost-value function (safety critic) from any state $s$ under policy $\pi$ is
\[
v_{C}^{\pi}(s)=\E_{\tau\sim\pi}\left[\sum_{t=0}^\infty\gamma_c^t\,c(s_t,a_t)\;\middle|\;s_0=s\right]
\]
for cost function $c$, with discount $\gamma_c$. Fixing a scalar threshold $d$, a state $s$ is called *feasible* under $\pi$ if $v_C^\pi(s) \leq d$; otherwise, it is *unsafe*. Certain states, denoted $\mathcal S_I$, are inherently infeasible if $v_C^\pi(s) > d$ for all policies $\pi$; the complement, $\mathcal S_F = \mathcal S \setminus \mathcal S_I$, constitutes the feasible region. For a subset of initial support $\mathcal I \subseteq \mathcal S$, the potentially feasible initial states are $\mathcal I_F = \mathcal I \cap \mathcal S_F$. The *statewise safety constraint* requires $v_C^\pi(s) \leq d$ for every $s \in \mathcal I_F$.

## 2. Optimization Problem and Statewise Lagrangian

The learning objective is to maximize the expected discounted return
\[
J(\pi) = \E_{\tau\sim\pi}\left[\sum_{t=0}^{\infty}\gamma^t r(s_t, a_t)\right]
\]
subject to the *infinite family of constraints*
\[
v_C^\pi(s) \leq d, \quad \forall s \in \mathcal I_F.
\tag{SP}
\]
FAC introduces a nonnegative, state-dependent Lagrange multiplier $\lambda: \mathcal S \to \mathbb{R}_+$, and forms the *statewise Lagrangian*
\[
L_{\mathrm{stw}}(\pi, \lambda) = \E_{s\sim d_0} \left[ -v^\pi(s) + \lambda(s)(v_C^\pi(s) - d) \right],
\tag{SL}
\]
where $d_0$ is the initial-state distribution. The corresponding saddle-point problem,
\[
\max_{\lambda(\cdot)\ge 0} \min_\pi L_{\mathrm{stw}}(\pi, \lambda),
\]
ensures, under mild conditions, that the optimal policy $\pi^*$ solves (SP).

## 3. Algorithmic Architecture and Updates

FAC extends standard off-policy actor-critic frameworks, particularly soft-actor-critic (SAC), by adding a multiplier network. The key components are:

- Two *reward Q-networks*: $Q_{\theta_1}(s,a), Q_{\theta_2}(s,a)$ (double-Q)
- *Cost Q-network*: $Q_{C,\theta_C}(s,a)$ (predicts cost-to-go)
- *Stochastic policy network*: $\pi_\phi(a|s)$ (e.g., Gaussian with $\tanh$ squashing)
- *Multiplier network*: $\lambda_\xi(s)\ge 0$

Each is updated as follows:

**Reward Q-update:**
\[
J_Q(\theta_i) = \E_{(s,a,r,s')\sim \mathcal B} \left[ \tfrac12 (Q_{\theta_i}(s,a) - y(s,a))^2 \right]
\]
where $y(s,a) = r + \gamma \E_{a'\sim\pi_\phi}\left[\min_j Q_{\bar{\theta}_j}(s', a') - \alpha \log \pi_\phi(a'|s')\right]$.

**Cost Q-update:** Mirrors reward Q, using cost $c$, discount $\gamma_c$, and no entropy penalty.

**Policy update:**
\[
J_\pi(\phi) = \E_{s\sim\mathcal B}\E_{a\sim\pi_\phi}\left[ \alpha \log \pi_\phi(a|s) - Q(s,a) + \lambda_\xi(s)(Q_C(s,a)-d) \right]
\]
with gradient estimated via reparameterization ($a = f_\phi(\epsilon; s)$).

**Multiplier update:**
\[
J_\lambda(\xi) = \E_{s\sim\mathcal B, a\sim\pi_\phi}\left[\lambda_\xi(s)(Q_C(s,a)-d)\right]
\]
with dual ascent on $\xi$.

**Entropy temperature $\alpha$:** Adapted as in SAC by gradient update.

A pseudocode summary is provided:

```
Initialize θ₁, θ₂, θ_C   ← Q-net weights
          ϕ              ← policy net weights
          ξ              ← multiplier net weights
          α              ← entropy temp
          τ              ← target-net smoothing
          B              ← empty replay buffer
for iteration = 1…M do
  for each environment step do
    a∼π_ϕ(·|s), observe (s,a,r,c,s′)
    store (s,a,r,c,s′) in B
  end for
  for gradient step = 1…N do
    Sample minibatch from B
    // 1) Update reward Q's
    θᵢ ← θᵢ - β_Q ∇_{θᵢ} ½[E(Q_{θᵢ}(s,a)-y)²],  i=1,2
    // 2) Update cost Q
    θ_C ← θ_C - β_Q ∇_{θ_C} ½[E(Q_C(s,a)-y_C)²]
    if step % m_π == 0 then
      // 3) Update policy ϕ
      ϕ ← ϕ - β_π ∇_ϕ E[ α logπ_ϕ(a|s) - Q(s,a) + λ_ξ(s)(Q_C(s,a)-d) ]
      // 4) Adjust α
      α ← α - β_α ∇_α E[ -α logπ_ϕ(a|s) - αŀ ]
    end if
    if step % m_λ == 0 then
      // 5) Update multiplier ξ
      ξ ← ξ + β_λ ∇_ξ E[ λ_ξ(s)(Q_C(s,a)-d) ]
    end if
    // 6) Soft update target networks
    θ̄ᵢ ← τ θᵢ + (1-τ)θ̄ᵢ ,    θ̄_C ← τ θ_C + (1-τ)θ̄_C
  end for
end for
```

## 4. Complementary Slackness and Feasibility Detection

At the optimum $(\pi^*, \lambda^*)$, the statewise complementary slackness conditions apply:
\[
\lambda^*(s) \ge 0,\quad v_C^{\pi^*}(s) \le d,\quad \lambda^*(s)(v_C^{\pi^*}(s)-d) = 0
\]
Thus, for $v_C(s) < d$ (strict safety), $\lambda^*(s) = 0$; for $v_C(s) = d$, $\lambda^*(s) > 0$. If $s$ is inherently infeasible, dual ascent causes $\lambda(s) \to +\infty$, so large values of $\lambda_\xi(s)$ can be used as a practical indicator of infeasibility. The trained multiplier network $\lambda_\xi$ thus serves as a learned feasibility oracle over $\mathcal S$.

## 5. Theoretical Properties

FAC's theoretical guarantees are as follows:

- *Equivalence of scaled/unscaled Lagrangians* (Theorem 3.1): Optimizing $\max_\lambda \min_\pi \E_{s\sim d_0}[\cdots]$ recovers the same $\pi^*$ as the infinite-constraint problem (SP).
- *Statewise implies expectation-based feasibility* (Theorem 3.2): Any policy $\pi$ satisfying $v_C^\pi(s) \le d$ for all $s$ in $\mathcal I$ also satisfies $\E_{s\sim d_0}[v_C^\pi(s)] \le d$.
- *Performance comparison* (Theorem 3.3): For the optimal statewise Lagrangian $\mathcal L^*_{\mathrm{stw}}$ and expectation Lagrangian $\mathcal L^*_{\mathrm{exp}}$, $\mathcal L^*_{\mathrm{stw}} \le \mathcal L^*_{\mathrm{exp}}$ implies $J^*_{\mathrm{stw}} \ge J^*_{\mathrm{exp}}$, i.e., FAC can only improve, or at least match, reward under tighter constraints compared to expectation-based methods.

## 6. Empirical Evaluation

FAC was assessed on both robot locomotion tasks with speed constraints (HalfCheetah, Walker2d, Ant) and safe exploration scenarios (Safety-Gym: Point-Button, Car-Goal). Key findings are:

- In robot locomotion tasks, FAC consistently respects the speed cap and achieves higher or comparable returns versus expectation-based baselines such as CPO, TRPO-Lagrangian, and PPO-Lagrangian, which exhibit greater constraint violations and instability.
- On Safety-Gym safe exploration, FAC yields an episode-cost rate well below the 10% threshold even at the tails of 95% confidence intervals and fewer dangerous episodes at test time, while scalar-multiplier baselines suffer approximately 50% dangerous runs.
- Feasibility indication emerges: $\lambda_\xi(s)$ sharply increases when the agent enters inherently infeasible regions (e.g., boxed-in by obstacles) and moderates when the agent returns to feasible zones.

## 7. Implementation Considerations and Practical Impact

FAC can be implemented in any off-policy actor-critic codebase by adding the multiplier network $\lambda_\xi(s)$, modifying the policy loss with the $\lambda(s)(Q_C(s,a)-d)$ term, and updating $\xi$ by dual ascent. This yields a policy that is statewise safe on all feasible initial states and a by-product classifier of infeasible ones. The learned multiplier network presents a direct means for practitioners to identify – in real time – whether specific initializations or encountered states admit any safe policy or are intrinsically unsafe [2105.10682].

Source: https://www.emergentmind.com/topics/zero-sum-multichain-lagrangian-formulation