---
title: Batch-Constrained Actor-Critic Methods
url: https://www.emergentmind.com/topics/batch-constrained-actor-critic-architectures
type: topic
---

# Batch-Constrained Actor-Critic Methods

Batch-constrained actor-critic architectures constitute a class of reinforcement learning (RL) algorithms that structure both actor and critic updates around batches of consecutive Markovian transitions, rather than single-step or fully i.i.d. samples. This architectural constraint addresses bias and variance in gradient estimates that arise in average-reward Markov Decision Processes (MDPs), especially in the absence of standard ergodicity assumptions. The NAC-B (Natural Actor-Critic with Batching) algorithm exemplifies the batch-constrained approach, achieving order-optimal $\tilde{O}(\sqrt{T})$ regret in infinite-horizon average-reward unichain MDPs, a setting that permits both transient states and system periodicity [2505.19986].

## 1. Foundations of Batch-Constrained Actor-Critic Methods

Batch-constrained actor-critic methods operate by alternating between actor (policy) and critic (value) updates, each performed over finite batches of $B$ contiguous state-action-reward-next state tuples extracted from a fixed policy roll-out. In contrast to classical actor-critic methods relying on single-step temporal-difference or stochastic policy gradients, batch-constrained variants average gradient estimates over each batch, reducing both transient-state and periodicity-induced biases as well as stochastic variance in both subroutines. NAC-B structures the overall optimization into $K$ outer epochs, each comprising $H$ inner loops for both critic and actor, processing $2HB$ consecutive transitions for each policy parameterization.

## 2. Algorithmic Structure and Update Mechanisms

The NAC-B architecture is organized as follows:

- **Policy Fixation:** In each epoch $k$, the current policy $\pi_{\theta_k}$ is fixed.
- **Critic Batch Updates:** Over $H$ batches of size $B$, critic parameters $\xi_k = [\eta_k, \zeta_k]^T$ are updated using a batched temporal-difference objective, incorporating all $B$ contiguous transitions at each inner step. The update utilizes
  $$
  \xi_{h+1} = \xi_h - \frac{\beta}{B} \sum_{b=1}^B [A_v(\theta_k, z_b)\xi_h - b_v(\theta_k, z_b)],
  $$
  where $A_v$ and $b_v$ encode the policy and feature-dependent matrices and rewards.
- **Batched Advantage Estimation:** The critic produces advantage estimates $\hat{A}(z; \xi_k)$ for use by the actor.
- **Actor Batch Updates:** Using identical batching and $H$ additional sub-loops, the actor computes a natural gradient estimate $\omega_k$ via least-squares fitting over the same batch transitions:
  $$
  \omega_{h+1} = \omega_h + \frac{\gamma}{B} \sum_{b=1}^B [A_u(\theta_k, z_b)\omega_h - b_u(\theta_k, \xi_k, z_b)].
  $$
- **Policy Update:** Policy parameters are updated as $\theta_{k+1} = \theta_k + \alpha\omega_k$.
- **Data Collection:** All batches are contiguous on a single trajectory, avoiding simulator resets and leveraging the Markov property to gather relevant transition statistics.

## 3. Formal Analysis: Bias, Variance, and Convergence

Classical RL regret analyses often assume geometric mixing (ergodicity) of the Markov chain for tractable bounds on bias and variance in empirical averages. NAC-B relaxes this via the unichain assumption, embracing periodic and transient states. The algorithm introduces two critical constants—$C_{\text{hit}}$ and $C_{\text{tar}}$:

- $C_{\text{hit}}$ quantifies the worst-case expected time to reach the chain’s recurrent class from any initial state.
- $C_{\text{tar}}$ measures the expected return time to a recurrent-class state sampled from the stationary distribution.

For any batch of $B$ contiguous transitions, the deviation of empirical state occupancy from the stationary distribution is $O(C/B)$ (Lemma 4.1), and variance in empirical mean estimates is $O(1/B)$ (Lemma 4.2). By setting $B = \Theta(\sqrt{T})$, both bias and variance are driven to $O(1/\sqrt{T})$, supporting overall order-optimal $\tilde{O}(\sqrt{T})$ regret scaling.

## 4. Pseudocode and Implementation Details

The implementation is formalized in Algorithm 1:

```
Input: θ₀, ξ₀, ω₀, stepsizes α,β,γ, critic-scale cβ, outer epochs K, inner loops H, batch B
s ← s₀~ρ
for k=0…K−1 do
  # Critic subroutine
  ξ←ξ_k
  for h=0…H−1 do
    collect batch {z_b=(s_b,a_b,s_{b+1}), b=1…B} under π_{θ_k}
    ξ ← ξ − (β/B) Σ_b [A_v(θ_k,z_b) ξ − b_v(θ_k,z_b)]
    s ← s_{B+1}
  end for
  ξ_k ← ξ

  # Actor subroutine
  ω←ω_k
  for h=0…H−1 do
    collect batch {z_b}^B under π_{θ_k}
    ω ← ω + (γ/B) Σ_b [A_u(θ_k,z_b) ω − b_u(θ_k,ξ_k,z_b)]
    s← s_{B+1}
  end for
  ω_k ← ω

  # Policy update
  θ_{k+1} ← θ_k + α ω_k
end for
```

Total trajectory length is $T=2KHB$. All transitions are contiguous, exploiting the structure of Markovian data and ensuring no resets.

## 5. Advantages and Limitations in Unichain MDPs

Batching in NAC-B offers several core benefits:

- **Mitigation of Non-Ergodicity:** Markov transition kernels in unichain MDPs may not mix exponentially fast; batching ensures empirical distributions over $B$ steps converge to stationary with error $O(C/B)$.
- **Bias and Variance Control:** Simultaneous reduction of transient/periodic bias and sampling variance enables rigorous high-probability error bounds, forming the basis of the regret analysis.
- **Elimination of Simulator Resets:** The fully online, single-trajectory nature with batching avoids the need for repeated stochastic restarts, characteristic of prior approaches under stronger assumptions.

A plausible implication is that batch-constrained architectures can generalize beyond the unichain assumption if analogous bounds on empirical stationary convergence can be established for broader MDP classes.

## 6. Theoretical Significance and Connections

NAC-B demonstrates that batch-constrained actor-critic methods can achieve non-asymptotic $\tilde{O}(\sqrt{T})$ regret in average-reward RL without strict mixing conditions. This approach “batches out” Markov noise similarly to variance reduction techniques in stochastic optimization, but is tailored for dependencies inherent in sequential data. The unichain assumption stands as one of the weakest under which the policy gradient theorem for average-reward remains valid. The batch-constrained framework thereby extends the applicability of scalable RL algorithms to environments where ergodicity is not assured, aligning theoretically justified updates with algorithmic practicality [2505.19986].

Source: https://www.emergentmind.com/topics/batch-constrained-actor-critic-architectures