---
title: Separately Controlled Chains
url: https://www.emergentmind.com/topics/separately-controlled-chains
type: topic
---

# Separately Controlled Chains

Searching arXiv for papers on “separately controlled chains” and closely related controlled Markov chain formulations.
Separately controlled chains are stochastic control or game-theoretic models in which the global system comprises multiple local Markov chains, each driven exclusively by its own local control, while coupling is introduced through an aggregate objective, information structure, or performance criterion rather than through the transition law itself. In the recent stochastic-games formulation, the defining structural property is a product-form transition kernel,
\[
P\bigl(\bm s_{t+1}\mid \bm s_t,\bm a_t\bigr)=\prod_{i=1}^n P^i\bigl(s'_{i}\mid s_i,a_i\bigr),
\]
together with local rewards of the form \(r_i(\bm s_t,\bm a_t)=r_i(s_{i,t},a_{i,t})\), so that each player’s internal state is controlled only by its own action [2507.22335]. Closely related controlled-Markov-chain work studies joint processes \((\theta_n,X_n)\) where the state chain and parameter chain admit separate Lyapunov drifts that can be combined into a single recurrence argument [1205.4181]. A different but conceptually adjacent line in partially observed control uses the term “separated” to denote the transformation of a partially observed controlled Markov chain into a fully observed control problem on the filtering state, governed by the Wonham–Zakai equations [2602.16392]. Taken together, these uses locate separately controlled chains within a broader family of decomposed stochastic systems in which dynamic decoupling coexists with objective- or inference-level coupling.

## 1. Formal model and defining structure

In the \(n\)-player stochastic-game formulation, player \(i\) has state space \(\mathcal S_i\), action space \(\mathcal A_i\), local transition kernel
\[
P^i(s_i' \mid s_i,a_i)=\Pr\{s_{i,t+1}=s_i' \mid s_{i,t}=s_i,a_{i,t}=a_i\},
\]
and immediate reward \(r_i(s_i,a_i)\) [2507.22335]. The global state is \(\bm s=(s_1,\dots,s_n)\), the joint action is \(\bm a=(a_1,\dots,a_n)\), and the key assumption is that the controlled dynamics factorize across players:
\[
P\bigl(\bm s_{t+1}\mid \bm s_t,\bm a_t\bigr)=\prod_{i=1}^n P^i\bigl(s'_{i}\mid s_i,a_i\bigr),
\qquad
r_i(\bm s_t,\bm a_t)=r_i(s_{i,t},a_{i,t}).
\]
This is the sense in which the chains are “controlled separately”: the evolution of player \(i\)’s local chain is unaffected by the actions of the other players at the transition level [2507.22335].

The same paper imposes a local-information architecture. Each player observes only its own history
\[
h_{i,t}=(s_{i,0},a_{i,0},r_{i,0},\dots,s_{i,t}),
\]
and uses a stationary deterministic policy \(u_i:\mathcal S_i\to\mathcal A_i\); the joint policy is \(u=(u_1,\dots,u_n)\) [2507.22335]. Thus, dynamic independence is paired with decentralized observability.

A related controlled-Markov-chain framework considers a family of kernels \(\{P_\theta:\theta\in\Theta\}\) on a state space \(X\), together with parameter updates \(\theta_{n+1}=\phi_{\gamma_{n+1}}(\theta_n,X_{n+1})\), yielding
\[
X_{n+1}\sim P_{\theta_n}(X_n,\cdot),\qquad
\theta_{n+1}=\phi_{\gamma_{n+1}}(\theta_n,X_{n+1}).
\]
Although this is not an \(n\)-player product-form model, it is a separately controlled-chain framework in the sense that the state and parameter components admit distinct control- and drift-analytic treatment before being recombined through a joint Lyapunov argument [1205.4181].

This suggests that “separately controlled chains” is best understood not as a single formalism but as a structural principle: one isolates subsystems whose one-step laws are locally controlled, and then studies how global behavior emerges once a common criterion or coupled analysis is imposed.

## 2. Objective coupling despite transition decoupling

The defining subtlety of separately controlled chains is that dynamic decoupling does not imply optimization decoupling. In the stochastic-games setting, all players share a common objective called the team variance [2507.22335]. The long-run average team reward is
\[
\mu^u
=
\frac1n\sum_{i=1}^n
\lim_{T\to\infty}\frac1T\sum_{t=0}^{T-1}
\mathbb E^u[r_i(s_{i,t},a_{i,t})],
\]
and the team variance is
\[
J^u_\sigma
=
\mathbb E^u\Bigl[
\sum_{i=1}^n
\lim_{T\to\infty}\frac1T\sum_{t=0}^{T-1}
\bigl(r_i(s_{i,t},a_{i,t})-\mu^u\bigr)^2
\Bigr].
\]
Although each reward term depends only on a local state-action pair, the centering quantity \(\mu^u\) depends on the entire joint policy, so the objective couples the players through a global statistic [2507.22335].

The paper explicitly identifies the consequence: the variance metric is not additive or Markovian, and the dynamic programming principle fails [2507.22335]. The instantaneous quantity \((r_i(s_{i,t},a_{i,t})-\mu^u)^2\) is not a standard stage cost because \(\mu^u\) is itself a functional of the policy. This breaks the Bellman decomposition ordinarily available for average-cost Markov decision processes.

An equivalent decomposition clarifies the structure:
\[
J^u_\sigma(y)
=
\sum_{i=1}^n
\mathbb E^u\Bigl[\lim_{T\to\infty}\tfrac1T\sum
(r_i(s_{i,t},a_{i,t})-y)^2\Bigr],
\]
where \(y\in\mathbb R\) is a pseudo team mean, and
\[
J^u_\sigma=\min_{y\in\mathbb R}J^u_\sigma(y),\qquad
J^u_\sigma(y)=\sum_{i=1}^n J^{u_i}_{\sigma,i}(y).
\]
For fixed \(y\), the problem decomposes into \(n\) independent average-cost MDPs, one per player; the coupling reappears only through the outer minimization over \(y\) [2507.22335].

A plausible implication is that separately controlled chains are especially natural when the system designer seeks fairness, risk balancing, or coordinated variability reduction rather than purely additive throughput or reward. In such cases, independence of transitions simplifies modeling, but the non-additive criterion remains the main analytical obstacle.

## 3. Sensitivity analysis and bilevel reformulation

Because dynamic programming fails for team variance, the recent theory develops a sensitivity-based formulation [2507.22335]. Introducing the pseudo team mean \(y\), one rewrites the global optimization problem as
\[
\min_{y\in\mathbb R}\;\sum_{i=1}^n \min_{u_i}\;J^{u_i}_{\sigma,i}(y).
\]
This yields a bilevel structure: an outer optimization over the coordinating scalar \(y\), and an inner level consisting of independent ergodic MDPs [2507.22335].

For a fixed \(y\) and player \(i\), the analysis uses the steady-state distribution \(\pi_i'\) under a comparison policy \(u_i'\), transition matrices \(P_i,P_i'\), immediate-cost vectors \(r_i,r_i'\) with entries \((r_i(s,a)-y)^2\), and a relative-value function \(g_i\) solving
\[
g_i=(r_i-y\mathbf1)^2-J^{u_i}_{\sigma,i}(y)\mathbf1+P_i g_i.
\]
The performance-difference formula is then
\[
J^{u_i'}_{\sigma,i}(y)-J^{u_i}_{\sigma,i}(y)
=
\pi_i'\bigl[(P_i'-P_i)g_i+r_i'-r_i\bigr].
\]
Aggregating across players and setting \(y=\mu^u\) produces the team-variance difference formula
\[
J^{u'}_\sigma - J^u_\sigma
=
\sum_{i=1}^n
\pi_i'\Bigl[(P_i'-P_i)g_i
+ (r_i'-\mu^u\mathbf1)^2 - (r_i-\mu^u\mathbf1)^2\Bigr]
-
n(\mu^{u'}-\mu^u)^2.
\]
A derivative formula is obtained by mixing policies through \(u^\delta=(1-\delta)u+\delta u'\):
\[
\frac{\partial}{\partial\delta}J^{u^\delta}_\sigma\Big|_{\delta=0}
=
\sum_{i=1}^n
\pi_i\Bigl[(P_i'-P_i)g_i
+ (r_i'-\mu^u\mathbf1)^2 - (r_i-\mu^u\mathbf1)^2\Bigr].
\]
These formulas supply the sensitivity information needed for decentralized improvement steps [2507.22335].

The algorithmic consequence is a decentralized bilevel optimization method. At iteration \(\ell\), each player evaluates the current team mean \(\mu^{(\ell)}=\mu^{u^{(\ell)}}\), solves the Poisson equation for the bias vector \(g_i^{(\ell)}\) of its own pseudo-MDP with running cost \((r_i-\mu^{(\ell)})^2\), and then updates its policy independently by
\[
u_i^{(\ell+1)}(s)
=
\arg\min_{a\in\mathcal A_i}
\Bigl\{(r_i(s,a)-\mu^{(\ell)})^2
+\sum_{s'}P^i(s'|s,a)\,g_i^{(\ell)}(s')\Bigr\}.
\]
Ties are broken by staying with the old action if possible, to avoid cycles [2507.22335].

The paper proves that this algorithm strictly decreases \(J^u_\sigma\) at each update, terminates in finitely many steps because the deterministic-policy space is finite and bounded below, and converges to a limit point satisfying first-order necessary conditions in the mixed-policy space; in most cases the limit is a strictly local minimum [2507.22335].

## 4. Equilibrium structure and recurrence analysis

For the team-variance game, the existence result is stated as follows: under the assumptions of separately controlled chains and local observations, there exists a joint policy \(u^*\) with each \(u_i^*\) stationary deterministic that minimizes \(J^u_\sigma\); equivalently, \(u^*\) is a stationary pure Nash equilibrium of the \(n\)-player game [2507.22335]. The proof strategy uses the fact that for fixed \(y\), each inner problem is a standard ergodic MDP with average-cost \((r_i-y)^2\), so each admits an optimal stationary-deterministic policy; convexity in \(y\) then yields a global minimizer [2507.22335].

A distinct but relevant stability perspective appears in the controlled-Markov-chain recurrence theory of Andrieu, Tadić, and Vihola [1205.4181]. There, the process \((\theta_n,X_n)\) is governed by a kernel update for \(X_n\) and a stepsize-indexed map for \(\theta_n\). The central idea is to derive separate Lyapunov drifts for the state and parameter components and combine them into a single Foster–Lyapunov inequality for the joint chain.

The state-chain drift is
\[
E[V(X_{n+1})\mid X_n=x,\theta_n=\theta]
\le V(x)-\Delta_V(\theta,x),
\]
with
\[
\Delta_V(\theta,x)=a^{-1}(\theta)V(x)^\iota 1_{x\notin C}-b(\theta)1_{x\in C},
\]
while the parameter-chain drift is
\[
E[w(\theta_{n+1})\mid X_n=x,\theta_n=\theta]
\le w(\theta)-\gamma_{n+1}\Delta_w(\theta,x),
\]
with
\[
\Delta_w(\theta,x)=w(\theta)\Delta\bigl(c(\theta)+V(x)^\beta/e(\theta)\bigr)1_{x\notin C}
+w(\theta)\Delta(d(\theta))1_{x\in C}.
\]
Defining
\[
L_n(\theta,x)=\lambda V(x)+\frac{w(\theta)}{\gamma_n},
\]
the paper proves that for suitable \(\lambda>0\) there exists \(\delta>0\) such that for all large \(n\) and all \((\theta,x)\notin W\times C\),
\[
E[L_{n+1}(\theta_{n+1},X_{n+1})\mid \theta_n=\theta,X_n=x]
\le
L_n(\theta,x)-\delta\Bigl[\frac{V(x)^\iota}{a(\theta)}+w(\theta)\Bigr].
\]
From this compound drift, one deduces that \((\theta_n,X_n)\) returns infinitely often to \(W\times C\) almost surely [1205.4181].

This recurrence theory does not use the same product-form transition structure as the stochastic-game model, but it exemplifies an allied methodological theme: separate control or drift structure at the component level can be recombined into a global stability statement. For separately controlled chains more broadly, this suggests that equilibrium analysis and recurrence analysis are complementary rather than competing perspectives.

## 5. Time-scale separation and partially observed separation

Time-scale separation is explicit in the recurrence analysis of controlled Markov chains [1205.4181]. When \(\gamma_n\to 0\), the parameter chain moves slowly relative to the state chain; the proof compares the slow \(w\)-drift, of order \(\gamma_n\), with the fast \(V\)-drift, of order \(1\), by rescaling \(w\) by \(1/\gamma_n\). The condition
\[
\limsup_n(\gamma_{n+1}^{-1}-\gamma_n^{-1})<\Delta(0)
\]
ensures that the rescaling does not grow too quickly and that the joint process remains stable even when the components evolve on different time scales [1205.4181].

Another notion of separation arises in partially observed controlled Markov chains [2602.16392]. Here the unobserved finite-state chain \(X_t^\alpha\) has controlled transition rates, while the controller observes only a diffusion-type process
\[
W_t=\int_0^t h(X_s^\alpha,\alpha_s,s)\,ds+B_t.
\]
Through a Girsanov-type change of measure, one introduces the unnormalized filter
\[
\rho_t^i=\mathbb E\bigl[1_{\{X_t^\alpha=i\}}Z_t^\alpha\mid \mathcal F_t^W\bigr],
\]
which solves the controlled Wonham–Zakai SDE
\[
d\rho_t^i
=
\sum_{j=1}^N \rho_t^j q(\alpha_t,t,j,i)\,dt
+\rho_t^i h(i,\alpha_t,t)\,dW_t.
\]
The resulting “separated optimal control problem” is fully observed in the filter state \(\rho_t\), and its value function satisfies HJB equations in viscosity form [2602.16392].

This use of “separated” is terminologically distinct from “separately controlled chains,” yet the connection is instructive. In both settings, control theory exploits a decomposition that converts an analytically difficult coupled problem into a tractable one on transformed coordinates or reduced subsystems. In the partially observed case the decomposition is epistemic, from hidden state to conditional law; in separately controlled chains it is dynamic, from joint transitions to local kernels.

A plausible implication is that future work may combine these two notions, for example by studying decentralized team objectives under local partial observations, with each player controlling a local hidden Markov chain and optimization performed on local or joint filtering states. The cited papers do not state such a model, but their technical ingredients are compatible in spirit.

## 6. Applications, limitations, and related interpretations

The most explicit application reported for separately controlled chains is decentralized energy management in a smart grid [2507.22335]. In the numerical experiment, \(n=3\) microgrids are modeled with wind-power states \(G_{i,t}\in\{0,1,\dots,5\}\), battery-level states \(B_{i,t}\in\{0,1,\dots,5\}\), actions \(a_{i,t}\in\{-2,-1,0,1,2\}\), and net exchanges
\[
r_{i,t}=G_{i,t}+a_{i,t}-D_i,
\]
with constant demand \(D_i\). The separately controlled-chains assumption holds because
\[
P^i(s_i'|s_i,\bm a)=P^i(s_i'|s_i,a_i).
\]
Applying the decentralized algorithm from a random initialization, the reported team variance \(J_\sigma\) drops from about \(10.12\) to \(4.34\) in \(6\) iterations; the team mean fluctuates mildly; and each microgrid’s pseudo variance \(\mathbb E[(r_i-\mu)^2]\) is monotonically reduced, while raw variance may transiently rise [2507.22335].

In adaptive MCMC, the recurrence framework applies to algorithms such as Adaptive Metropolis and coerced-acceptance-probability random-walk Metropolis [1205.4181]. For Adaptive Metropolis,
\[
X_{n+1}\sim \mathrm{Metropolis}(\pi,N(\mu_n,\Sigma_n)),
\]
with updates
\[
\mu_{n+1}=\mu_n+\gamma_{n+1}(X_{n+1}-\mu_n),
\]
\[
\Sigma_{n+1}=\Sigma_n+\gamma_{n+1}\bigl[(X_{n+1}-\mu_n)(X_{n+1}-\mu_n)^T-\Sigma_n\bigr].
\]
Under suitable drift and moment conditions, the compound-drift theorem yields recurrence of \((\mu_n,\Sigma_n,X_n)\) to a compact \(W\times C\) [1205.4181]. Although this is not a decentralized game, it shows that separately analyzable controlled components arise naturally in computational statistics.

Several common misconceptions can be ruled out. First, separately controlled chains do not mean that the optimization problem is decentralized in the strong sense of being fully independent across subsystems. In the team-variance model, the objective couples all players through \(\mu^u\), and the authors emphasize that dynamic programming fails precisely because of this non-Markovian coupling [2507.22335]. Second, separate control does not by itself guarantee tractable verification of optimality or global convergence; the cited algorithm converges to a first-order stationary point in the mixed-policy space and, in most cases, a strictly local minimum, not necessarily a global optimum over all mixed policies [2507.22335]. Third, the term should not be conflated with unrelated uses of “chain” in robotics or blockchain systems. For example, “control chains” in hierarchical robot control refer to ordered tuples of low-level controllers coordinated for sequential manipulation [2205.04362], while “decoupled consensus between chains” concerns blockchain sidechains with independent consensus mechanisms [1812.05441]. These are distinct domains and definitions.

In present research usage, separately controlled chains therefore denote a class of structured stochastic systems where local transition dynamics are autonomous under local control, yet global analysis remains nontrivial because coupling re-enters through risk-sensitive criteria, equilibrium conditions, shared statistics, or stability arguments. The concept is valuable precisely because it isolates where complexity resides: not in the product-form dynamics themselves, but in the global objective or in the analytical machinery required to recombine the local pieces into a coherent control theory [2507.22335][1205.4181][2602.16392].

Source: https://www.emergentmind.com/topics/separately-controlled-chains