---
title: Reward Decomposition & Exogenous Filtering
url: https://www.emergentmind.com/topics/reward-decomposition-and-exogenous-filtering
type: topic
---

# Reward Decomposition & Exogenous Filtering

Reward decomposition and exogenous filtering are techniques in reinforcement learning (RL) and bandit problems that explicitly split observed variation into reward-relevant (endogenous) and reward-irrelevant (exogenous) components of the state, transition, or reward process. This structural insight enables algorithms to filter out exogenous noise, yielding sample complexity and regret guarantees scaling only in the (typically much lower-dimensional) endogenous subspace. Such decompositions are realized via reward-relevant support recovery, subspace discovery, and causal/statistical independence constraints, and apply across offline and online RL, as well as in specialized bandit problems. The decision-theoretic significance is that the optimal policy always depends only on the endogenous part, even when full-state transitions depend on exogenous factors.

## 1. Formal Decomposition: Exogenous and Endogenous State and Reward

In a general Markov Decision Process (MDP) or bandit model, the observed state $x$ can often be partitioned as $x = (s, e)$, with $s$ "endogenous" (reward- or action-relevant) and $e$ "exogenous" (reward-irrelevant) [2401.12934, 2303.12957, 1806.01584]. Exogenous variables $e$ evolve stochastically and independently of agent action, while $s$ can depend on both $x$, the action $a$, and possibly the exogenous variable. 

A decomposition is formally characterized by:
- **Transition factorization**: $P(x' \mid x, a) = P(s' \mid s, a) \cdot P(e' \mid s, e, a)$ with $e' \perp s' \mid (s, a)$ and $e \perp r \mid s$ [2401.12934].
- **Additive reward decomposition**: $R(x, a) = R_x(e) + R_s(s, a)$ or more generally $R(s, e, a) = R_{\rm exo}(e) + R_{\rm end}(s,e,a)$, with $R_{\rm exo}$ actions independent [2303.12957, 1806.01584].
- **Causal exogeneity**: $e$ satisfies $P(e_{t+1} \mid s_t, e_t, a_t) = P(e_{t+1} \mid e_t)$ (actions do not influence $e$'s future) [2303.12957].

**Key Theorem:** The value function and Bellman equations decompose as
$$
V(s, e; h) = V_{\rm exo}(e; h) + V_{\rm end}(s, e; h)
$$
where $V_{\rm exo}(e; h)$ is the (action-independent) value induced by exogenous reward and $V_{\rm end}$ is the solution to the endogenous MDP. Thus, the optimal policy depends only on the endogenous component [1806.01584, 2303.12957].

## 2. Causal and Statistical Criteria for Exogeneity

Causal and statistical independence conditions are central to guarantees of exogenous filtering. In the formalism of [2303.12957], exogeneity holds if, for $X$ an exogenous subvector,
$$
P(X_{t+1}, ..., X_H \mid X_t, \operatorname{do}(A_t = a_t)) = P(X_{t+1}, ..., X_H \mid X_t)
$$
or equivalently, in two-time-slice DBN terms,
$$
P(E', X' \mid E, X, A) = P(X' \mid X) P(E' \mid E, X, A, X')
$$
Statistical exogeneity is identified through a factorization and conditional independence constraints, typically via zero (conditional) correlation coefficients or conditional mutual information minimization with respect to candidate exogenous spaces [2303.12957, 1806.01584].

Blockwise and "diachronic" independence assumptions—such as $e_t \perp s_{t+1}\mid(s_t, a_t)$ and $s_t \perp e_t \mid (s_{t-1}, a_{t-1})$—formalize reward-irrelevant pathways [2401.12934].

## 3. Algorithmic Approaches: Support Recovery and Subspace Discovery

### Linear and High-dimensional Decompositions

When the state is high-dimensional, reward decomposition identifies a sparse reward-relevant subspace via support recovery, typically using thresholded Lasso or coefficient thresholding:
1. Lasso or sparse regression of the reward on features $\phi(x,a)$.
2. Identify nonzero coefficients exceeding threshold $\tau$; these index the reward-relevant “support.”
3. Restrict further value function estimation to this support for efficient policy evaluation [2401.12934].

Offline RL with function approximation then regresses the Bellman target only on endogenous features, iterating via Fitted Q-Iteration on the learned support. Sample complexity and policy suboptimality then depend only on $|\text{support}|=|s|$, not the ambient $d$ [2401.12934].

### Latent Exogeneity Discovery

When the exogenous/endogenous split is unknown, recovery proceeds via linear projections or matrix decompositions:
- Seek $W_{\rm exo}\in\mathbb{R}^{d\times d_{\rm exo}}$ maximizing the variance of $x=W_{\rm exo}^\top s$ subject to $x$ being conditionally independent of $(e, a)$ given $x$.
- Use conditional correlation coefficient (CCC) or partial correlation constraints to make the conditional independence tractable [2303.12957, 1806.01584].

Two principal algorithms exist:
- Global rank-descending: Solve for the largest $d_{\rm exo}$ for which $\operatorname{Covc}(\cdot)\le \epsilon$.
- Stepwise rank-ascending: Build up $W_{\rm exo}$ one vector at a time, greedily extending while maintaining the conditional independence property.

Once $W_{\rm exo}$ is found, endogenous rewards are isolated via regression, and exogenous reward is subtracted for all subsequent learning [2303.12957].

## 4. Sample Efficiency and Regret Guarantees

Exogenous filtering dramatically improves sample complexity and regret bounds by reducing effective statistical dimension:

- In Exo-MDPs (finite-state, stochastic exogenous dynamics), both rewards and transitions decompose linearly with respect to the unknown law $\Xi\in\Delta^{d}$. RL reduces to estimating only $\Xi$ rather than full transition and reward tables. This yields regret upper bounds of $O(H^{3/2} d \sqrt{K})$ for horizon $H$ and $K$ episodes (no observation), and $O(H^{3/2}\sqrt{dK})$ in the full-observation regime, where $d$ is the exogenous state cardinality [2409.14557].

- In offline RL with linear transitions, statistical complexity and policy-value suboptimality scale with $|s|$, the dimension of the reward-relevant (endorogenous) subspace, and not the full ambient dimension $d$ [2401.12934]. Specifically, the suboptimality bound is $O(T\sqrt{|s|(\log d)/n})$ for $n$ trajectories.

- Similar reductions apply in bandit settings; in the RMAB under exogenous Markov process context, regret is shown to be $O(\log T)$, as the exploration focuses only on the relevant local means and global transitions, with exogenous variation properly filtered out [2112.09484].

## 5. Algorithms in Practice: Applications and Empirical Findings

Reward decomposition and exogenous filtering have been applied in a variety of RL and bandit scenarios:

- **High-dimensional synthetic MDPs:** Filtering out the exogenous component reduces mean-squared error and accelerates policy learning significantly, with empirical results showing order-of-magnitude improvements in sample efficiency [2401.12934, 2303.12957].
- **Inventory control with lead time and lost sales:** Exo-MDP formulation leads to RL algorithms that beat bandit baselines and approach the performance of an oracle with full exogenous information [2409.14557].
- **Wireless network configuration and control:** Removing exogenous signal via reward filtering allows Q-learning to converge with fewer samples [1806.01584].
- **Restless multi-armed bandits:** The LEMP algorithm exploits exogenous filtering, outperforming classical DSEE and simple baselines both in theoretical regret and empirical benchmarks [2112.09484].

Empirical validation consistently shows that reward decomposition, when exogeneity and additivity conditions hold even approximately, results in improved sample efficiency and learning speed. In cases of violation (e.g., strong anti-correlation between exogenous and endogenous returns), the predicted speedup may be absent, as theoretically characterized via variance–covariance criteria [1806.01584].

## 6. Practical Guidelines, Limitations, and Open Questions

To apply exogenous filtering in new settings [1806.01584, 2303.12957], the following workflow is typical:

1. Collect initial transition data under exploratory policy.
2. Recover or specify an exogenous subspace using conditional correlation constraints (often via manifold optimization or stepwise search).
3. Regress observed rewards onto exogenous state, subtract this component to construct endogenous reward.
4. Train or continue RL using only endogenous reward, optionally repeating decomposition as more data accrue.

**Limitations and open questions:**
- The additivity and independence assumptions underlying reward decomposition must approximately hold. Strong violations (e.g., non-additive or fully enmeshed rewards) preclude effective filtering.
- Conditional independence may break down in nonlinear or non-Gaussian settings. The bulk of existing methodology uses linear projections; extending to nonlinear feature maps or deep learning architectures is ongoing.
- PCC/CCC are only approximations to conditional mutual information; finite-sample behavior and appropriate thresholds require additional care.
- Exploration strategies tailored to decomposition discovery, especially in the online RL case, remain underdeveloped [1806.01584].
- Rank reduction of the information matrix yields further gains for Exo-MDPs, but realistic identification of minimal-dimensional endogenous spaces remains challenging [2409.14557].

**Summary Table: Core Decomposition and Filtering Principles**

| Principle                        | Formulation/Algorithm                          | Key Reference   |
|-----------------------------------|-----------------------------------------------|-----------------|
| Additive reward decomposition     | $R(x,a) = R_{exo}(e) + R_{end}(s,e,a)$        | [2303.12957]    |
| Transition factorization          | $P(x'|x,a) = P(s'|s,a)P(e'|s,e,a)$            | [2401.12934]    |
| Thresholded-Lasso support         | Filter features via $|\beta| \geq \tau$       | [2401.12934]    |
| Subspace discovery (PCC/CCC)      | Manifold or stepwise optimization             | [2303.12957]    |
| Exo-MDP linear mixture            | Transition/reward as $⟨\phi,\Xi⟩$             | [2409.14557]    |
| Endo-MDP policy optimality        | Policy depends only on endogenous part        | [1806.01584]    |

The consensus emerging from this research area is that identifying and exploiting reward decomposition and exogenous structure is a powerful way to achieve sample-efficient RL, especially in high-dimensional settings where only a small subset of the state-action space is reward-relevant. Continued progress depends on advances in non-linear discovery and robustification to model assumption violations.

Source: https://www.emergentmind.com/topics/reward-decomposition-and-exogenous-filtering