---
title: 'BellNet: Graph-Filtered MDP Solver'
url: https://www.emergentmind.com/topics/bellnet
type: topic
---

# BellNet: Graph-Filtered MDP Solver

BellNet is a parametric, learnable model architecture that reinterprets policy iteration for Markov Decision Processes (MDPs) through the lens of graph signal processing, casting the value function update process as a truncated cascade of nonlinear graph filters. Designed to minimize the Bellman error from random value function initializations, BellNet unifies, compresses, and generalizes dynamic programming (DP) solvers by stacking a finite sequence of truncated, learnable policy-evaluation steps and differentiable policy-improvement steps. Each layer of BellNet acts as a nonlinear graph filter on the state-action transition graph, facilitating efficient, fully differentiable inference and providing explicit control over computational complexity via tunable depth and filter order [2507.21705].

## 1. Architecture and Computation

BellNet is constructed by unrolling $K$ policy-evaluation steps and a single, differentiable policy-improvement step into an elementary block. Multiple such blocks ($L+1$ in total) are stacked, forming a deep network. The architecture is governed by the following operations at each block $\ell \rightarrow \ell+1$:

- **Graph-filtered evaluation:**
  $$
  \widetilde Q^{(\ell+1)} =
  \sum_{j=0}^{K} h_j^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^j\;\mathrm{vec}(R)
  + h_{K+1}^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^{K+1} Q^{(\ell)}
  $$
  where $h_j^{(\ell)}$ are learnable filter coefficients per layer.

- **Differentiable policy improvement:**
  $$
  \pi^{(\ell+1)} = \sigma_\tau(\mathrm{unvec}(\widetilde Q^{(\ell+1)})), \quad
  [\sigma_\tau(Q)]_{s,a} = \frac{\exp(Q_{s,a}/\tau)}{\sum_{a'} \exp(Q_{s,a'}/\tau)}
  $$
  The nonlinearity is introduced via the softmax operator, parameterized by temperature $\tau$ for exploration-exploitation tradeoff.

After $L+1$ such blocks, the network outputs the terminal value function estimate $\hat Q$ and policy $\hat \pi$, yielding the parametric map $(\hat Q, \hat \pi) = \mathrm{BellNet}(Q^{(0)};\{h_j^{(\ell)}\})$.

## 2. Graph Filter Interpretation

BellNet draws from graph signal processing by treating the MDP’s (policy-induced) transition probability matrix $P^\pi$ as the adjacency matrix $A$ of a weighted directed graph over state–action pairs. Each layer implements a polynomial graph filter:
$$
H(A) = \sum_{j=0}^K h_j A^j,
$$
where the $h_j$ are layer-specific, learnable parameters. The $h_{K+1}A^{K+1}Q^{(\ell)}$ term provides a residual-bias mechanism akin to skip connections. As $\pi^{(\ell)}$ evolves with each layer, BellNet constitutes a cascade of nonlinear (since $P^{\pi^{(\ell)}}$ depends on the intermediate Q-value) graph filters, collectively enabling nonlinear, policy-adaptive propagation and aggregation of reward and value information.

## 3. Training Objective and Optimization

BellNet is trained to minimize the squared Bellman error for the optimality operator
$$
\mathcal{B}[Q] = \mathrm{vec}(R) + \gamma P^{\pi[Q]} Q
$$
where policy selection is relaxed from $\arg\max$ to softmax for differentiability. For a random initial value function $Q^{(0)}$,
$$
\mathcal{L}(h) = \|\,\mathrm{vec}(R) + \gamma P^{\hat \pi}\hat Q - \hat Q\|_2^2
$$
is computed, and all filter coefficients $\{h_j^{(\ell)}\}$ are updated via backpropagation through $L+1$ unrolled blocks using standard optimizers such as SGD or Adam. Random restarts of $Q^{(0)}$ across epochs are used to enhance robustness to initialization.

## 4. Algorithmic Recipe and Practical Considerations

Key procedural elements of BellNet include:

- **Initialization:** Randomize filter coefficients $\{h_j^{(\ell)}\}$ and sample initial $Q^{(0)}$ per epoch.
- **Forward passes:** At each unrolled layer, construct $A := P^{\pi}$, where $\pi = \mathrm{softmax}(Q/\tau)$, apply graph-filtered evaluation, followed by differentiable policy improvement.
- **Loss computation:** Compute the squared Bellman error at the final block.
- **Backward pass:** Apply backpropagation to train $\{h_j^{(\ell)}\}$.
- **Inference:** For deployment, propagate any $Q^{(0)}$ (e.g., zero or uniform) through $L+1$ blocks with the trained coefficients. Optional weight sharing across layers ($h_j^{(\ell)} \equiv h_j$) allows arbitrary-depth inference and minimizes parameter count.

Additional practices involve tuning softmax temperature $\tau$, employing early stopping for stabilization, and leveraging optional weight sharing for improved transfer capability and reduced variance.

## 5. Computational Complexity

Let $N=|S|\cdot|A|$ denote the dimension of the value function and $M$ the number of nonzero entries in $P^{\pi^{(\ell)}}$ ($M=O(N)$ for sparse graphs). The cost of applying $A^j$ to a vector is $O(M)$; thus, a full graph filtering step in one BellNet layer (order $K$) is $O(KM)$. The subsequent softmax incurs $O(N|A|)$. For $L$ layers, overall inference complexity is $O(LKM)$.

For comparison:

- **Value iteration**: $O(M/(1-\gamma))$ for $\epsilon$-accuracy.
- **Classic policy iteration**: $O(M)$ per step, but $L_{\text{classic}} \gg L_{\text{BellNet}}$ to convergence in typical scenarios.

Empirical results report that BellNet achieves near-optimal performance with $L\approx 4$, where classic policy iteration requires at least $10$ full iterations [2507.21705].

## 6. Empirical Performance and Transferability

In grid-world experiments (Gymnasium’s “cliff-walking” MDP, per-step reward $-1$, large negative cliff penalty), BellNet is assessed using normalized value function (VF) error:
$$
nerr(Q, Q^*) = \|\frac{Q}{\|Q\|_2} - \frac{Q^*}{\|Q^*\|_2}\|_2^2
$$

Key results with $K=5$ and $L=4$ include:

- Sub-3% normalized error using only four BellNet blocks, versus $\sim$15% error for policy iteration after 10 cycles.
- Weight sharing (BN-WS) reduces VF error variance and improves median error.
- Increasing filter order $K$ to $10$ yields marginal accuracy gains beyond $L\geq5$.
- In transfer settings (e.g., mirroring the environment/cliff), BellNet’s error degrades only slightly without retraining, unlike classical methods which necessitate full reruns.

This suggests BellNet’s learned, parametric graph-filter-based updates offer significant efficiency and transfer advantages by decoupling update complexity from the number of classical DP iterations [2507.21705].

Source: https://www.emergentmind.com/topics/bellnet