Papers
Topics
Authors
Recent
Search
2000 character limit reached

BellNet: Graph-Filtered MDP Solver

Updated 3 July 2026
  • BellNet is a parametric model that reinterprets policy iteration for MDPs using nonlinear graph filtering to update value functions.
  • The architecture unrolls layers of graph-filtered evaluation and softmax-based policy improvement to minimize the Bellman error.
  • Empirical results show near-optimal performance with reduced iterations and robust transferability compared to classical dynamic programming solvers.

BellNet is a parametric, learnable model architecture that reinterprets policy iteration for Markov Decision Processes (MDPs) through the lens of graph signal processing, casting the value function update process as a truncated cascade of nonlinear graph filters. Designed to minimize the Bellman error from random value function initializations, BellNet unifies, compresses, and generalizes dynamic programming (DP) solvers by stacking a finite sequence of truncated, learnable policy-evaluation steps and differentiable policy-improvement steps. Each layer of BellNet acts as a nonlinear graph filter on the state-action transition graph, facilitating efficient, fully differentiable inference and providing explicit control over computational complexity via tunable depth and filter order (Rozada et al., 29 Jul 2025).

1. Architecture and Computation

BellNet is constructed by unrolling KK policy-evaluation steps and a single, differentiable policy-improvement step into an elementary block. Multiple such blocks (L+1L+1 in total) are stacked, forming a deep network. The architecture is governed by the following operations at each block ℓ→ℓ+1\ell \rightarrow \ell+1:

  • Graph-filtered evaluation:

Q~(ℓ+1)=∑j=0Khj(ℓ)(Pπ(ℓ))j  vec(R)+hK+1(ℓ)(Pπ(ℓ))K+1Q(ℓ)\widetilde Q^{(\ell+1)} = \sum_{j=0}^{K} h_j^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^j\;\mathrm{vec}(R) + h_{K+1}^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^{K+1} Q^{(\ell)}

where hj(ℓ)h_j^{(\ell)} are learnable filter coefficients per layer.

  • Differentiable policy improvement:

π(ℓ+1)=στ(unvec(Q~(ℓ+1))),[στ(Q)]s,a=exp⁡(Qs,a/τ)∑a′exp⁡(Qs,a′/τ)\pi^{(\ell+1)} = \sigma_\tau(\mathrm{unvec}(\widetilde Q^{(\ell+1)})), \quad [\sigma_\tau(Q)]_{s,a} = \frac{\exp(Q_{s,a}/\tau)}{\sum_{a'} \exp(Q_{s,a'}/\tau)}

The nonlinearity is introduced via the softmax operator, parameterized by temperature τ\tau for exploration-exploitation tradeoff.

After L+1L+1 such blocks, the network outputs the terminal value function estimate Q^\hat Q and policy π^\hat \pi, yielding the parametric map L+1L+10.

2. Graph Filter Interpretation

BellNet draws from graph signal processing by treating the MDP’s (policy-induced) transition probability matrix L+1L+11 as the adjacency matrix L+1L+12 of a weighted directed graph over state–action pairs. Each layer implements a polynomial graph filter:

L+1L+13

where the L+1L+14 are layer-specific, learnable parameters. The L+1L+15 term provides a residual-bias mechanism akin to skip connections. As L+1L+16 evolves with each layer, BellNet constitutes a cascade of nonlinear (since L+1L+17 depends on the intermediate Q-value) graph filters, collectively enabling nonlinear, policy-adaptive propagation and aggregation of reward and value information.

3. Training Objective and Optimization

BellNet is trained to minimize the squared Bellman error for the optimality operator

L+1L+18

where policy selection is relaxed from L+1L+19 to softmax for differentiability. For a random initial value function ℓ→ℓ+1\ell \rightarrow \ell+10,

ℓ→ℓ+1\ell \rightarrow \ell+11

is computed, and all filter coefficients ℓ→ℓ+1\ell \rightarrow \ell+12 are updated via backpropagation through ℓ→ℓ+1\ell \rightarrow \ell+13 unrolled blocks using standard optimizers such as SGD or Adam. Random restarts of ℓ→ℓ+1\ell \rightarrow \ell+14 across epochs are used to enhance robustness to initialization.

4. Algorithmic Recipe and Practical Considerations

Key procedural elements of BellNet include:

  • Initialization: Randomize filter coefficients ℓ→ℓ+1\ell \rightarrow \ell+15 and sample initial ℓ→ℓ+1\ell \rightarrow \ell+16 per epoch.
  • Forward passes: At each unrolled layer, construct ℓ→ℓ+1\ell \rightarrow \ell+17, where ℓ→ℓ+1\ell \rightarrow \ell+18, apply graph-filtered evaluation, followed by differentiable policy improvement.
  • Loss computation: Compute the squared Bellman error at the final block.
  • Backward pass: Apply backpropagation to train ℓ→ℓ+1\ell \rightarrow \ell+19.
  • Inference: For deployment, propagate any Q~(ℓ+1)=∑j=0Khj(ℓ)(Pπ(ℓ))j  vec(R)+hK+1(ℓ)(Pπ(ℓ))K+1Q(ℓ)\widetilde Q^{(\ell+1)} = \sum_{j=0}^{K} h_j^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^j\;\mathrm{vec}(R) + h_{K+1}^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^{K+1} Q^{(\ell)}0 (e.g., zero or uniform) through Q~(ℓ+1)=∑j=0Khj(ℓ)(Pπ(ℓ))j  vec(R)+hK+1(ℓ)(Pπ(ℓ))K+1Q(ℓ)\widetilde Q^{(\ell+1)} = \sum_{j=0}^{K} h_j^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^j\;\mathrm{vec}(R) + h_{K+1}^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^{K+1} Q^{(\ell)}1 blocks with the trained coefficients. Optional weight sharing across layers (Q~(ℓ+1)=∑j=0Khj(ℓ)(Pπ(ℓ))j  vec(R)+hK+1(ℓ)(Pπ(ℓ))K+1Q(ℓ)\widetilde Q^{(\ell+1)} = \sum_{j=0}^{K} h_j^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^j\;\mathrm{vec}(R) + h_{K+1}^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^{K+1} Q^{(\ell)}2) allows arbitrary-depth inference and minimizes parameter count.

Additional practices involve tuning softmax temperature Q~(ℓ+1)=∑j=0Khj(ℓ)(Pπ(ℓ))j  vec(R)+hK+1(ℓ)(Pπ(ℓ))K+1Q(ℓ)\widetilde Q^{(\ell+1)} = \sum_{j=0}^{K} h_j^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^j\;\mathrm{vec}(R) + h_{K+1}^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^{K+1} Q^{(\ell)}3, employing early stopping for stabilization, and leveraging optional weight sharing for improved transfer capability and reduced variance.

5. Computational Complexity

Let Q~(ℓ+1)=∑j=0Khj(ℓ)(Pπ(ℓ))j  vec(R)+hK+1(ℓ)(Pπ(ℓ))K+1Q(ℓ)\widetilde Q^{(\ell+1)} = \sum_{j=0}^{K} h_j^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^j\;\mathrm{vec}(R) + h_{K+1}^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^{K+1} Q^{(\ell)}4 denote the dimension of the value function and Q~(ℓ+1)=∑j=0Khj(ℓ)(Pπ(ℓ))j  vec(R)+hK+1(ℓ)(Pπ(ℓ))K+1Q(ℓ)\widetilde Q^{(\ell+1)} = \sum_{j=0}^{K} h_j^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^j\;\mathrm{vec}(R) + h_{K+1}^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^{K+1} Q^{(\ell)}5 the number of nonzero entries in Q~(ℓ+1)=∑j=0Khj(ℓ)(Pπ(ℓ))j  vec(R)+hK+1(ℓ)(Pπ(ℓ))K+1Q(ℓ)\widetilde Q^{(\ell+1)} = \sum_{j=0}^{K} h_j^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^j\;\mathrm{vec}(R) + h_{K+1}^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^{K+1} Q^{(\ell)}6 (Q~(ℓ+1)=∑j=0Khj(ℓ)(Pπ(ℓ))j  vec(R)+hK+1(ℓ)(Pπ(ℓ))K+1Q(ℓ)\widetilde Q^{(\ell+1)} = \sum_{j=0}^{K} h_j^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^j\;\mathrm{vec}(R) + h_{K+1}^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^{K+1} Q^{(\ell)}7 for sparse graphs). The cost of applying Q~(ℓ+1)=∑j=0Khj(ℓ)(Pπ(ℓ))j  vec(R)+hK+1(ℓ)(Pπ(ℓ))K+1Q(ℓ)\widetilde Q^{(\ell+1)} = \sum_{j=0}^{K} h_j^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^j\;\mathrm{vec}(R) + h_{K+1}^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^{K+1} Q^{(\ell)}8 to a vector is Q~(ℓ+1)=∑j=0Khj(ℓ)(Pπ(ℓ))j  vec(R)+hK+1(ℓ)(Pπ(ℓ))K+1Q(ℓ)\widetilde Q^{(\ell+1)} = \sum_{j=0}^{K} h_j^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^j\;\mathrm{vec}(R) + h_{K+1}^{(\ell)} \left(P^{\pi^{(\ell)}}\right)^{K+1} Q^{(\ell)}9; thus, a full graph filtering step in one BellNet layer (order hj(ℓ)h_j^{(\ell)}0) is hj(ℓ)h_j^{(\ell)}1. The subsequent softmax incurs hj(ℓ)h_j^{(\ell)}2. For hj(ℓ)h_j^{(\ell)}3 layers, overall inference complexity is hj(ℓ)h_j^{(\ell)}4.

For comparison:

  • Value iteration: hj(ℓ)h_j^{(\ell)}5 for hj(ℓ)h_j^{(\ell)}6-accuracy.
  • Classic policy iteration: hj(ℓ)h_j^{(\ell)}7 per step, but hj(ℓ)h_j^{(\ell)}8 to convergence in typical scenarios.

Empirical results report that BellNet achieves near-optimal performance with hj(ℓ)h_j^{(\ell)}9, where classic policy iteration requires at least π(ℓ+1)=στ(unvec(Q~(ℓ+1))),[στ(Q)]s,a=exp⁡(Qs,a/τ)∑a′exp⁡(Qs,a′/τ)\pi^{(\ell+1)} = \sigma_\tau(\mathrm{unvec}(\widetilde Q^{(\ell+1)})), \quad [\sigma_\tau(Q)]_{s,a} = \frac{\exp(Q_{s,a}/\tau)}{\sum_{a'} \exp(Q_{s,a'}/\tau)}0 full iterations (Rozada et al., 29 Jul 2025).

6. Empirical Performance and Transferability

In grid-world experiments (Gymnasium’s “cliff-walking” MDP, per-step reward π(ℓ+1)=στ(unvec(Q~(ℓ+1))),[στ(Q)]s,a=exp⁡(Qs,a/τ)∑a′exp⁡(Qs,a′/τ)\pi^{(\ell+1)} = \sigma_\tau(\mathrm{unvec}(\widetilde Q^{(\ell+1)})), \quad [\sigma_\tau(Q)]_{s,a} = \frac{\exp(Q_{s,a}/\tau)}{\sum_{a'} \exp(Q_{s,a'}/\tau)}1, large negative cliff penalty), BellNet is assessed using normalized value function (VF) error:

π(ℓ+1)=στ(unvec(Q~(ℓ+1))),[στ(Q)]s,a=exp⁡(Qs,a/τ)∑a′exp⁡(Qs,a′/τ)\pi^{(\ell+1)} = \sigma_\tau(\mathrm{unvec}(\widetilde Q^{(\ell+1)})), \quad [\sigma_\tau(Q)]_{s,a} = \frac{\exp(Q_{s,a}/\tau)}{\sum_{a'} \exp(Q_{s,a'}/\tau)}2

Key results with π(ℓ+1)=στ(unvec(Q~(ℓ+1))),[στ(Q)]s,a=exp⁡(Qs,a/τ)∑a′exp⁡(Qs,a′/τ)\pi^{(\ell+1)} = \sigma_\tau(\mathrm{unvec}(\widetilde Q^{(\ell+1)})), \quad [\sigma_\tau(Q)]_{s,a} = \frac{\exp(Q_{s,a}/\tau)}{\sum_{a'} \exp(Q_{s,a'}/\tau)}3 and π(ℓ+1)=στ(unvec(Q~(ℓ+1))),[στ(Q)]s,a=exp⁡(Qs,a/τ)∑a′exp⁡(Qs,a′/τ)\pi^{(\ell+1)} = \sigma_\tau(\mathrm{unvec}(\widetilde Q^{(\ell+1)})), \quad [\sigma_\tau(Q)]_{s,a} = \frac{\exp(Q_{s,a}/\tau)}{\sum_{a'} \exp(Q_{s,a'}/\tau)}4 include:

  • Sub-3% normalized error using only four BellNet blocks, versus π(ℓ+1)=στ(unvec(Q~(ℓ+1))),[στ(Q)]s,a=exp⁡(Qs,a/τ)∑a′exp⁡(Qs,a′/τ)\pi^{(\ell+1)} = \sigma_\tau(\mathrm{unvec}(\widetilde Q^{(\ell+1)})), \quad [\sigma_\tau(Q)]_{s,a} = \frac{\exp(Q_{s,a}/\tau)}{\sum_{a'} \exp(Q_{s,a'}/\tau)}515% error for policy iteration after 10 cycles.
  • Weight sharing (BN-WS) reduces VF error variance and improves median error.
  • Increasing filter order π(ℓ+1)=στ(unvec(Q~(ℓ+1))),[στ(Q)]s,a=exp⁡(Qs,a/τ)∑a′exp⁡(Qs,a′/τ)\pi^{(\ell+1)} = \sigma_\tau(\mathrm{unvec}(\widetilde Q^{(\ell+1)})), \quad [\sigma_\tau(Q)]_{s,a} = \frac{\exp(Q_{s,a}/\tau)}{\sum_{a'} \exp(Q_{s,a'}/\tau)}6 to π(ℓ+1)=στ(unvec(Q~(ℓ+1))),[στ(Q)]s,a=exp⁡(Qs,a/τ)∑a′exp⁡(Qs,a′/τ)\pi^{(\ell+1)} = \sigma_\tau(\mathrm{unvec}(\widetilde Q^{(\ell+1)})), \quad [\sigma_\tau(Q)]_{s,a} = \frac{\exp(Q_{s,a}/\tau)}{\sum_{a'} \exp(Q_{s,a'}/\tau)}7 yields marginal accuracy gains beyond π(ℓ+1)=στ(unvec(Q~(ℓ+1))),[στ(Q)]s,a=exp⁡(Qs,a/τ)∑a′exp⁡(Qs,a′/τ)\pi^{(\ell+1)} = \sigma_\tau(\mathrm{unvec}(\widetilde Q^{(\ell+1)})), \quad [\sigma_\tau(Q)]_{s,a} = \frac{\exp(Q_{s,a}/\tau)}{\sum_{a'} \exp(Q_{s,a'}/\tau)}8.
  • In transfer settings (e.g., mirroring the environment/cliff), BellNet’s error degrades only slightly without retraining, unlike classical methods which necessitate full reruns.

This suggests BellNet’s learned, parametric graph-filter-based updates offer significant efficiency and transfer advantages by decoupling update complexity from the number of classical DP iterations (Rozada et al., 29 Jul 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to BellNet.