---
title: Feudal Reinforcement Learning
url: https://www.emergentmind.com/topics/feudal-reinforcement-learning
type: topic
---

# Feudal Reinforcement Learning

Feudal Reinforcement Learning (FRL) is a class of hierarchical reinforcement learning (HRL) frameworks in which high-level “manager” policies decompose long-horizon planning into temporally extended goals (subgoals) issued to lower-level “worker” policies. Each level in the hierarchy operates at a distinct spatial or temporal abstraction, with managers setting goals for workers, which in turn either enact low-level primitive actions or further delegate via sub-managers. FRL fundamentally enables reward and credit assignment decomposition, temporal abstraction, and modularity, and has become a central paradigm in scalable RL for both single-agent and multi-agent systems.

## 1. Hierarchical Architectures and Formal Description

FRL is generally instantiated as a multi-level hierarchy, with the classical structure comprising at least a manager and a set of workers. At each timestep, the manager observes a coarse abstraction of the environment and outputs a latent goal $g_t$ or subgoal $g_{t}^{m\to s}$ for a sub-manager or $g_{t}^{m\to w}$ for a worker. The workers, conditioned on both their local observation and the goal, then choose primitive actions $a_t^w$. This process is formalized as:

- Three-level hierarchy: a top-level manager $m$, multiple sub-managers $\{s\}$, and workers $\{w\}$ [2507.23604].
    - At each higher level, goals are issued every $K\alpha_s$ or $\alpha_s$ steps, introducing temporal abstraction.
- Policies at each level:
    $$
    \begin{aligned}
    \pi_m &: \mathcal{Z}_m \rightarrow \mathcal{N}(\mu_m(\cdot), \Sigma_m(\cdot)) \quad (\text{sample } g_t^{m\to s}) \\
    \pi_s &: \mathcal{Z}_s \rightarrow \mathcal{N}(\mu_s(\cdot), \Sigma_s(\cdot)) \quad (\text{sample } g_t^{s\to w}) \\
    \pi_w &: \mathcal{Z}_w \rightarrow \mathcal{P}(\mathcal{A}) \quad (\text{sample } a_t^w)
    \end{aligned}
    $$
    where each level uses its own latent embedding $\mathcal{Z}_m, \mathcal{Z}_s, \mathcal{Z}_w$ composed via message-passing or local observation.

This framework enables decomposition of planning and control across both time and space, and can be implemented for partially observable stochastic games (POSG) or centralized MDPs [2507.23604, 2310.05695].

## 2. Temporal and Spatial Abstraction

Core to FRL is the separation of decision-making into different resolutions:

- **Temporal abstraction:** High-level managers operate at a coarser timescale, issuing commands that persist over multiple worker actions (parameter $k$ or $c$). Workers act at every environment step under fixed goals until the next high-level decision [1703.01161, 2310.05695, 2511.17351].
- **Spatial abstraction:** Manager and workers may have access to distinct abstract state spaces. In gridworlds, the manager might observe downsampled (e.g., $2\times2$) grids, while workers operate on the full $4\times4$ environment [2310.05695]. In modular agent morphologies, hierarchical policy graphs reflect body-part clustering, with higher layers controlling higher-order submodules [2304.05099, 2507.23604].

Abstraction at these axes provides both faster learning and more stable credit assignment, as workers focus only on achieving local objectives, while managers optimize over longer-term, sparser extrinsic reward signals.

## 3. Reward Decomposition and Intrinsic Motivation

Reward assignment in FRL decouples local and global credit assignment by introducing intrinsic rewards at lower levels and extrinsic rewards at higher levels:

- **Intrinsic rewards**: Workers receive shaped intrinsic rewards that measure alignment with manager goals, often via state-difference cosine similarity or advantage function of the parent policy [1703.01161, 2507.23604, 2307.06742]. For example, the worker reward may be:
  $$
  r_t^w = \frac{1}{\alpha_s} A^{\pi_s}(z_t^{s\to w}, g_t^{s\to w})
  $$
  where $A^{\pi_s}$ is the advantage function of the sub-manager [2507.23604].
- **Extrinsic rewards**: Managers are optimized directly on the (possibly sparse) environment reward, e.g., average return or win rate, aggregated over appropriate timescales.

Theoretical results show that, under suitable assumptions (e.g., $\gamma \approx 1$), maximizing the expected return at each level leads to maximization of the global extrinsic return [2507.23604].

## 4. Policy Learning and Optimization Algorithms

Learning in FRL comprises parallel, often coordinated training at each level. Typical learning algorithms include:

- **Two-timescale Q-learning**: High-level and low-level Q-functions are updated via coupled stochastic approximation, with workers' updates on a faster timescale than managers [2511.17351, 2310.05695]. Stability and convergence can be established via the ODE method under appropriate assumptions.
- **Actor-Critic methods**: Both managers and workers use policy-gradient methods, such as A2C or PPO, with level-specific baselines and possibly decentralized value functions [2507.23604, 2307.06742].
- **Black-box optimization**: Hierarchies with complex or unstable gradients may use evolutionary strategies (e.g., CMA-ES) at each level independently [2304.05099].

Central to most FRL algorithms is a pseudocode loop where, at each episode, the manager generates goals or macro-actions, workers execute policies for a set number of steps or until goal attainment, and each policy is updated from level-specific or global transitions [2507.23604, 2511.17351, 2310.05695].

## 5. Hierarchical Graphs and Multi-Agent Extensions

Modern FRL approaches use explicit graph structures to organize agents and communication:

- **Graph-based message-passing**: Hierarchies are encoded as layered directed graphs, with message-passing among same-level nodes to enable robust coordination under partial observability [2507.23604, 2304.05099].
- **Partition- or region-based hierarchies**: MARL settings often divide the environment into dynamic regions managed by higher-level policies, with workers controlling local agents. Adaptive partitioning uses GNNs (e.g., DiffPool) or MCTS to optimize the decomposition [2205.13836].
- **Multi-agent feudal hierarchies**: Manager(s) coordinate multiple workers with intrinsic rewards; in decentralized settings, managers set subgoals for cohorts of agents, enhancing scalability and stability [1901.08492, 2507.23604].

This explicit graph structure captures morphological or spatial relations and enables scalable policy transfer and compositionality.

## 6. Applications and Empirical Performance

FRL has been applied to a variety of domains:

- **Multi-agent coordination**: StarCraft Multi-Agent Challenge (SMAC), Level-Based Foraging, and room clearance tasks all demonstrated superior coordination and scalability versus flat RL baselines [2507.23604, 2105.11328, 1901.08492].
- **Resource allocation and logistics**: In intercity ride-pooling and traffic signal control, feudal MARL with partitioned managers and intrinsic reward shaping decisively improved supply-demand balance, order-fulfillment rates, and system profit [2307.06742, 2205.13836].
- **Sequence prediction and control**: Feudal Q-learning and its DQN variants achieved faster and more consistent learning on tasks ranging from stock trading to vehicle steering, with explicit multi-resolution decomposition [2310.05695].
- **Robotics and modular agents**: Hierarchical graph FRL produced more natural, modular control strategies and generalizes well to unseen morphologies or larger agent populations [2304.05099].
- **Language-guided tasks**: Hierarchical manager–worker architectures bridge text-level reasoning and low-level control, achieving near-optimal performance without manual curriculum learning [2110.06477].

## 7. Theoretical Guarantees and Limitations

Recent research has addressed the convergence, stability, and game-theoretic properties of FRL algorithms:

- **Convergence and stability**: Under bounded rewards, finite state/action spaces, Lipschitz policies, on-policy sampling, and suitable step-size schedules, two-timescale Feudal Q-learning converges almost surely to a Nash/Stackelberg equilibrium of the coupled Bellman equations [2511.17351].
- **Sample efficiency**: Temporal and spatial abstraction confers significant sample efficiency: FRL can reduce the sample requirements to reach high performance by up to 40% or more compared to flat baselines on a range of domains [2310.05695, 2507.23604].
- **Limitations**: Fixed subgoal sets, dependency on manual abstraction or subgoal representation, and instability with deep, joint end-to-end optimization remain open challenges [2304.05099, 1901.08492].

## 8. Outlook and Future Directions

FRL frameworks are advancing toward more autonomous subgoal discovery and symbolic abstraction, leveraging set-based reachability, adaptively partitioned networks, and learned goal representations [2309.07675, 2205.13836]. Promising future directions include:

- **Discoverable/learned subgoal spaces**: Autonomous abstraction of symbolic or continuous goals for generalization and transferability.
- **Deeper and more flexible hierarchies**: Moving beyond two- or three-level hierarchies to multi-level, dynamically reconfigurable structures.
- **Integration with centralized critics and richer communication**: Hybrid methods enhancing stability and coordination in MARL.
- **Application to language, vision, and real-world robotics**: Bridging high-level reasoning and low-level control across domains and sensory modalities.

Feudal RL thus represents a principled, theoretically grounded, and empirically validated framework for hierarchical policy decomposition, scalable credit assignment, and improved learning efficiency in complex RL environments [2507.23604, 2511.17351, 2304.05099, 2310.05695].

Source: https://www.emergentmind.com/topics/feudal-reinforcement-learning