Feudal Reinforcement Learning
- Feudal Reinforcement Learning is a hierarchical framework that decomposes decision-making by letting high-level managers assign subgoals to lower-level workers.
- It leverages temporal and spatial abstraction to streamline learning and improve credit assignment in single-agent and multi-agent settings.
- By integrating intrinsic reward mechanisms, graph-based coordination, and two-timescale learning, FRL enhances sample efficiency and scalability in complex environments.
Feudal Reinforcement Learning (FRL) is a class of hierarchical reinforcement learning (HRL) frameworks in which high-level “manager” policies decompose long-horizon planning into temporally extended goals (subgoals) issued to lower-level “worker” policies. Each level in the hierarchy operates at a distinct spatial or temporal abstraction, with managers setting goals for workers, which in turn either enact low-level primitive actions or further delegate via sub-managers. FRL fundamentally enables reward and credit assignment decomposition, temporal abstraction, and modularity, and has become a central paradigm in scalable RL for both single-agent and multi-agent systems.
1. Hierarchical Architectures and Formal Description
FRL is generally instantiated as a multi-level hierarchy, with the classical structure comprising at least a manager and a set of workers. At each timestep, the manager observes a coarse abstraction of the environment and outputs a latent goal or subgoal for a sub-manager or for a worker. The workers, conditioned on both their local observation and the goal, then choose primitive actions . This process is formalized as:
- Three-level hierarchy: a top-level manager , multiple sub-managers , and workers (Marzi et al., 31 Jul 2025).
- At each higher level, goals are issued every or steps, introducing temporal abstraction.
- Policies at each level:
where each level uses its own latent embedding 0 composed via message-passing or local observation.
This framework enables decomposition of planning and control across both time and space, and can be implemented for partially observable stochastic games (POSG) or centralized MDPs (Marzi et al., 31 Jul 2025, Johnson et al., 2023).
2. Temporal and Spatial Abstraction
Core to FRL is the separation of decision-making into different resolutions:
- Temporal abstraction: High-level managers operate at a coarser timescale, issuing commands that persist over multiple worker actions (parameter 1 or 2). Workers act at every environment step under fixed goals until the next high-level decision (Vezhnevets et al., 2017, Johnson et al., 2023, Manenti et al., 21 Nov 2025).
- Spatial abstraction: Manager and workers may have access to distinct abstract state spaces. In gridworlds, the manager might observe downsampled (e.g., 3) grids, while workers operate on the full 4 environment (Johnson et al., 2023). In modular agent morphologies, hierarchical policy graphs reflect body-part clustering, with higher layers controlling higher-order submodules (Marzi et al., 2023, Marzi et al., 31 Jul 2025).
Abstraction at these axes provides both faster learning and more stable credit assignment, as workers focus only on achieving local objectives, while managers optimize over longer-term, sparser extrinsic reward signals.
3. Reward Decomposition and Intrinsic Motivation
Reward assignment in FRL decouples local and global credit assignment by introducing intrinsic rewards at lower levels and extrinsic rewards at higher levels:
- Intrinsic rewards: Workers receive shaped intrinsic rewards that measure alignment with manager goals, often via state-difference cosine similarity or advantage function of the parent policy (Vezhnevets et al., 2017, Marzi et al., 31 Jul 2025, Si et al., 2023). For example, the worker reward may be:
5
where 6 is the advantage function of the sub-manager (Marzi et al., 31 Jul 2025).
- Extrinsic rewards: Managers are optimized directly on the (possibly sparse) environment reward, e.g., average return or win rate, aggregated over appropriate timescales.
Theoretical results show that, under suitable assumptions (e.g., 7), maximizing the expected return at each level leads to maximization of the global extrinsic return (Marzi et al., 31 Jul 2025).
4. Policy Learning and Optimization Algorithms
Learning in FRL comprises parallel, often coordinated training at each level. Typical learning algorithms include:
- Two-timescale Q-learning: High-level and low-level Q-functions are updated via coupled stochastic approximation, with workers' updates on a faster timescale than managers (Manenti et al., 21 Nov 2025, Johnson et al., 2023). Stability and convergence can be established via the ODE method under appropriate assumptions.
- Actor-Critic methods: Both managers and workers use policy-gradient methods, such as A2C or PPO, with level-specific baselines and possibly decentralized value functions (Marzi et al., 31 Jul 2025, Si et al., 2023).
- Black-box optimization: Hierarchies with complex or unstable gradients may use evolutionary strategies (e.g., CMA-ES) at each level independently (Marzi et al., 2023).
Central to most FRL algorithms is a pseudocode loop where, at each episode, the manager generates goals or macro-actions, workers execute policies for a set number of steps or until goal attainment, and each policy is updated from level-specific or global transitions (Marzi et al., 31 Jul 2025, Manenti et al., 21 Nov 2025, Johnson et al., 2023).
5. Hierarchical Graphs and Multi-Agent Extensions
Modern FRL approaches use explicit graph structures to organize agents and communication:
- Graph-based message-passing: Hierarchies are encoded as layered directed graphs, with message-passing among same-level nodes to enable robust coordination under partial observability (Marzi et al., 31 Jul 2025, Marzi et al., 2023).
- Partition- or region-based hierarchies: MARL settings often divide the environment into dynamic regions managed by higher-level policies, with workers controlling local agents. Adaptive partitioning uses GNNs (e.g., DiffPool) or MCTS to optimize the decomposition (Ma et al., 2022).
- Multi-agent feudal hierarchies: Manager(s) coordinate multiple workers with intrinsic rewards; in decentralized settings, managers set subgoals for cohorts of agents, enhancing scalability and stability (Ahilan et al., 2019, Marzi et al., 31 Jul 2025).
This explicit graph structure captures morphological or spatial relations and enables scalable policy transfer and compositionality.
6. Applications and Empirical Performance
FRL has been applied to a variety of domains:
- Multi-agent coordination: StarCraft Multi-Agent Challenge (SMAC), Level-Based Foraging, and room clearance tasks all demonstrated superior coordination and scalability versus flat RL baselines (Marzi et al., 31 Jul 2025, Charlesworth et al., 2021, Ahilan et al., 2019).
- Resource allocation and logistics: In intercity ride-pooling and traffic signal control, feudal MARL with partitioned managers and intrinsic reward shaping decisively improved supply-demand balance, order-fulfillment rates, and system profit (Si et al., 2023, Ma et al., 2022).
- Sequence prediction and control: Feudal Q-learning and its DQN variants achieved faster and more consistent learning on tasks ranging from stock trading to vehicle steering, with explicit multi-resolution decomposition (Johnson et al., 2023).
- Robotics and modular agents: Hierarchical graph FRL produced more natural, modular control strategies and generalizes well to unseen morphologies or larger agent populations (Marzi et al., 2023).
- Language-guided tasks: Hierarchical manager–worker architectures bridge text-level reasoning and low-level control, achieving near-optimal performance without manual curriculum learning (Wang et al., 2021).
7. Theoretical Guarantees and Limitations
Recent research has addressed the convergence, stability, and game-theoretic properties of FRL algorithms:
- Convergence and stability: Under bounded rewards, finite state/action spaces, Lipschitz policies, on-policy sampling, and suitable step-size schedules, two-timescale Feudal Q-learning converges almost surely to a Nash/Stackelberg equilibrium of the coupled Bellman equations (Manenti et al., 21 Nov 2025).
- Sample efficiency: Temporal and spatial abstraction confers significant sample efficiency: FRL can reduce the sample requirements to reach high performance by up to 40% or more compared to flat baselines on a range of domains (Johnson et al., 2023, Marzi et al., 31 Jul 2025).
- Limitations: Fixed subgoal sets, dependency on manual abstraction or subgoal representation, and instability with deep, joint end-to-end optimization remain open challenges (Marzi et al., 2023, Ahilan et al., 2019).
8. Outlook and Future Directions
FRL frameworks are advancing toward more autonomous subgoal discovery and symbolic abstraction, leveraging set-based reachability, adaptively partitioned networks, and learned goal representations (Zadem et al., 2023, Ma et al., 2022). Promising future directions include:
- Discoverable/learned subgoal spaces: Autonomous abstraction of symbolic or continuous goals for generalization and transferability.
- Deeper and more flexible hierarchies: Moving beyond two- or three-level hierarchies to multi-level, dynamically reconfigurable structures.
- Integration with centralized critics and richer communication: Hybrid methods enhancing stability and coordination in MARL.
- Application to language, vision, and real-world robotics: Bridging high-level reasoning and low-level control across domains and sensory modalities.
Feudal RL thus represents a principled, theoretically grounded, and empirically validated framework for hierarchical policy decomposition, scalable credit assignment, and improved learning efficiency in complex RL environments (Marzi et al., 31 Jul 2025, Manenti et al., 21 Nov 2025, Marzi et al., 2023, Johnson et al., 2023).