Discounted Empowerment in Reinforcement Learning
- Discounted empowerment is an information-theoretic measure that aggregates multi-step channel capacities with a discount factor to balance immediate and longer-term control.
- It is applied as an intrinsic reward in RL pre-training, enabling policies to achieve data-efficient adaptation in both deterministic and stochastic environments.
- The approach reduces the hyperparameter burden of horizon selection by discounting extended horizons, thereby mitigating variance and improving sample efficiency.
Searching arXiv for the cited and related empowerment literature to ground the article with current paper identifiers. I’m checking arXiv entries for discounted empowerment and closely related intrinsic-motivation work. Discounted empowerment is an information-theoretic measure of an agent’s potential influence on its environment that extends classical empowerment by aggregating multi-step channel capacities across horizons with a discount factor. In the formulation introduced for reinforcement-learning pre-training, the quantity is defined as , where balances short- and long-term control and is the finite episode horizon. This construction is used as an intrinsic reward for policy initialization, with the stated goal of producing broadly competent policies that adapt data-efficiently to downstream tasks. Related control-capacity results for partially observed continuous-time systems show that empowerment is constrained by system time constants and that only a near-past window of control contributes effectively to future influence (Schneider et al., 7 Oct 2025, Tiomkin et al., 2017).
1. Classical empowerment and the discounted extension
Classical empowerment treats the transition kernel as a channel with input and output . For a state , the contextual one-step empowerment is the channel capacity
Because the maximization is over input distributions , is a property of the environment at 0, not of any particular policy (Schneider et al., 7 Oct 2025).
The multi-step generalization measures control over longer timescales. For horizon 1,
2
This quantifies the mutual information between an action sequence and a terminal state at a fixed horizon. A related horizon-3 channel-capacity form is
4
which is consistent with the 5-step expression when conditioning is handled with respect to the initial state and the policy-induced state sequence.
Discounted empowerment aggregates these horizon-specific capacities: 6 The construction is explicitly a horizon-weighted variant, not per-step information along a trajectory. Its stated purpose is to avoid committing to a single horizon while retaining the controllability semantics of empowerment.
2. Discounting, saturation, and controllability across time scales
The primary motivation for discounting is that very long-horizon empowerment can become nearly uniform when most states are reachable. In that regime, empowerment loses its ability to differentiate states by controllability. Discounting down-weights very long horizons while still accounting for longer-term influence. The factor 7 plays a role analogous to 8 in RL returns: larger 9 emphasizes long-term control, whereas smaller 0 emphasizes immediate control (Schneider et al., 7 Oct 2025).
The limiting cases are explicit. As 1, discounted empowerment approaches one-step empowerment,
2
As 3 with finite 4, it approaches the unweighted sum 5. In the reported experiments, 6 is fixed to the episode length, 7, which both bounds the quantity and keeps computation tractable.
An information-geometric interpretation is also given for the one-step case: 8 Under this view, maximizing mutual information selects an average next-state distribution 9 that minimizes the worst-case KL divergence to the dynamics-conditioned outputs 0. The capacity-achieving policy 1 induces an output distribution near the center of the “information ball” of achievable next-state distributions. The paper argues that this yields an initialization that is, on average, close in an information-theoretic sense to many downstream policies, which can enhance adaptability (Schneider et al., 7 Oct 2025).
Continuous-time control-capacity analysis provides a complementary time-scale justification. For partially observable linear-Gaussian systems, empowerment is the control-to-output information capacity over a finite horizon. The reported result is that the capacity between the control signal and the system output does not grow without limits with the length of the control signal; only the near-past window contributes effectively, and empowerment depends on a time constant of the dynamic system (Tiomkin et al., 2017). This suggests a physical interpretation of discounting: the effective influence horizon is shaped by dynamical memory rather than by an arbitrary truncation alone.
3. Intrinsic-reward pre-training and estimation procedures
The pre-training objective assigns intrinsic reward equal to discounted empowerment: 2 and maximizes the expected discounted return
3
Here 4 is the standard RL discount factor used by the learning algorithm, whereas 5 is the empowerment discount. In the reported experiments, 6 and 7 are fixed across environments (Schneider et al., 7 Oct 2025).
Empowerment estimation is environment-dependent. In deterministic gridworlds, one-step empowerment reduces to the logarithm of the number of distinct next states reachable, and 8-step empowerment becomes
9
In stochastic gridworlds, the channel capacity is computed with the Blahut–Arimoto algorithm for each horizon, yielding both 0 and the capacity-achieving source distributions 1. No variational mutual-information critic is trained; empowerment values are precomputed offline and used as a static intrinsic reward map.
The paper distinguishes two pre-training regimes. In capacity-achieving pre-training, 2 is obtained independently for each state through Blahut–Arimoto, and a policy network is initialized by behavior cloning these per-state action distributions. In capacity-maximizing pre-training, the policy is optimized by RL to navigate toward states with high 3, that is, to maximize cumulative empowerment reward. This distinction is important: the first tracks local capacity-achieving action distributions, whereas the second learns trajectories that seek globally high-empowerment regions.
The implementation sequence is explicit. First, empowerment maps are computed for every state and every horizon 4, then aggregated into 5. Second, a standard RL algorithm—REINFORCE, Actor–Critic, PPO, or DQN—is used with 6 as intrinsic reward, together with entropy regularization. Third, fine-tuning replaces the intrinsic reward with the downstream extrinsic reward while keeping the policy architecture and the same RL algorithm (Schneider et al., 7 Oct 2025).
4. Implementation regime, algorithmic machinery, and computational considerations
The reported implementations use neural networks as policy parameterizations: MLPs for low-dimensional states and CNNs for image-based PPO and DQN. On-policy methods—REINFORCE, Actor–Critic, and PPO—use entropy bonuses, and PPO uses generalized advantage estimation. Off-policy DQN uses a replay buffer, target network updates, and 7-greedy exploration. The overall implementation uses JAX, PPO and DQN are implemented via Stable-Baselines3, and Adam is used throughout (Schneider et al., 7 Oct 2025).
The paper fixes 8 to the episode length, 9, and reports that performance is largely insensitive to 0, which is fixed at 1 across experiments. By contrast, plain 2-step empowerment is sensitive to the choice of a single 3, and the discounted sum is therefore presented as reducing horizon-selection burden. Truncation at the episode length keeps computation tractable and mitigates the flattening of empowerment at very long horizons, while discounting reduces variance from overly long horizons but preserves useful long-term control structure.
Computational cost is dominated by empowerment estimation. Blahut–Arimoto must be run per state and per horizon in stochastic environments, which is expensive in large spaces. Deterministic reachability counts are comparatively efficient. Precomputing empowerment maps avoids training a noisy MI critic and thereby reduces variance during pre-training. PPO and DQN training used GPUs, whereas REINFORCE and Actor–Critic used CPUs; entropy regularization is reported to stabilize learning.
The paper also records concrete hyperparameter settings. REINFORCE and Actor–Critic use MLPs with two hidden layers of 256 units, Adam learning rate approximately 4, entropy coefficient 5, batch size 6, multiple random seeds, and pre-training and fine-tuning of up to 7 environment steps across parallel environments. PPO uses a CNN backbone as in DQN Atari, learning rate approximately 8, clip range 9, GAE-0, entropy coefficient 1, batch size 2, 3 pre-training steps, and 4 fine-tuning steps. DQN uses a replay buffer of 5, learning rate approximately 6, 7-greedy exploration from 8 to 9 over 0 steps, target update every 1 steps with 2, 3 pre-training steps, and 4 fine-tuning steps. The RL discount factor is fixed at 5 (Schneider et al., 7 Oct 2025).
5. Empirical behavior in downstream adaptation
The empirical study uses deterministic and stochastic gridworlds with a goal-reaching downstream task. The evaluated algorithms are REINFORCE with and without baseline, Actor–Critic with a TD critic, PPO with GAE, and DQN. The reported metrics are fine-tuning learning curves in mean return versus steps, sample efficiency, and final performance (Schneider et al., 7 Oct 2025).
The principal finding is that empowerment-based pre-training consistently accelerates adaptation compared with training from scratch, especially for REINFORCE and Actor–Critic. Capacity-maximizing pre-training is reported to be more data-efficient than capacity-achieving behavior cloning, although both outperform training from scratch. Discounted empowerment outperforms fixed 6-step empowerment without horizon tuning, and the benefits of plain 7-step empowerment diminish when 8 is too short or too long. PPO shows small or no gains, which the paper attributes to already strong variance reduction, but empowerment pre-training is described as safe in the sense that it does not harm performance. DQN benefits in both faster learning and higher final performance. In stochastic gridworlds, empowerment pre-training still improves sample efficiency, indicating robustness under uncertainty (Schneider et al., 7 Oct 2025).
The ablation results reinforce the role of discounting. Varying 9 in plain 0-step empowerment shows sensitivity to horizon selection, whereas discounted empowerment reduces that hyperparameter burden. The fixed setting 1 performs well across experiments, and performance is reported to be largely insensitive to 2. Taken together, these findings position discounted empowerment as a general-purpose initialization strategy rather than a task-specific shaping signal.
6. Relations to adjacent intrinsic objectives, limitations, and open directions
Discounted empowerment is positioned against both standard empowerment and other intrinsic-motivation paradigms. Relative to undiscounted long-horizon empowerment, its stated advantage is that it avoids trivial or nearly uniform empowerment landscapes while retaining the controllability interpretation. Relative to entropy maximization, such as APT, the distinction is that entropy maximization encourages visiting diverse states but does not directly maximize action-conditioned controllability. Relative to skill-discovery methods such as DIAYN, VIC, DADS, and CIC, the distinction is that those methods maximize mutual information between skills and terminal states, whereas empowerment focuses on actions-to-states mutual information and does not require a skill identifier. Discounted empowerment is therefore presented as preferable when a single, generic pre-training objective is needed that balances short- and long-term control without skill conditioning or horizon tuning (Schneider et al., 7 Oct 2025).
Undiscounted 3 is not excluded. The paper notes that small, short-horizon tasks in which reachability does not flatten and the relevant horizon 4 is clear can use 5 directly. The discounted form is argued to be more robust for general unsupervised pre-training across tasks.
The current formulation also has explicit limitations. Precomputing empowerment requires either known dynamics or tractable estimation, and scaling Blahut–Arimoto to large, continuous, high-dimensional environments is challenging. Empowerment estimates can be biased in complex stochastic settings, and sensitivity remains to RL hyperparameters and architecture choices. PPO’s small gains are described as suggesting diminishing returns for algorithms that already have strong variance reduction. Failure modes include sparse or deceptive dynamics that produce misleading empowerment gradients, such as local traps with high short-term empowerment but poor long-term utility, and corruption of empowerment maps through estimator collapse or world-model error if learned models are used (Schneider et al., 7 Oct 2025).
Several extension paths are identified. For continuous domains, prior work is said to suggest Gaussian-channel approximations, convex estimators, or world-model-based approaches. Future directions listed for discounted empowerment include model-based empowerment with learned world models for large-scale visual RL, hierarchical discounted empowerment for temporally extended skills, improved MI estimators and scalable channel-capacity approximations, extension to partial observability through directed information and BA variants with feedback, and foundation-model-style pre-training over large video datasets. Continuous-time control-capacity analysis further suggests a principled route for choosing discounts from system dynamics: with dominant time constant 6, one may use a continuous kernel 7 and the discrete-time correspondence 8, aligning the effective empowerment horizon with physical memory decay (Tiomkin et al., 2017).