Optimal Objective for Representation Learning in Image-Based World Models
Determine the optimal objective function for learning latent state representations in image-based model-based reinforcement learning world models, such as those based on the Recurrent State-Space Model (RSSM), so that the learned representations emphasize task-essential information while avoiding overfitting to irrelevant visual details.
References
While architectures like the Recurrent State-Space Model (RSSM) have achieved remarkable success (Hafner et al., 2025), a fundamental question remains open: What is the optimal objective function for learning the representation itself?
— R2-Dreamer: Redundancy-Reduced World Models without Decoders or Augmentation
(2603.18202 - Morihira et al., 18 Mar 2026) in Section 1 (Introduction)
Whether training against a reachability-aware criterion would produce a representation whose plain L2 is usable is an open and more interesting question.
— The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use
(2608.12959 - Singh, 13 Aug 2026) in Section 9, Limitations