Joint Training and Evaluation of World Models in Non-Stationary Environments

Investigate joint training, continual updating, and rigorous evaluation protocols for world models used by large language model-based agents in non-stationary environments, and ascertain the causal impact of these world models on downstream planning reliability.

Background

World-model-based agents aim to mitigate myopic reasoning via internal simulation and lookahead. Although model-based RL systems like DreamerV3 show the effectiveness of imagined rollouts, current LLM-based agents often rely on ad hoc representations trained on short-horizon, environment-specific data.

Only a few efforts explore co-evolving world models and agents over time. Establishing methods to jointly train, update, and evaluate world models under non-stationarity—and to quantify their causal influence on planning—remains a core challenge.

References

An open problem is how to jointly train, update, and evaluate world models in non-stationary environments, and how to assess their causal impact on downstream planning reliability.

— Agentic Reasoning for Large Language Models  (2601.12538 - Wei et al., 18 Jan 2026) in Section 7.3

It cannot determine whether correctly specified R/V supervision helps or hurts, or explain the gap causally.

— Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning  (2608.18746 - Wang et al., 19 Aug 2026) in Appendix, Experimental Details, subsection “Reward and value proxy targets”

On the physical AI side, multi-modal temporal alignment, physical validity checking, and coverage-aware sampling (Section~\ref{sec:open_challenges}) require enrichment and indexing techniques with no direct digital AI analogue, and evaluation itself is an open problem: unlike text QA, there is no simple ground truth for whether a retrieved training batch improves a world model.

— UniK: Universal Knowledge Perception for Digital and Physical AI  (2609.23971 - Desai et al., 21 Sep 2026) in Conclusion, 'Open directions'

Divergence is an outcome of applying the update, so the estimand counts every attempted fork; restricted to the $693$ of $720$ that did not collapse, CartPole and Walker are unchanged in sign ($-113.4$ and $-82.1$) and Cheetah becomes unresolved ($-3.9$; $[-17.5,+13.0]$).

— Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation  (2609.10954 - Li et al., 10 Sep 2026) in Abstract; Section 4.1, “A fixed update rule loses return on all three tasks”