Joint Training and Evaluation of World Models in Non-Stationary Environments
Investigate joint training, continual updating, and rigorous evaluation protocols for world models used by large language model-based agents in non-stationary environments, and ascertain the causal impact of these world models on downstream planning reliability.
References
An open problem is how to jointly train, update, and evaluate world models in non-stationary environments, and how to assess their causal impact on downstream planning reliability.
It cannot determine whether correctly specified R/V supervision helps or hurts, or explain the gap causally.
On the physical AI side, multi-modal temporal alignment, physical validity checking, and coverage-aware sampling (Section~\ref{sec:open_challenges}) require enrichment and indexing techniques with no direct digital AI analogue, and evaluation itself is an open problem: unlike text QA, there is no simple ground truth for whether a retrieved training batch improves a world model.
Divergence is an outcome of applying the update, so the estimand counts every attempted fork; restricted to the $693$ of $720$ that did not collapse, CartPole and Walker are unchanged in sign ($-113.4$ and $-82.1$) and Cheetah becomes unresolved ($-3.9$; $[-17.5,+13.0]$).