Learning procedures and reliability thresholds for LLM-based world models
Develop and evaluate training methodologies for large language model–based world models, and establish criteria that determine when these models are sufficiently reliable to improve performance of downstream agents.
References
While prior work has explored LLMs as simulators, experience generators, or planning interfaces \citep{chen2025scalingagentlearningexperience,li2025simulatingenvironmentsreasoningmodels,wu2025rlvrworldtrainingworldmodels,gu2025llmsecretlyworldmodel,wang2025world,he2025pretrained}, it remains unclear how to learn a world model and when it is reliable enough to improve downstream agents.
This shift is important but not yet settled. Latent futures may improve planning, but they can also become inaccurate precisely in rare situations where planning most needs reliability.